Module & Package Architecture

Reviewed & published by Brayan K

Learn how to structure and scale large Python applications. Master the architectural patterns used by Django, FastAPI, Airflow, and enterprise teams to build maintainable codebases with thousands of files.

Part of the free Python course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

Module & Package Architecture for Large Codebases

When your project grows beyond a few files, the MOST important factor for long-term success is structure. A great architecture makes your code:

This lesson teaches you exactly how professional Python teams structure + scale large applications.

🔥 1. Why Python Projects Need Good Architecture

Small scripts can look like this:

But large systems suffer without structure:

⚙️ 2. How Python Imports Actually Work

Understanding the import system is key.

When Python imports a module:

👉 Keep import side-effects to zero. (No DB connections, no heavy computation.)

📦 3. What Is a Package? (And Why It Matters)

A package is a folder containing an __init__.py.

Without __init__.py, Python treats folders as namespace packages.

With it → proper isolated packages.

Use regular packages unless you need distributed namespace packages.

🧱 4. Standard Large-Scale Project Structure

Professional Python projects (Django, Flask, Airflow, FastAPI) use this format:

This is clean because:

This scales to 100K+ lines.

🧩 5. The Layered Architecture (Most Common Design)

FastAPI routers, Flask routes, CLI commands.

Business logic, domain rules, coordination.

Database, cache, filesystem, external APIs.

4. Core / Shared Components

🔌 6. Dependency Direction (The #1 Rule)

High-level modules must NOT depend on low-level modules.

Instead → low-level depends on high-level interfaces.

🔍 7. Avoiding Circular Imports

Circular imports happen when two modules import each other:

# instead of importing across layers
# from domain.user_service import get_user

# restructure using interfaces
from typing import Protocol

class UserGetter(Protocol):
    def get_user(self, user_id: str): ...

# Now domain depends on interface, not concrete implementation
class UserService:
    def __init__(self, user_getter: UserGetter):
        self.user_getter = user_getter
    
    def get_user_profile(self, user_id: str):
        return self.user_getter.get_user(user_id)

print("Interface-based design avoids circular imports!")

# ✅ Expected output:
# Interface-based design avoids circular imports!

🧠 8. Internal Package APIs (__all__ and public API)

Every package should define what it exposes:

# __init__.py example:

# Simulating what would be in separate files
class UserService:
    def get_user(self, user_id):
        return f"User {user_id}"

class User:
    def __init__(self, name):
        self.name = name

# Export only the public API
__all__ = ["UserService", "User"]

# Usage demonstration
service = UserService()
user = User("Alice")
print(service.get_user("123"))
print(f"Created user: {user.name}")

# ✅ Expected output:
# User 123
# Created user: Alice

🧪 Worked Example — Build a Real Package and Watch It Import

Architecture is hard to feel from diagrams. This program creates a genuine two-file package on disk, adds its folder to sys.path, imports it, and then proves three things you have just read about: __init__.py defines the public API, __all__ names it, and imports are cached so a module's code runs exactly once.

Everything the program writes is ordinary Python you would normally type into files by hand — it is written in code here only so you can run the whole thing in one go.

import sys
from pathlib import Path

# ------------------------------------------------------------------
# 1. Create the package:  shop/__init__.py  and  shop/pricing.py
# ------------------------------------------------------------------
pkg = Path("shop")
pkg.mkdir(exist_ok=True)      # a package is a folder with an __init__.py in it

# pricing.py is a plain module: pure logic, no printing, no database calls.
# That is the rule — importing a module must be cheap and must not do work.
(pkg / "pricing.py").write_text('''
VAT_RATE = 0.20

def with_vat(amount):
    return round(amount * (1 + VAT_RATE), 2)

def _internal_helper():
    return "private - other layers must not touch this"
''')

# __init__.py turns the folder into a package AND decides its public API:
# the names other layers are allowed to depend on. Everything else stays
# an implementation detail you are free to rename tomorrow.
(pkg / "__init__.py").write_text('''
from .pricing import with_vat, VAT_RATE

__all__ = ["with_vat", "VAT_RATE"]
''')

# ------------------------------------------------------------------
# 2. Make the folder importable, then import it
# ------------------------------------------------------------------
# Python only imports from directories listed in sys.path. Adding the current
# folder is exactly what running your project from its root does for you.
sys.path.insert(0, str(Path.cwd()))

import shop

print(shop.with_vat(10))   # reached through the package API, not shop.pricing
print(shop.VAT_RATE)

# ------------------------------------------------------------------
# 3. Imports are cached — a module's code executes once per process
# ------------------------------------------------------------------
import shop as shop_again
print(shop is shop_again)              # same module object, not a second read
print("shop" in sys.modules)           # the cache Python checks first
print("shop.pricing" in sys.modules)   # submodules are cached under their full name

# ------------------------------------------------------------------
# 4. __all__ is the published surface
# ------------------------------------------------------------------
print(shop.__all__)
print(hasattr(shop, "_internal_helper"))   # never re-exported, so not visible here

# ✅ Expected output:
# 12.0
# 0.2
# True
# True
# True
# ['with_vat', 'VAT_RATE']
# False

The shop is shop_again line is the one worth remembering. Because Python caches modules in sys.modules, any work you do at the top level of a module happens once and then never again — which is why a database connection or a slow computation at import time is so hard to debug when it goes wrong.

🎯 Your Turn — Publish One Name, Hide the Rest

The module is written for you; you write the __init__.py. Two blanks: the relative import that reaches the module next door, and the single name this package is willing to publish.

import sys
from pathlib import Path

# 🎯 YOUR TURN — finish the package's public API (two ___ blanks below)

pkg = Path("billing")
pkg.mkdir(exist_ok=True)

# This module is already written for you. Note the two functions:
# total() is meant to be public, _debug_dump() is meant to stay internal.
(pkg / "invoices.py").write_text('''
def total(lines):
    return sum(lines)

def _debug_dump(lines):
    return "internal only"
''')

# 👉 blank 1: the module name to import from. The leading dot means
#             "from this same package" — a relative import.
# 👉 blank 2: the one name this package publishes.
(pkg / "__init__.py").write_text('''
from .___ import total

__all__ = ["___"]
''')

sys.path.insert(0, str(Path.cwd()))
import billing

print(billing.total([10, 20, 5]))
print(billing.__all__)
print(hasattr(billing, "_debug_dump"))

# ✅ Expected output:
# 35
# ['total']
# False

A ModuleNotFoundError: No module named 'billing.___' means blank 1 is still a blank — it needs the file name without the .py.

🚀 9. Configuration Architecture

NEVER scatter config constants in files.

Avoid hard-coding secrets, URLs, DB credentials.

# config/settings.py
import os
from pathlib import Path

# In a real settings.py this is how you find the project root:
#     BASE_DIR = Path(__file__).resolve().parent.parent
# __file__ only exists when Python is running an actual file, so this runnable
# demo falls back to the current working directory instead.
BASE_DIR = Path.cwd()

DEBUG = os.getenv("DEBUG", "False") == "True"
DATABASE_URL = os.getenv("DATABASE_URL", "sqlite:///db.sqlite3")
SECRET_KEY = os.getenv("SECRET_KEY", "change-me-in-production")

# Feature flags
ENABLE_ANALYTICS = os.getenv("ENABLE_ANALYTICS", "True") == "True"
ENABLE_CACHING = os.getenv("ENABLE_CACHING", "False") == "True"

print(f"DEBUG mode: {DEBUG}")
print(f"Database: {DATABASE_URL}")
print(f"Analytics enabled: {ENABLE_ANALYTICS}")

# ✅ Expected output:
# DEBUG mode: False
# Database: sqlite:///db.sqlite3
# Analytics enabled: True

📚 10. Module Naming Conventions (Professional Standard)

🧪 11. Testing Architecture

Test mirrors app structure:

Use factories for test data.

Separate unit and integration tests.

🔥 12. Domain-Driven Design (DDD) in Python

DDD is a system-design method focused on modelling business rules, not framework limitations.

Why DDD fits Python:

Goal: Domains should NOT depend on frameworks or external APIs. Only infrastructure depends on domains.

⚙️ 13. Domain Layer Responsibilities

The domain layer should contain:

Objects that have identity across time (e.g., User, Order).

Objects identified by value, not identity (e.g., Price, Email, Coordinates).

Logic that doesn't naturally belong to any one entity.

Business validation, decision logic.

Errors specific to the domain.

What domain must NOT contain:

This keeps the code clean, portable, and testable.

🧱 14. Infrastructure Layer (DB, Cache, External Systems)

Infrastructure is where all "real-world" systems live:

Instead → infrastructure implements interfaces defined in the domain.

# domain/users/interfaces.py
from typing import Protocol

class UserRepository(Protocol):
    def save(self, user): ...
    def find_by_email(self, email: str): ...

# Infrastructure implementation:
# infrastructure/db/user_repository.py
class SQLUserRepository:
    """Implements UserRepository interface"""
    
    def save(self, user):
        # actual SQL logic
        print(f"Saving user to database: {user}")
    
    def find_by_email(self, email: str):
        # actual SQL query
        print(f"Finding user by email: {email}")
        return {"email": email, "name": "Found User"}

# Usage
repo = SQLUserRepository()
repo.save({"name": "Alice", "email": "[email protected]"})
print(repo.find_by_email("[email protected]"))

# ✅ Expected output:
# Saving user to database: {'name': 'Alice', 'email': '[email protected]'}
# Finding user by email: [email protected]
# {'email': '[email protected]', 'name': 'Found User'}

🧩 15. The API Layer

This is the "edge" of the app. Usually contains:

API should talk ONLY to domain services.

🔥 16. The Three Master Architectures for Large Python Projects

There are 3 real-world architectures used by engineering teams once codebases reach 50K+ lines:

✔ 1. Layered Architecture (most common)

Clear vertical layers with strict dependency rules.

✔ 2. Clean/Hexagonal Architecture (enterprise-grade)

Domain is isolated and stable. Adapters wrap external systems. Application orchestrates flows.

✔ 3. Plugin / Modular Monolith (like Django)

Each feature is an "app" with its own mini-architecture inside.

This is the architecture used by:

⚙️ 17. How Django Organises Massive Codebases

Django uses a modular app structure:

Why this matters: If you model your website like this, you can scale to 200+ pages and thousands of functions without losing control.

🔥 18. How FastAPI Organises Modern Backend Projects

FastAPI encourages a clean, layered layout:

Perfect for scalable SaaS backends.

🎉 Final Conclusion

Across the three parts, you've now learned:

You now have the knowledge to architect and scale Python applications from startup MVPs to enterprise systems handling millions of users.

🎯 Mini-Challenge: A Module That Is Both Importable and Runnable

Every professional module can be imported by other code and run directly with python report.py. The if __name__ == "__main__": guard is what keeps those two modes apart — code inside it runs only when the file is the program being run, never when it is imported. Prove it. Only the outline is given.

# 🎯 MINI-CHALLENGE: prove the __main__ guard does not fire on import
#
# 1. import sys, and  from pathlib import Path
#
# 2. Write a file called report.py containing exactly this Python:
#
#        def summarise(numbers):
#            return str(len(numbers)) + " values, total " + str(sum(numbers))
#
#        if __name__ == "__main__":
#            print("running as a script")
#            print(summarise([1, 2, 3]))
#
#    (write it with Path("report.py").write_text('''...'''), the same trick
#     the worked example used)
#
# 3. sys.path.insert(0, str(Path.cwd()))   so Python can find it
# 4. import report
# 5. print("imported, and the __main__ block did not run")
# 6. print(report.summarise([4, 5]))
#
# ✅ Expected output:
# imported, and the __main__ block did not run
# 2 values, total 9

# your code here

If running as a script appears in your output, the guard is missing or misspelled — it is two underscores on each side of both name and main.

📋 Quick Reference — Module Architecture

ConceptWhat it means
__init__.pyMakes a directory a Python package
__all__ = [...]Control what's exported from a module
from . import moduleRelative import within a package
importlib.import_module()Dynamic import at runtime
src/ layoutBest-practice project structure

🎉 Great work! You've completed this lesson.

You can now architect large Python codebases with proper packages, relative imports, and clean module boundaries.

Practice quiz

What traditionally makes a directory a regular Python package?

  • Containing a setup.py file
  • Being listed in sys.path
  • Containing an __init__.py file
  • Ending its name with .pkg

Answer: Containing an __init__.py file. A regular package is a folder containing an __init__.py; without it Python may treat the folder as a namespace package.

After importing a module once, where does Python cache it so it isn't re-executed?

  • sys.modules
  • sys.path
  • __pycache__ only
  • os.environ

Answer: sys.modules. Imported modules are cached in sys.modules; subsequent imports return the cached module object.

What does the __all__ list in a module control?

  • The order methods run in
  • The module's dependencies
  • The Python version required
  • Which names are treated as the module's public API (and exported by 'from module import *')

Answer: Which names are treated as the module's public API (and exported by 'from module import *'). __all__ defines the public API and the names exported on a wildcard import.

What is a circular import?

  • A module that imports itself in a loop forever
  • Two modules that import each other, which can cause runtime errors
  • Importing the same module twice
  • Importing a package without __init__.py

Answer: Two modules that import each other, which can cause runtime errors. When a.py imports b.py and b.py imports a.py, partial initialization can trigger ImportError/AttributeError.

Which is a recommended way to break a circular import?

  • Introduce an interface/Protocol or move shared logic into a core module
  • Duplicate the code in both modules
  • Delete one of the __init__.py files
  • Import everything at the top of every file

Answer: Introduce an interface/Protocol or move shared logic into a core module. Depending on an interface (Protocol) or shared core, or using local imports, removes the cycle.

In layered architecture, what is the #1 dependency-direction rule?

  • High-level modules should depend directly on low-level modules
  • Every module depends on every other module
  • Low-level modules depend on high-level interfaces, not the reverse
  • The API layer holds all the business logic

Answer: Low-level modules depend on high-level interfaces, not the reverse. Dependencies should point toward abstractions: low-level code implements interfaces defined by higher layers.

Which layer should contain the core business logic and domain rules?

  • The API layer
  • The domain/service layer
  • The infrastructure/data layer
  • The configuration layer

Answer: The domain/service layer. Business rules belong in the domain/service layer, kept independent of frameworks and databases.

What is the recommended rule about side effects when a module is imported?

  • Open database connections at import time for speed
  • Run all tests on import
  • Print debug logs on every import
  • Keep import-time side effects to zero (no DB connections, no heavy computation)

Answer: Keep import-time side effects to zero (no DB connections, no heavy computation). Modules execute on first import, so keep that fast and side-effect-free to avoid surprises and slow startup.

Per the lesson's naming conventions, where do business rules belong?

  • models.py
  • services.py
  • routes.py
  • exceptions.py

Answer: services.py. services.py holds business logic; models.py holds data classes and routes.py holds API endpoints.

What is the repository pattern's purpose in the infrastructure layer?

  • To store git history
  • To define HTTP routes
  • To implement a data-access interface defined by the domain, hiding SQL/DB details
  • To replace the domain layer entirely

Answer: To implement a data-access interface defined by the domain, hiding SQL/DB details. A repository implements a domain-defined interface for persistence, so the domain stays free of DB specifics.

Continue this course