Module & Package Architecture
Reviewed & published by Brayan K
Learn how to structure and scale large Python applications. Master the architectural patterns used by Django, FastAPI, Airflow, and enterprise teams to build maintainable codebases with thousands of files.
Part of the free Python course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
Module & Package Architecture for Large Codebases
When your project grows beyond a few files, the MOST important factor for long-term success is structure. A great architecture makes your code:
- ✔ easier to understand
- ✔ easier to test
- ✔ easier to extend
- ✔ easier to debug
- ✔ easier to onboard new developers
- ✔ scale to thousands of files
This lesson teaches you exactly how professional Python teams structure + scale large applications.
🔥 1. Why Python Projects Need Good Architecture
Small scripts can look like this:
But large systems suffer without structure:
- ❌ circular imports
- ❌ duplicated logic
- ❌ unclear module responsibilities
- ❌ "god files" with 5,000–20,000 lines
- ❌ impossible navigation
- ❌ hard-to-test logic
- ❌ breaking one feature breaks everything
- modular isolation
- clear domain boundaries
- strong naming conventions
- proper folder hierarchy
- dependency direction
- package-level APIs
⚙️ 2. How Python Imports Actually Work
Understanding the import system is key.
When Python imports a module:
- It searches directories from sys.path
- Looks for a package folder OR .py file
- Executes the module once
- Caches it in sys.modules
- ✔ imports are cached
- ✔ circular imports cause runtime errors
- ✔ module execution on import can be expensive
👉 Keep import side-effects to zero. (No DB connections, no heavy computation.)
📦 3. What Is a Package? (And Why It Matters)
A package is a folder containing an __init__.py.
Without __init__.py, Python treats folders as namespace packages.
With it → proper isolated packages.
Use regular packages unless you need distributed namespace packages.
🧱 4. Standard Large-Scale Project Structure
Professional Python projects (Django, Flask, Airflow, FastAPI) use this format:
This is clean because:
- domain layer contains business logic
- infrastructure contains external systems
- api layer exposes HTTP / CLI interface
- core holds shared primitives
- config holds settings
This scales to 100K+ lines.
🧩 5. The Layered Architecture (Most Common Design)
FastAPI routers, Flask routes, CLI commands.
Business logic, domain rules, coordination.
Database, cache, filesystem, external APIs.
4. Core / Shared Components
- abstractions
- ✔ isolates business logic
- ✔ minimizes circular imports
- ✔ pluggable infrastructure (swap DB easily)
- ✔ easier unit testing
🔌 6. Dependency Direction (The #1 Rule)
High-level modules must NOT depend on low-level modules.
Instead → low-level depends on high-level interfaces.
- ✔ dependency injection
- ✔ testing with mock DB
- ✔ clean separation
- ✔ avoiding circular imports
🔍 7. Avoiding Circular Imports
Circular imports happen when two modules import each other:
- ✔ moving shared logic into core
- ✔ using local imports inside functions
- ✔ introducing interfaces
- ✔ separating pure logic from IO
# instead of importing across layers
# from domain.user_service import get_user
# restructure using interfaces
from typing import Protocol
class UserGetter(Protocol):
def get_user(self, user_id: str): ...
# Now domain depends on interface, not concrete implementation
class UserService:
def __init__(self, user_getter: UserGetter):
self.user_getter = user_getter
def get_user_profile(self, user_id: str):
return self.user_getter.get_user(user_id)
print("Interface-based design avoids circular imports!")
# ✅ Expected output:
# Interface-based design avoids circular imports!🧠 8. Internal Package APIs (__all__ and public API)
Every package should define what it exposes:
# __init__.py example:
# Simulating what would be in separate files
class UserService:
def get_user(self, user_id):
return f"User {user_id}"
class User:
def __init__(self, name):
self.name = name
# Export only the public API
__all__ = ["UserService", "User"]
# Usage demonstration
service = UserService()
user = User("Alice")
print(service.get_user("123"))
print(f"Created user: {user.name}")
# ✅ Expected output:
# User 123
# Created user: Alice- ✔ clean external imports
- ✔ hides internal details
- ✔ stable API for other modules
🧪 Worked Example — Build a Real Package and Watch It Import
Architecture is hard to feel from diagrams. This program creates a genuine two-file package on disk, adds its folder to sys.path, imports it, and then proves three things you have just read about: __init__.py defines the public API, __all__ names it, and imports are cached so a module's code runs exactly once.
Everything the program writes is ordinary Python you would normally type into files by hand — it is written in code here only so you can run the whole thing in one go.
import sys
from pathlib import Path
# ------------------------------------------------------------------
# 1. Create the package: shop/__init__.py and shop/pricing.py
# ------------------------------------------------------------------
pkg = Path("shop")
pkg.mkdir(exist_ok=True) # a package is a folder with an __init__.py in it
# pricing.py is a plain module: pure logic, no printing, no database calls.
# That is the rule — importing a module must be cheap and must not do work.
(pkg / "pricing.py").write_text('''
VAT_RATE = 0.20
def with_vat(amount):
return round(amount * (1 + VAT_RATE), 2)
def _internal_helper():
return "private - other layers must not touch this"
''')
# __init__.py turns the folder into a package AND decides its public API:
# the names other layers are allowed to depend on. Everything else stays
# an implementation detail you are free to rename tomorrow.
(pkg / "__init__.py").write_text('''
from .pricing import with_vat, VAT_RATE
__all__ = ["with_vat", "VAT_RATE"]
''')
# ------------------------------------------------------------------
# 2. Make the folder importable, then import it
# ------------------------------------------------------------------
# Python only imports from directories listed in sys.path. Adding the current
# folder is exactly what running your project from its root does for you.
sys.path.insert(0, str(Path.cwd()))
import shop
print(shop.with_vat(10)) # reached through the package API, not shop.pricing
print(shop.VAT_RATE)
# ------------------------------------------------------------------
# 3. Imports are cached — a module's code executes once per process
# ------------------------------------------------------------------
import shop as shop_again
print(shop is shop_again) # same module object, not a second read
print("shop" in sys.modules) # the cache Python checks first
print("shop.pricing" in sys.modules) # submodules are cached under their full name
# ------------------------------------------------------------------
# 4. __all__ is the published surface
# ------------------------------------------------------------------
print(shop.__all__)
print(hasattr(shop, "_internal_helper")) # never re-exported, so not visible here
# ✅ Expected output:
# 12.0
# 0.2
# True
# True
# True
# ['with_vat', 'VAT_RATE']
# FalseThe shop is shop_again line is the one worth remembering. Because Python caches modules in sys.modules, any work you do at the top level of a module happens once and then never again — which is why a database connection or a slow computation at import time is so hard to debug when it goes wrong.
🎯 Your Turn — Publish One Name, Hide the Rest
The module is written for you; you write the __init__.py. Two blanks: the relative import that reaches the module next door, and the single name this package is willing to publish.
import sys
from pathlib import Path
# 🎯 YOUR TURN — finish the package's public API (two ___ blanks below)
pkg = Path("billing")
pkg.mkdir(exist_ok=True)
# This module is already written for you. Note the two functions:
# total() is meant to be public, _debug_dump() is meant to stay internal.
(pkg / "invoices.py").write_text('''
def total(lines):
return sum(lines)
def _debug_dump(lines):
return "internal only"
''')
# 👉 blank 1: the module name to import from. The leading dot means
# "from this same package" — a relative import.
# 👉 blank 2: the one name this package publishes.
(pkg / "__init__.py").write_text('''
from .___ import total
__all__ = ["___"]
''')
sys.path.insert(0, str(Path.cwd()))
import billing
print(billing.total([10, 20, 5]))
print(billing.__all__)
print(hasattr(billing, "_debug_dump"))
# ✅ Expected output:
# 35
# ['total']
# FalseA ModuleNotFoundError: No module named 'billing.___' means blank 1 is still a blank — it needs the file name without the .py.
🚀 9. Configuration Architecture
NEVER scatter config constants in files.
- settings_local.py
- environment variables
Avoid hard-coding secrets, URLs, DB credentials.
# config/settings.py
import os
from pathlib import Path
# In a real settings.py this is how you find the project root:
# BASE_DIR = Path(__file__).resolve().parent.parent
# __file__ only exists when Python is running an actual file, so this runnable
# demo falls back to the current working directory instead.
BASE_DIR = Path.cwd()
DEBUG = os.getenv("DEBUG", "False") == "True"
DATABASE_URL = os.getenv("DATABASE_URL", "sqlite:///db.sqlite3")
SECRET_KEY = os.getenv("SECRET_KEY", "change-me-in-production")
# Feature flags
ENABLE_ANALYTICS = os.getenv("ENABLE_ANALYTICS", "True") == "True"
ENABLE_CACHING = os.getenv("ENABLE_CACHING", "False") == "True"
print(f"DEBUG mode: {DEBUG}")
print(f"Database: {DATABASE_URL}")
print(f"Analytics enabled: {ENABLE_ANALYTICS}")
# ✅ Expected output:
# DEBUG mode: False
# Database: sqlite:///db.sqlite3
# Analytics enabled: True📚 10. Module Naming Conventions (Professional Standard)
- models.py — classes representing data
- services.py — business logic
- repository.py — DB access
- routes.py — API routes
- tasks.py — background jobs
- exceptions.py — error definitions
- utils.py — ONLY for generic helpers
- ❌ helpers_mixed.py
- ❌ random_functions.py
🧪 11. Testing Architecture
Test mirrors app structure:
Use factories for test data.
Separate unit and integration tests.
- domain tested heavily with unit tests
- infrastructure tested with mocks
- API tested with integration tests
🔥 12. Domain-Driven Design (DDD) in Python
DDD is a system-design method focused on modelling business rules, not framework limitations.
- ✔ Code structure mirrors business structure
- ✔ Each domain is isolated
- ✔ Logic belongs to the domain, not to API or DB layers
- ✔ Domain objects represent real-world concepts
Why DDD fits Python:
- Python's dynamic nature simplifies domain modelling
- Dataclasses + type hints make models clean
- Package isolation prevents circular imports
- Encourages clean business rules without tech dependencies
Goal: Domains should NOT depend on frameworks or external APIs. Only infrastructure depends on domains.
⚙️ 13. Domain Layer Responsibilities
The domain layer should contain:
Objects that have identity across time (e.g., User, Order).
Objects identified by value, not identity (e.g., Price, Email, Coordinates).
Logic that doesn't naturally belong to any one entity.
Business validation, decision logic.
Errors specific to the domain.
What domain must NOT contain:
- ❌ database code
- ❌ API framework code
- ❌ external library calls
- ❌ logging / caching
- ❌ infrastructure details
This keeps the code clean, portable, and testable.
🧱 14. Infrastructure Layer (DB, Cache, External Systems)
Infrastructure is where all "real-world" systems live:
- actual SQL queries
- Redis connections
- HTTP clients
- integrations with Stripe, AWS, etc.
- ❌ contain business rules
- ❌ call domain services
Instead → infrastructure implements interfaces defined in the domain.
# domain/users/interfaces.py
from typing import Protocol
class UserRepository(Protocol):
def save(self, user): ...
def find_by_email(self, email: str): ...
# Infrastructure implementation:
# infrastructure/db/user_repository.py
class SQLUserRepository:
"""Implements UserRepository interface"""
def save(self, user):
# actual SQL logic
print(f"Saving user to database: {user}")
def find_by_email(self, email: str):
# actual SQL query
print(f"Finding user by email: {email}")
return {"email": email, "name": "Found User"}
# Usage
repo = SQLUserRepository()
repo.save({"name": "Alice", "email": "[email protected]"})
print(repo.find_by_email("[email protected]"))
# ✅ Expected output:
# Saving user to database: {'name': 'Alice', 'email': '[email protected]'}
# Finding user by email: [email protected]
# {'email': '[email protected]', 'name': 'Found User'}🧩 15. The API Layer
This is the "edge" of the app. Usually contains:
- Flask blueprints
- FastAPI routers
- Django views
- CLI commands (Click, Typer)
- Websocket handlers
- ✔ parsing requests
- ✔ converting domain errors → HTTP codes
- ✔ authentication
- ✔ response formatting
- ❌ contain business decisions
- ❌ contain SQL queries
- ❌ talk directly to infrastructure
- ❌ hold domain logic
API should talk ONLY to domain services.
🔥 16. The Three Master Architectures for Large Python Projects
There are 3 real-world architectures used by engineering teams once codebases reach 50K+ lines:
✔ 1. Layered Architecture (most common)
Clear vertical layers with strict dependency rules.
✔ 2. Clean/Hexagonal Architecture (enterprise-grade)
Domain is isolated and stable. Adapters wrap external systems. Application orchestrates flows.
✔ 3. Plugin / Modular Monolith (like Django)
Each feature is an "app" with its own mini-architecture inside.
This is the architecture used by:
- Many enterprise monoliths
⚙️ 17. How Django Organises Massive Codebases
Django uses a modular app structure:
- ✓ easy testing
- ✓ independent teams
- ✓ plugin marketplace (reusable apps)
Why this matters: If you model your website like this, you can scale to 200+ pages and thousands of functions without losing control.
🔥 18. How FastAPI Organises Modern Backend Projects
FastAPI encourages a clean, layered layout:
- Fast startup
- Domain + repository pattern
- Event-driven hooks
Perfect for scalable SaaS backends.
🎉 Final Conclusion
Across the three parts, you've now learned:
- ✔ Clean architecture
- ✔ Domain-driven design
- ✔ Layered module organization
- ✔ Plugin-based modular monolith
- ✔ Event-driven communication
- ✔ Dependency inversion
- ✔ Container-based dependency wiring
- ✔ API versioning for long-term stability
- ✔ Scaling to hundreds of modules
- ✔ How real companies structure Python systems
You now have the knowledge to architect and scale Python applications from startup MVPs to enterprise systems handling millions of users.
🎯 Mini-Challenge: A Module That Is Both Importable and Runnable
Every professional module can be imported by other code and run directly with python report.py. The if __name__ == "__main__": guard is what keeps those two modes apart — code inside it runs only when the file is the program being run, never when it is imported. Prove it. Only the outline is given.
# 🎯 MINI-CHALLENGE: prove the __main__ guard does not fire on import
#
# 1. import sys, and from pathlib import Path
#
# 2. Write a file called report.py containing exactly this Python:
#
# def summarise(numbers):
# return str(len(numbers)) + " values, total " + str(sum(numbers))
#
# if __name__ == "__main__":
# print("running as a script")
# print(summarise([1, 2, 3]))
#
# (write it with Path("report.py").write_text('''...'''), the same trick
# the worked example used)
#
# 3. sys.path.insert(0, str(Path.cwd())) so Python can find it
# 4. import report
# 5. print("imported, and the __main__ block did not run")
# 6. print(report.summarise([4, 5]))
#
# ✅ Expected output:
# imported, and the __main__ block did not run
# 2 values, total 9
# your code hereIf running as a script appears in your output, the guard is missing or misspelled — it is two underscores on each side of both name and main.
📋 Quick Reference — Module Architecture
| Concept | What it means |
|---|---|
| __init__.py | Makes a directory a Python package |
| __all__ = [...] | Control what's exported from a module |
| from . import module | Relative import within a package |
| importlib.import_module() | Dynamic import at runtime |
| src/ layout | Best-practice project structure |
🎉 Great work! You've completed this lesson.
You can now architect large Python codebases with proper packages, relative imports, and clean module boundaries.
Practice quiz
What traditionally makes a directory a regular Python package?
- Containing a setup.py file
- Being listed in sys.path
- Containing an __init__.py file
- Ending its name with .pkg
Answer: Containing an __init__.py file. A regular package is a folder containing an __init__.py; without it Python may treat the folder as a namespace package.
After importing a module once, where does Python cache it so it isn't re-executed?
- sys.modules
- sys.path
- __pycache__ only
- os.environ
Answer: sys.modules. Imported modules are cached in sys.modules; subsequent imports return the cached module object.
What does the __all__ list in a module control?
- The order methods run in
- The module's dependencies
- The Python version required
- Which names are treated as the module's public API (and exported by 'from module import *')
Answer: Which names are treated as the module's public API (and exported by 'from module import *'). __all__ defines the public API and the names exported on a wildcard import.
What is a circular import?
- A module that imports itself in a loop forever
- Two modules that import each other, which can cause runtime errors
- Importing the same module twice
- Importing a package without __init__.py
Answer: Two modules that import each other, which can cause runtime errors. When a.py imports b.py and b.py imports a.py, partial initialization can trigger ImportError/AttributeError.
Which is a recommended way to break a circular import?
- Introduce an interface/Protocol or move shared logic into a core module
- Duplicate the code in both modules
- Delete one of the __init__.py files
- Import everything at the top of every file
Answer: Introduce an interface/Protocol or move shared logic into a core module. Depending on an interface (Protocol) or shared core, or using local imports, removes the cycle.
In layered architecture, what is the #1 dependency-direction rule?
- High-level modules should depend directly on low-level modules
- Every module depends on every other module
- Low-level modules depend on high-level interfaces, not the reverse
- The API layer holds all the business logic
Answer: Low-level modules depend on high-level interfaces, not the reverse. Dependencies should point toward abstractions: low-level code implements interfaces defined by higher layers.
Which layer should contain the core business logic and domain rules?
- The API layer
- The domain/service layer
- The infrastructure/data layer
- The configuration layer
Answer: The domain/service layer. Business rules belong in the domain/service layer, kept independent of frameworks and databases.
What is the recommended rule about side effects when a module is imported?
- Open database connections at import time for speed
- Run all tests on import
- Print debug logs on every import
- Keep import-time side effects to zero (no DB connections, no heavy computation)
Answer: Keep import-time side effects to zero (no DB connections, no heavy computation). Modules execute on first import, so keep that fast and side-effect-free to avoid surprises and slow startup.
Per the lesson's naming conventions, where do business rules belong?
- models.py
- services.py
- routes.py
- exceptions.py
Answer: services.py. services.py holds business logic; models.py holds data classes and routes.py holds API endpoints.
What is the repository pattern's purpose in the infrastructure layer?
- To store git history
- To define HTTP routes
- To implement a data-access interface defined by the domain, hiding SQL/DB details
- To replace the domain layer entirely
Answer: To implement a data-access interface defined by the domain, hiding SQL/DB details. A repository implements a domain-defined interface for persistence, so the domain stays free of DB specifics.
Continue this course
- Previous: Design Patterns in Python (Singleton, Factory, Strategy)
- Next: Logging, Debugging & Error Handling at Scale — Add structured logging, use pdb, and handle errors in production
- Quick reference: Python cheat sheet