Memory Management & Garbage Collection

Reviewed & published by Brayan K

Python may look simple on the surface, but underneath it has a powerful and complex memory management system. To write high-performance Python — whether you're building ML pipelines, backend servers, or tools that process millions of objects — you must understand how Python allocates and frees memory, reference counting, garbage collection cycles, memory fragmentation, and how to track leaks and optimize memory usage.

Part of the free Python course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

🔥 1. How Python Allocates Memory

Python uses a private memory manager (PyMalloc) layered on top of the OS allocator.

There are three layers:

LayerWhat It DoesSpeed
Object-specific allocatorsCustom optimized allocators for ints, lists, dicts, strings⚡ Fastest
Python memory managerHandles small object pools, caches freed memory⚡ Fast
OS-level allocatormalloc(), free() — used for large blocks🐢 Slow

Python tries to avoid calling the OS too often, because OS allocations are slow.

⚙️ 2. Reference Counting — The Core Mechanism

Every Python object has an internal counter: how many references point to it.

You can inspect it:

import sys
x = []
print(sys.getrefcount(x))

# Whenever you do:
y = x
# refcount increases.

# Whenever something goes out of scope:
del y
# refcount decreases.

# ✅ Expected output:
# 2

When it reaches 0, Python immediately frees the memory.

ActionEffect on RefcountExample
Create object+1x = []
Assign to another variable+1y = x
Delete reference-1del y
Leave function scope-1Local variables cleaned up

But it has one big problem…

🧠 3. The Problem: Reference Cycles

a = []
b = []
a.append(b)
b.append(a)

# Both objects reference each other.
# Even if you del a and del b, refcount never reaches 0.

🌀 4. Garbage Collection for Cycles

Python's cyclic garbage collector scans container objects (lists, dicts, sets, classes) to find reference cycles.

GenerationContainsChecked
Gen 0Newest objectsMost frequently
Gen 1Survived 1+ collectionsLess often
Gen 2Long-lived objectsRarely

The cycle detector periodically:

This prevents memory leaks caused by circular references.

🧑‍🏫 Worked example: watch memory actually being freed

Sections 2 to 4 are much easier to believe once you see them happen. The program below gives every object a __del__ method — a hook Python calls at the exact moment the last reference to an object disappears — so each birth and death prints a line. Read the comments, run it, and watch the three steps: refcounting freeing something instantly, a cycle refusing to die, and the cycle collector finishing the job.

# WORKED EXAMPLE: watch Python free memory, in real time.
# Every object below announces its own birth and death, so you can SEE
# refcounting work - and see the one case where it cannot.
import gc
import sys

class Tracked:
    def __init__(self, name):
        self.name = name
        self.partner = None            # used to build a cycle in step 2
        print(f"  + {name} created")

    def __del__(self):
        # __del__ runs the instant the last reference to this object disappears.
        print(f"  - {self.name} freed")

print("STEP 1 - refcounting frees objects immediately")
a = Tracked("solo")
b = a                                  # a SECOND name for the SAME object
# getrefcount always counts its own temporary argument, so subtract 1.
print("  references to solo:", sys.getrefcount(a) - 1)
del a                                  # one reference gone; object survives
print("  deleted 'a' - still alive because 'b' points at it")
del b                                  # last reference gone -> freed on this line
print("  deleted 'b'")

print()
print("STEP 2 - a cycle that refcounting cannot free")
gc.collect()                           # clean slate before the demo
gc.disable()                           # switch the cycle collector OFF to expose the problem
x = Tracked("left")
y = Tracked("right")
x.partner = y                          # left  -> right
y.partner = x                          # right -> left   (now they keep each other alive)
del x
del y
print("  both names deleted - and nothing was freed")

print()
print("STEP 3 - the cycle collector cleans up what refcounting missed")
gc.enable()
gc.collect()                           # finds unreachable cycles and frees them
print("  done")

# ✅ Output:
# STEP 1 - refcounting frees objects immediately
#   + solo created
#   references to solo: 2
#   deleted 'a' - still alive because 'b' points at it
#   - solo freed
#   deleted 'b'
#
# STEP 2 - a cycle that refcounting cannot free
#   + left created
#   + right created
#   both names deleted - and nothing was freed
#
# STEP 3 - the cycle collector cleans up what refcounting missed
#   - left freed
#   - right freed
#   done
#
# Look closely at step 1: "solo freed" prints BEFORE "deleted 'b'", because the
# object dies on the del line itself, not at the end of the program.

# ✅ Expected output:
# STEP 1 - refcounting frees objects immediately
#   + solo created
#   references to solo: 2
#   deleted 'a' - still alive because 'b' points at it
#   - solo freed
#   deleted 'b'
#
# STEP 2 - a cycle that refcounting cannot free
#   + left created
#   + right created
#   both names deleted - and nothing was freed
#
# STEP 3 - the cycle collector cleans up what refcounting missed
#   - left freed
#   - right freed
#   done

⚡ 5. Viewing & Controlling the GC

You can inspect thresholds:

import gc
print(gc.get_threshold())

# Typical default: (700, 10, 10)
# Meaning:
# - collect gen0 every 700 allocations
# - collect gen1 every 10 gen0 collections
# - collect gen2 every 10 gen1 collections

# Force a collection:
gc.collect()

# Disable GC (not recommended unless profiling):
gc.disable()

# ✅ Expected output:
# (700, 10, 10)

📦 6. Memory Fragmentation

Python memory isn't always "returned" to the OS immediately.

ReasonWhat Happens
Freed blocks stay in poolsPython keeps them for reuse
Partially used arenasOS can't reclaim until completely empty
Long-lived objectsCreate "holes" in memory
Extension modulesAllocate outside Python's control

This is why a Python process may appear large even after freeing objects.

Tools like Heapy, tracemalloc, and Pympler help inspect fragmentation.

🧪 7. Detecting Memory Leaks

Python leaks often happen because:

import tracemalloc

tracemalloc.start()

# run code
snapshot = tracemalloc.take_snapshot()
top = snapshot.statistics('lineno')

for item in top[:5]:
    print(item)

# Shows exactly where memory increases.
# pip install objgraph

import objgraph
objgraph.show_growth()

# Helpful for finding leaking objects over time.

🧩 8. Efficient Memory Techniques

✅ Prefer generators over lists

# Use generators:
nums = range(1000000)
gen = (x*x for x in nums)

# Instead of:
# lst = [x*x for x in nums]

# Generators use much less memory!
print(next(gen))
print(next(gen))

# ✅ Expected output:
# 0
# 1

✅ Use __slots__ for large class collections

class Point:
    __slots__ = ("x", "y")
    
    def __init__(self, x, y):
        self.x = x
        self.y = y

# Saves memory by removing per-object dict.
p = Point(1, 2)
print(p.x, p.y)

# ✅ Expected output:
# 1 2

✅ Avoid unnecessary large structures

✅ Reuse objects when possible

Instead of allocating repeatedly inside loops.

✅ Clean up large variables manually

import gc

big_object = [i for i in range(1000000)]
# ... use big_object ...

del big_object
gc.collect()

# Useful in ML/data pipelines.
print("Memory cleaned")

# ✅ Expected output:
# Memory cleaned

🔥 9. Memory & Speed Tradeoffs

Optimising memory may reduce speed, and vice-versa.

Your optimisation depends on whether the bottleneck is:

Profiling tells you which.

🔥 10. How Python Stores Objects in Memory (Deep Internal View)

Every Python object is stored in a PyObject structure that contains:

But different types store additional metadata:

Small integers (from –5 to 256) are pre-allocated and reused → "integer cache".

a = 5
b = 5

# Both point to the same object in memory.
print(a is b)  # True

# But large integers are different objects:
c = 500
d = 500
print(c is d)  # May be False

# ✅ Expected output:
# True
# True

Short strings and identifiers are interned (cached forever) to speed up comparisons.

Python lists are dynamic arrays with over-allocation (extra capacity) to avoid constant resizing.

This saves CPU time but increases memory use.

Use open addressing with "sparse tables".

They resize when the load factor gets too high (~66%).

🧬 11. Arena Allocation (The Deepest Python Memory Detail)

CPython allocates memory in units:

LayerSizePurpose
Arena~256 KBLarge chunk from OS
Pool4 KBFor objects of same size
BlockvariableIndividual object

This explains two things:

🧠 12. Why Lists & Dicts "Grow" in Memory

When you append items, the list grows faster than needed.

Example internal growth pattern:

0 → 4 → 8 → 16 → 25 → 35 → 46 → ...

When keys increase, it resizes to maintain fast O(1) access.

Understanding this helps you design efficient data structures.

🧨 13. Object Lifetimes — From Creation to Deallocation

Python is deterministic for most objects… …but not for cycles.

🔍 14. Memory Leak Patterns in Real Python Code

Here are the 7 most common memory leak patterns seen in production:

cache = []
def add(x):
    cache.append(x)

# This cache grows forever!
for i in range(100):
    add(i)
print(len(cache))

# ✅ Expected output:
# 100
def make():
    big = [1] * 1_000_000
    def inner(): 
        return len(big)
    return inner

# 'big' is kept alive by the closure!
fn = make()
print(fn())

# ✅ Expected output:
# 1000000

🎯 Your turn: the cache that refuses to let go

Pattern 2 above — a cache that never expires — is the leak you are most likely to write yourself. The fix is a weak reference: a reference that lets you find an object but does not keep it alive. Fill in the three ___ blanks below and run it to prove the difference.

# 🎯 YOUR TURN - prove why a plain dict cache leaks
# Replace each ___ below, then press Run.

import gc
import sys
import weakref

class Session:
    def __init__(self, user):
        self.user = user
    def __del__(self):
        print(f"Session({self.user}) freed")

strong_cache = {}          # an ordinary dict - it holds objects alive forever
weak_cache = ___()         # 👉 replace ___ with weakref.WeakValueDictionary

s = Session("ada")
strong_cache["ada"] = s
weak_cache["ada"] = s      # storing in a weak dict does NOT add a reference

# getrefcount counts its own temporary argument, so take one off.
print("references to the session:", sys.getrefcount(s) ___ 1)   # 👉 replace ___ with the operator that removes it

del s                      # drop the only plain name pointing at the Session
gc.___()                   # 👉 replace ___ with the function that runs a cycle collection

print("strong cache holds:", len(strong_cache), "item(s)")
print("weak cache holds:", len(weak_cache), "item(s)")

del strong_cache["ada"]    # now drop the last strong reference
gc.collect()
print("weak cache now holds:", len(weak_cache), "item(s)")

# ✅ Expected output once the blanks are filled in:
# references to the session: 2
# strong cache holds: 1 item(s)
# weak cache holds: 1 item(s)
# Session(ada) freed
# weak cache now holds: 0 item(s)
#
# Read the order carefully: the Session is only freed when the STRONG cache lets go.
# The weak entry then vanishes by itself - that is the leak fix in one line.

🧪 15. Real-World Debugging — Finding a Leak in a Web Server

Imagine you run a FastAPI app, and memory keeps rising.

Step 1: Enable tracemalloc

import tracemalloc
tracemalloc.start()
print("tracemalloc started")

# ✅ Expected output:
# tracemalloc started
import tracemalloc
tracemalloc.start()
# ... run some code ...
s1 = tracemalloc.take_snapshot()
print("Snapshot taken")

# ✅ Expected output:
# Snapshot taken

Step 3: Compare after operations

import tracemalloc
tracemalloc.start()

# Initial state
s1 = tracemalloc.take_snapshot()

# Simulate some allocations
data = [i for i in range(100000)]

# New state
s2 = tracemalloc.take_snapshot()
stats = s2.compare_to(s1, "lineno")
print(stats[:3])

Step 4: Identify dangling references

# pip install objgraph
import objgraph
objgraph.show_most_common_types()

# To see what's holding a reference:
# objgraph.show_backrefs(target_object)

This reveals why something never got garbage-collected.

⚡ 16. Avoiding Fragmentation in Large Applications

Memory fragmentation is a silent killer for long-running apps.

Techniques used by big companies:

Major apps like Instagram and Dropbox use multi-process setups for this exact reason.

📦 17. Working With Huge Data Without Crashing RAM

—you must avoid loading everything at once.

def read_chunks(file, chunk=1024):
    while True:
        data = file.read(chunk)
        if not data:
            break
        yield data

# Usage:
# with open("bigfile.txt") as f:
#     for chunk in read_chunks(f):
#         process(chunk)
def stream_data():
    for i in range(1000000):
        yield {"id": i, "value": i * 2}

# Process one at a time:
for row in stream_data():
    if row["id"] < 5:
        print(row)

# ✅ Expected output:
# {'id': 0, 'value': 0}
# {'id': 1, 'value': 2}
# {'id': 2, 'value': 4}
# {'id': 3, 'value': 6}
# {'id': 4, 'value': 8}

Memory-map files to avoid RAM explosion.

🔧 18. Advanced Optimisation Tools

Compiles Python code to C → 10×–200× speed + fixed memory layout.

JIT compiler for numeric loops.

Alternative Python interpreter with a fast JIT.

Great for long-running loops.

Compiles typed Python into C extensions.

📊 19. Memory & Performance Profiling Workflow (Professional Method)

Here's the exact workflow used in production:

This is the method used by performance engineers at scale.

🧠 20. Python Memory Myths (Corrected)

❌ Myth: Python returns memory to OS when freed

✔ Truth: Almost never — only FULL arenas are returned.

❌ Myth: Garbage collection slows Python

✔ Truth: GC rarely runs unless many container objects are created.

❌ Myth: Variables disappear after function exit

✔ Truth: Closures, globals, caches can keep them alive forever.

❌ Myth: Generators are slower than lists

✔ Truth: They are massively more memory-efficient and usually faster for pipelines.

🎓 21. Final Summary of Python Memory Mastery

This knowledge puts you way above normal Python developers — this is senior-level backend engineer understanding.

🔥 Practical Engineering Summary

How Python Actually Manages Memory (High-Level Recap)

Python uses a multi-layered system:

🧠 The Biggest Causes of Memory Problems in Real Systems

These are the real culprits when you see "Python memory leak".

⚙️ Practical Checklist for Writing Memory-Safe Python

Here is what senior engineers follow:

This checklist alone prevents 90% of real-world problems.

🎯 Mini Challenge: A Cache That Cannot Leak

The single most common Python memory leak is a cache with no upper bound. Fix it properly: give the cache a maximum size and evict the oldest entry whenever it overflows. Python dictionaries remember the order keys were inserted, which is all the bookkeeping you need.

No blanks this time — just the brief and an empty editor. Use the exact print formats in the outline and your output will match the expected output line for line.

# 🎯 MINI-CHALLENGE: a bounded cache

requests = ["a", "b", "c", "a", "d", "e", "b"]
MAX_ITEMS = 3

# 1. Create an empty dict called cache
#
# 2. Loop over requests. For each key:
#      already in cache -> print(f"HIT   {key}")   (three spaces after HIT)
#      not in cache     -> print(f"MISS  {key}")   (two spaces after MISS)
#                          then store cache[key] = key.upper()
#
# 3. Straight after storing a new key, if the cache is now bigger than MAX_ITEMS,
#    evict the oldest entry:
#      - the oldest key is the first one a dict yields:  next(iter(cache))
#      - delete it from the cache
#      - print(f"EVICT {oldest}")
#
# 4. When the loop finishes, print(f"cache: {list(cache)}")

# your code here

# ✅ Expected output:
# MISS  a
# MISS  b
# MISS  c
# HIT   a
# MISS  d
# EVICT a
# MISS  e
# EVICT b
# MISS  b
# EVICT c
# cache: ['d', 'e', 'b']
#
# Note that "a" was evicted even though it had just been a HIT. Fixing that -
# moving a key back to the end on every hit - turns this into a true LRU cache,
# which is exactly what functools.lru_cache does for you.

🔥 Ultimate Takeaways (The "If You Remember Only 10 Things…" List)

Memorise this list — it's the essence of professional Python memory engineering:

If you follow these principles, you'll never struggle with memory issues again.

🎉 Final Conclusion — You Now Understand Memory Like a Senior Engineer

By completing all parts, you now understand:

This is deep Python internals knowledge that most developers never learn.

With this mastery, you're ready to:

📋 Quick Reference — Memory Management

Concept / ToolWhat it does
sys.getrefcount(obj)Check reference count of an object
gc.collect()Manually trigger garbage collection
weakref.ref(obj)Hold reference without preventing GC
__slots__Reduce per-instance memory overhead
tracemallocTrace memory allocations

🎉 Great work! You've completed this lesson.

You now understand reference counting, the garbage collector, and how to avoid memory leaks in long-running Python programs.

Practice quiz

What is Python's primary, most immediate memory-reclamation mechanism?

  • Mark-and-sweep on every allocation
  • Manual free() calls by the programmer
  • Reference counting
  • The operating system

Answer: Reference counting. Every object tracks how many references point to it; when that count hits 0 the memory is freed immediately.

When does reference counting alone FAIL to free objects?

  • When objects form a reference cycle (they refer to each other)
  • When objects are very large
  • When objects are integers
  • When objects are created inside functions

Answer: When objects form a reference cycle (they refer to each other). In a cycle, objects keep each other's refcount above zero, so the cyclic garbage collector is needed to reclaim them.

What does Python's cyclic garbage collector specifically handle?

  • Freeing every object immediately
  • Returning all memory to the OS
  • Counting references
  • Detecting and collecting unreachable reference cycles

Answer: Detecting and collecting unreachable reference cycles. The generational cyclic GC scans container objects to find and free unreachable cycles that refcounting misses.

How many generations does CPython's cyclic garbage collector use?

  • 1
  • 3
  • 2
  • 5

Answer: 3. There are three generations (0, 1, 2); younger generations are collected more frequently.

What is the main benefit of defining __slots__ on a class?

  • It removes the per-instance __dict__, reducing memory per object
  • It makes attribute access raise errors
  • It enables multiple inheritance
  • It speeds up the garbage collector only

Answer: It removes the per-instance __dict__, reducing memory per object. __slots__ drops the per-instance dictionary, cutting per-object memory (often 40 to 70 percent).

Which function lets you inspect an object's current reference count?

  • gc.count()
  • obj.refcount()
  • sys.getrefcount(obj)
  • weakref.count(obj)

Answer: sys.getrefcount(obj). sys.getrefcount(obj) returns the count (slightly inflated by the temporary argument reference).

Which standard-library module traces memory allocations to help find leaks?

  • timeit
  • tracemalloc
  • logging
  • pickle

Answer: tracemalloc. tracemalloc records where allocations happen so you can take and compare snapshots.

Why are small integers like 5 often the SAME object in memory?

  • All integers are singletons
  • Integers are never garbage collected
  • Because they use __slots__
  • CPython pre-allocates and caches small integers (about -5 to 256)

Answer: CPython pre-allocates and caches small integers (about -5 to 256). CPython interns small integers in that range, so a = 5; b = 5 gives 'a is b' as True.

Compared with building a full list, what is the memory advantage of a generator?

  • It stores all items twice for safety
  • It produces items one at a time instead of holding them all in memory
  • It is always faster but uses more RAM
  • It cannot be iterated

Answer: It produces items one at a time instead of holding them all in memory. A generator yields values lazily, so it avoids holding the entire sequence in memory at once.

Why can a Python process stay large even after objects are freed?

  • Python never frees any memory
  • The OS forbids freeing memory
  • Freed memory often stays in Python's pools/arenas rather than returning to the OS
  • Reference counts can go negative

Answer: Freed memory often stays in Python's pools/arenas rather than returning to the OS. CPython keeps freed blocks in pools/arenas for reuse and only returns full, empty arenas to the OS.

Continue this course