Profiling & Optimising Performance
Reviewed & published by Brayan K
High-performance Python isn't about writing "faster code" — it's about finding bottlenecks and eliminating them with scientific precision. You cannot optimise what you do not measure.
Part of the free Python course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
What You'll Learn in This Lesson
- • How to profile CPU usage with cProfile and line_profiler
- • How to measure memory usage with tracemalloc and memory_profiler
- • How to find the real bottleneck in your code (it's rarely where you think)
- • Practical optimisation techniques: caching, algorithm choice, data structures
- • How to write benchmarks using timeit and interpret results correctly
- • Production-level performance patterns used in real Python systems
🔥 1. Why Profiling Matters
Beginners try to "guess" what's slow. Advanced developers measure what's slow.
| Approach | Method | Result |
|---|---|---|
| ❌ Guessing | "This loop looks slow" | Waste time optimizing the wrong code |
| ✔ Profiling | Measure actual execution time | Find and fix real bottlenecks |
20% of code → 80% of runtime
Optimizing the wrong 80% gives no improvement!
⚙️ 2. Timing Functions with time.perf_counter()
import time
start = time.perf_counter()
# code block
end = time.perf_counter()
print("Elapsed:", end - start)Use this for comparing:
- two ways of looping
- two algorithms
- two function implementations
But for full programs, we need real profilers.
🧪 Worked Example: Measure, Don't Guess
Here is a complete benchmark you can run right now. It adds up the same ten million numbers twice — once with a hand-written loop, once with the built-in sum() — and proves two things at once: the answers are identical, and one of them is far quicker. Read every comment before you run it.
import time
N = 10_000_000 # ten million. Underscores are just readability - Python ignores them.
# --- Version 1: a hand-written Python loop -------------------------------
start = time.perf_counter() # perf_counter = the most precise clock available
total_loop = 0
for i in range(N):
total_loop += i # every one of these additions is interpreted Python
loop_seconds = time.perf_counter() - start # subtract the two readings to get elapsed time
# --- Version 2: the built-in sum(), which loops in C ---------------------
start = time.perf_counter()
total_builtin = sum(range(N)) # same work, but the loop lives inside CPython itself
builtin_seconds = time.perf_counter() - start
print("loop total :", total_loop) # 49999995000000
print("sum() total :", total_builtin) # 49999995000000 - identical, as it must be
print("same answer? :", total_loop == total_builtin) # True
# Format to 3 decimals so the numbers stay readable.
print(f"loop took {loop_seconds:.3f}s")
print(f"sum() took {builtin_seconds:.3f}s")
print("sum() was faster:", builtin_seconds < loop_seconds) # True
# Expected output (the two timing lines show different numbers on every
# machine - the run below was roughly 0.63s vs 0.10s):
# loop total : 49999995000000
# sum() total : 49999995000000
# same answer? : True
# loop took 0.630s
# sum() took 0.095s
# sum() was faster: True🧠 3. Profiling With cProfile — The Standard Tool
Run a script with profiling from command line:
python -m cProfile myscript.py⚠️ Requires local Python installation • Download Python
Or profile a specific function in your code:
import cProfile
def slow():
# A loop that runs 5 million times
for _ in range(5_000_000):
pass
# Profile this specific function call
cProfile.run("slow()")
# Output shows: number of calls, total time, time per call| Column | What It Shows |
|---|---|
| ncalls | Number of times function was called |
| tottime | Total time in this function (excluding subcalls) |
| cumtime | Cumulative time (including subcalls) |
📊 4. Making Results Readable With pstats
import cProfile, pstats
profiler = cProfile.Profile()
profiler.enable()
# code…
for _ in range(3_000_000):
pass
profiler.disable()
stats = pstats.Stats(profiler)
stats.sort_stats("tottime").print_stats(10)
# ✅ Expected output:
# 1 function calls in 0.000 seconds
#
# Ordered by: internal time
#
# ncalls tottime percall cumtime percall filename:lineno(function)
# 1 0.000 0.000 0.000 0.000 {method 'disable' of '_lsprof.Profiler' objects}| Sort By | What It Shows | Best For Finding |
|---|---|---|
| "tottime" | Time in function itself | The actual slow functions |
| "cumtime" | Time including all sub-calls | Functions that call slow things |
| "ncalls" | Number of times called | Unexpectedly hot loops |
- "tottime" → slowest total functions
- "cumtime" → functions including subcalls
- "ncalls" → most-called functions
🧠 5. Line-by-Line Profiling With line_profiler
pip install line_profiler@profile
def slow():
total = 0
for i in range(10_000_000):
total += ikernprof -l myscript.py
python myscript.py.lprof⚠️ Command-line tool - requires local setup
Shows exactly which line is slow.
This is invaluable for:
- nested loops
- ML preprocessing
- tight functions
- recursive code
🧩 6. Memory Profiling
pip install memory_profilerfrom memory_profiler import profile
@profile
def load_items():
items = [i for i in range(5_000_000)]
return items| Memory Issue | Symptom | Common Cause |
|---|---|---|
| Memory spike | Sudden +500MB on one line | Loading large dataset at once |
| Memory leak | Memory grows over time | Data accumulating in loops |
| High baseline | Program starts with 200MB+ | Heavy imports (pandas, tensorflow) |
Shows memory growth line-by-line.
- ✔ large lists
- ✔ numpy allocations
- ✔ memory leaks
- ✔ generators vs lists performance
⚡ 7. Real Techniques for Faster Python
1. Use built-ins over manual loops
sum(list) # faster
vs
total = 0
for x in list: total += xBuilt-ins use C-level optimisations.
2. Prefer list comprehensions
[x*x for x in nums] # faster
vs
result = []
for x in nums:
result.append(x*x)3. Use generators for large data
(x*x for x in nums)saves huge amounts of memory.
4. Use numpy for heavy math
Pure Python loops are slow. NumPy performs operations in C — often 50–200× faster.
5. Cache Results With functools.lru_cache
from functools import lru_cache
@lru_cache(None)
def fib(n):
if n < 2: return n
return fib(n-1) + fib(n-2)Transforms slow recursive functions → instant.
6. Use Multiprocessing for CPU
from multiprocessing import Pool
with Pool() as pool:
pool.map(func, items)⚠️ Works best with local Python installation
Runs tasks on multiple cores.
🎯 Your Turn: Cache a Slow Recursion
Without caching, fib(30) calls itself over 2.6 million times. With lru_cache ("least recently used" cache — it remembers what each input returned) the function body runs once per distinct input and never again. Fill in the two blanks and run it: the counter proves how few real calls are left.
import functools
# 🎯 YOUR TURN — replace each ___ and run it.
calls = 0 # counts how many times the function BODY actually executes
@functools.___(maxsize=None) # 👉 the decorator that remembers past results
def fib(n):
global calls
calls += 1 # only reached on a cache MISS
if n < 2:
return n # fib(0) = 0, fib(1) = 1
return fib(n - 1) + fib(n - ___) # 👉 the classic rule: the two previous numbers
print("fib(30) =", fib(30))
print("function bodies actually run:", calls)
print(fib.cache_info()) # hits / misses / how many results are stored
# ✅ Expected output:
# fib(30) = 832040
# function bodies actually run: 31
# CacheInfo(hits=28, misses=31, maxsize=None, currsize=31)🏎️ 8. Avoiding the Biggest Performance Mistakes
| ❌ Mistake | Why It's Slow | ✅ Better Approach |
|---|---|---|
| Unnecessary list copies | Copies entire list in memory | Use slices or itertools |
| Python loops for math | Interpreted = slow | NumPy vectorized operations |
| String concatenation in loop | Creates new string each time | Use ''.join(list) |
| Opening files repeatedly | Disk I/O is expensive | Open once, read/write many |
| Blocking I/O in async | Blocks the entire event loop | Use run_in_executor() |
Summary of Common Traps:
- ❌ Unnecessary list copies
- ❌ Using Python loops for math
- ❌ Excessive string concatenation
- ❌ Opening files repeatedly
- ❌ Overuse of classes when simple functions work
- ❌ Blocking I/O in async code
🧪 9. Real-World Example: Speeding Up JSON Parsing
import json
data = [json.loads(x) for x in lines]import orjson
data = [orjson.loads(x) for x in lines]⚠️ Requires: pip install orjson • Download Python
orjson is 5–20× faster than Python's JSON parser.
🎉 Conclusion
By mastering profiling and optimisation, you gain the ability to:
✔ Build faster APIs and scripts
✔ Save CPU & memory in production
✔ Think like a performance engineer
Performance comes from measure → diagnose → optimise, not guessing.
🎯 Mini-Challenge: Fix the Slow Lookup
Checking if item in some_list scans the list from the start every single time. Checking if item in some_set jumps straight to the answer. Your job is to measure the difference yourself, and to prove the faster version still gives the same answer. Only the brief is given below — the code is yours.
# 🎯 MINI-CHALLENGE: make a lookup dramatically faster
#
# 1. import time
# 2. order_ids = list(range(0, 20000, 2)) # every EVEN id from 0 to 19998
# to_check = list(range(5000)) # the ids you must look up
# 3. Write count_known_list(ids): count how many of ids are in order_ids
# (search the LIST directly - this is the slow version)
# 4. Write count_known_set(ids): convert order_ids to a set ONCE, then count
# the same way against the set
# 5. Time each call with time.perf_counter() the way the worked example did
# 6. Print both counts, then both timings
#
# ✅ The counts are the part you can check exactly - they must match:
# list version: 2500
# set version: 2500
#
# The two timings differ on every machine, so no fixed number is given here.
# What matters is the gap: the set version should finish many times faster.
# your code here📋 Quick Reference — Profiling & Performance
| Tool / Syntax | What it does |
|---|---|
| cProfile.run('fn()') | Profile function call counts and time |
| timeit.timeit('expr', number=1000) | Benchmark small code snippets |
| line_profiler | Profile line-by-line execution time |
| memory_profiler | Track memory usage per line |
| __slots__ | Reduce class memory footprint |
🎉 Great work! You've completed this lesson.
You now know how to measure, diagnose, and fix performance bottlenecks — the professional workflow every senior engineer uses.
Practice quiz
What is the core idea behind profiling before optimising?
- Always optimise the longest function first
- Rewrite everything in C
- Measure what's slow instead of guessing
- Add more print statements
Answer: Measure what's slow instead of guessing. You cannot optimise what you do not measure — profiling replaces guessing with data.
Which standard-library tool profiles function call counts and time?
- cProfile
- tracemalloc
- asyncio
- logging
Answer: cProfile. cProfile records every call, how long each takes, and how many times it runs.
In cProfile output, what does 'tottime' measure?
- Total time including all sub-calls
- Number of times the function was called
- Total program runtime
- Time in the function itself, excluding sub-calls
Answer: Time in the function itself, excluding sub-calls. tottime is time spent in the function body alone; cumtime includes time in sub-calls.
Which sort key best reveals the actually-slow functions?
- "ncalls"
- "tottime"
- "cumtime"
- "name"
Answer: "tottime". Sort by tottime to find where time is really spent; then use cumtime to trace callers.
What does functools.lru_cache do to a recursive fib function?
- Caches results so repeated calls are instant
- Slows it down by caching
- Runs it on multiple cores
- Converts it to a loop
Answer: Caches results so repeated calls are instant. lru_cache memoizes results, turning exponential recursion into near-instant lookups.
What does timeit.timeit('expr', number=1000) return?
- The result of the expression
- A profile report object
- The total time (a float) to run it that many times
- The number of calls
Answer: The total time (a float) to run it that many times. timeit returns the total elapsed time as a float for the given number of executions.
Why use a generator expression (x*x for x in nums) over a list for large data?
- It is always faster to build
- It saves memory by yielding items lazily
- It sorts the data
- It runs in parallel
Answer: It saves memory by yielding items lazily. Generators produce items one at a time instead of building the whole list in memory.
What is the recommended fix for slow string concatenation in a loop?
- Use s = s + x each time
- Use a global string
- Use print() to build it
- Use ''.join(list_of_strings)
Answer: Use ''.join(list_of_strings). ''.join builds the result in one pass; repeated + creates a new string object every iteration.
What does the 80/20 rule say about performance?
- 80% of code runs in 20% of the time
- Roughly 20% of code accounts for ~80% of runtime
- Optimise 80% of functions
- 20% of bugs cause 80% of crashes
Answer: Roughly 20% of code accounts for ~80% of runtime. A small fraction of code dominates runtime, so optimising the wrong 80% gives no improvement.
Which tool tracks memory usage line-by-line?
- cProfile
- timeit
- memory_profiler
- pstats
Answer: memory_profiler. memory_profiler shows memory growth per line; cProfile and timeit measure time, not memory.
Continue this course
- Previous: Parallelism with concurrent.futures
- Next: Memory Management & Garbage Collection Internals — How Python allocates, tracks, and frees memory under the hood
- Quick reference: Python cheat sheet
- From the blog: Python Decorators: A Practical Guide