Profiling & Optimising Performance

Reviewed & published by Brayan K

High-performance Python isn't about writing "faster code" — it's about finding bottlenecks and eliminating them with scientific precision. You cannot optimise what you do not measure.

Part of the free Python course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

🔥 1. Why Profiling Matters

Beginners try to "guess" what's slow. Advanced developers measure what's slow.

ApproachMethodResult
❌ Guessing"This loop looks slow"Waste time optimizing the wrong code
✔ ProfilingMeasure actual execution timeFind and fix real bottlenecks

20% of code → 80% of runtime

Optimizing the wrong 80% gives no improvement!

⚙️ 2. Timing Functions with time.perf_counter()

import time

start = time.perf_counter()
# code block
end = time.perf_counter()

print("Elapsed:", end - start)

Use this for comparing:

But for full programs, we need real profilers.

🧪 Worked Example: Measure, Don't Guess

Here is a complete benchmark you can run right now. It adds up the same ten million numbers twice — once with a hand-written loop, once with the built-in sum() — and proves two things at once: the answers are identical, and one of them is far quicker. Read every comment before you run it.

import time

N = 10_000_000    # ten million. Underscores are just readability - Python ignores them.

# --- Version 1: a hand-written Python loop -------------------------------
start = time.perf_counter()      # perf_counter = the most precise clock available
total_loop = 0
for i in range(N):
    total_loop += i              # every one of these additions is interpreted Python
loop_seconds = time.perf_counter() - start   # subtract the two readings to get elapsed time

# --- Version 2: the built-in sum(), which loops in C ---------------------
start = time.perf_counter()
total_builtin = sum(range(N))    # same work, but the loop lives inside CPython itself
builtin_seconds = time.perf_counter() - start

print("loop total   :", total_loop)      # 49999995000000
print("sum() total  :", total_builtin)   # 49999995000000 - identical, as it must be
print("same answer? :", total_loop == total_builtin)   # True

# Format to 3 decimals so the numbers stay readable.
print(f"loop  took {loop_seconds:.3f}s")
print(f"sum() took {builtin_seconds:.3f}s")
print("sum() was faster:", builtin_seconds < loop_seconds)   # True

# Expected output (the two timing lines show different numbers on every
# machine - the run below was roughly 0.63s vs 0.10s):
# loop total   : 49999995000000
# sum() total  : 49999995000000
# same answer? : True
# loop  took 0.630s
# sum() took 0.095s
# sum() was faster: True

🧠 3. Profiling With cProfile — The Standard Tool

Run a script with profiling from command line:

python -m cProfile myscript.py

⚠️ Requires local Python installation • Download Python

Or profile a specific function in your code:

import cProfile

def slow():
    # A loop that runs 5 million times
    for _ in range(5_000_000):
        pass

# Profile this specific function call
cProfile.run("slow()")
# Output shows: number of calls, total time, time per call
ColumnWhat It Shows
ncallsNumber of times function was called
tottimeTotal time in this function (excluding subcalls)
cumtimeCumulative time (including subcalls)

📊 4. Making Results Readable With pstats

import cProfile, pstats

profiler = cProfile.Profile()
profiler.enable()

# code…
for _ in range(3_000_000):
    pass

profiler.disable()
stats = pstats.Stats(profiler)
stats.sort_stats("tottime").print_stats(10)

# ✅ Expected output:
#          1 function calls in 0.000 seconds
#
#    Ordered by: internal time
#
#    ncalls  tottime  percall  cumtime  percall filename:lineno(function)
#         1    0.000    0.000    0.000    0.000 {method 'disable' of '_lsprof.Profiler' objects}
Sort ByWhat It ShowsBest For Finding
"tottime"Time in function itselfThe actual slow functions
"cumtime"Time including all sub-callsFunctions that call slow things
"ncalls"Number of times calledUnexpectedly hot loops

🧠 5. Line-by-Line Profiling With line_profiler

pip install line_profiler
@profile
def slow():
    total = 0
    for i in range(10_000_000):
        total += i
kernprof -l myscript.py
python myscript.py.lprof

⚠️ Command-line tool - requires local setup

Shows exactly which line is slow.

This is invaluable for:

🧩 6. Memory Profiling

pip install memory_profiler
from memory_profiler import profile

@profile
def load_items():
    items = [i for i in range(5_000_000)]
    return items
Memory IssueSymptomCommon Cause
Memory spikeSudden +500MB on one lineLoading large dataset at once
Memory leakMemory grows over timeData accumulating in loops
High baselineProgram starts with 200MB+Heavy imports (pandas, tensorflow)

Shows memory growth line-by-line.

⚡ 7. Real Techniques for Faster Python

1. Use built-ins over manual loops

sum(list)     # faster

vs

total = 0
for x in list: total += x

Built-ins use C-level optimisations.

2. Prefer list comprehensions

[x*x for x in nums]  # faster

vs

result = []
for x in nums:
    result.append(x*x)

3. Use generators for large data

(x*x for x in nums)

saves huge amounts of memory.

4. Use numpy for heavy math

Pure Python loops are slow. NumPy performs operations in C — often 50–200× faster.

5. Cache Results With functools.lru_cache

from functools import lru_cache

@lru_cache(None)
def fib(n):
    if n < 2: return n
    return fib(n-1) + fib(n-2)

Transforms slow recursive functions → instant.

6. Use Multiprocessing for CPU

from multiprocessing import Pool

with Pool() as pool:
    pool.map(func, items)

⚠️ Works best with local Python installation

Runs tasks on multiple cores.

🎯 Your Turn: Cache a Slow Recursion

Without caching, fib(30) calls itself over 2.6 million times. With lru_cache ("least recently used" cache — it remembers what each input returned) the function body runs once per distinct input and never again. Fill in the two blanks and run it: the counter proves how few real calls are left.

import functools

# 🎯 YOUR TURN — replace each ___ and run it.

calls = 0    # counts how many times the function BODY actually executes

@functools.___(maxsize=None)    # 👉 the decorator that remembers past results
def fib(n):
    global calls
    calls += 1                  # only reached on a cache MISS
    if n < 2:
        return n                # fib(0) = 0, fib(1) = 1
    return fib(n - 1) + fib(n - ___)   # 👉 the classic rule: the two previous numbers

print("fib(30) =", fib(30))
print("function bodies actually run:", calls)
print(fib.cache_info())         # hits / misses / how many results are stored

# ✅ Expected output:
# fib(30) = 832040
# function bodies actually run: 31
# CacheInfo(hits=28, misses=31, maxsize=None, currsize=31)

🏎️ 8. Avoiding the Biggest Performance Mistakes

❌ MistakeWhy It's Slow✅ Better Approach
Unnecessary list copiesCopies entire list in memoryUse slices or itertools
Python loops for mathInterpreted = slowNumPy vectorized operations
String concatenation in loopCreates new string each timeUse ''.join(list)
Opening files repeatedlyDisk I/O is expensiveOpen once, read/write many
Blocking I/O in asyncBlocks the entire event loopUse run_in_executor()

Summary of Common Traps:

🧪 9. Real-World Example: Speeding Up JSON Parsing

import json

data = [json.loads(x) for x in lines]
import orjson

data = [orjson.loads(x) for x in lines]

⚠️ Requires: pip install orjson • Download Python

orjson is 5–20× faster than Python's JSON parser.

🎉 Conclusion

By mastering profiling and optimisation, you gain the ability to:

✔ Build faster APIs and scripts

✔ Save CPU & memory in production

✔ Think like a performance engineer

Performance comes from measure → diagnose → optimise, not guessing.

🎯 Mini-Challenge: Fix the Slow Lookup

Checking if item in some_list scans the list from the start every single time. Checking if item in some_set jumps straight to the answer. Your job is to measure the difference yourself, and to prove the faster version still gives the same answer. Only the brief is given below — the code is yours.

# 🎯 MINI-CHALLENGE: make a lookup dramatically faster
#
# 1. import time
# 2. order_ids = list(range(0, 20000, 2))   # every EVEN id from 0 to 19998
#    to_check  = list(range(5000))          # the ids you must look up
# 3. Write count_known_list(ids): count how many of ids are in order_ids
#    (search the LIST directly - this is the slow version)
# 4. Write count_known_set(ids): convert order_ids to a set ONCE, then count
#    the same way against the set
# 5. Time each call with time.perf_counter() the way the worked example did
# 6. Print both counts, then both timings
#
# ✅ The counts are the part you can check exactly - they must match:
# list version: 2500
# set  version: 2500
#
# The two timings differ on every machine, so no fixed number is given here.
# What matters is the gap: the set version should finish many times faster.

# your code here

📋 Quick Reference — Profiling & Performance

Tool / SyntaxWhat it does
cProfile.run('fn()')Profile function call counts and time
timeit.timeit('expr', number=1000)Benchmark small code snippets
line_profilerProfile line-by-line execution time
memory_profilerTrack memory usage per line
__slots__Reduce class memory footprint

🎉 Great work! You've completed this lesson.

You now know how to measure, diagnose, and fix performance bottlenecks — the professional workflow every senior engineer uses.

Practice quiz

What is the core idea behind profiling before optimising?

  • Always optimise the longest function first
  • Rewrite everything in C
  • Measure what's slow instead of guessing
  • Add more print statements

Answer: Measure what's slow instead of guessing. You cannot optimise what you do not measure — profiling replaces guessing with data.

Which standard-library tool profiles function call counts and time?

  • cProfile
  • tracemalloc
  • asyncio
  • logging

Answer: cProfile. cProfile records every call, how long each takes, and how many times it runs.

In cProfile output, what does 'tottime' measure?

  • Total time including all sub-calls
  • Number of times the function was called
  • Total program runtime
  • Time in the function itself, excluding sub-calls

Answer: Time in the function itself, excluding sub-calls. tottime is time spent in the function body alone; cumtime includes time in sub-calls.

Which sort key best reveals the actually-slow functions?

  • "ncalls"
  • "tottime"
  • "cumtime"
  • "name"

Answer: "tottime". Sort by tottime to find where time is really spent; then use cumtime to trace callers.

What does functools.lru_cache do to a recursive fib function?

  • Caches results so repeated calls are instant
  • Slows it down by caching
  • Runs it on multiple cores
  • Converts it to a loop

Answer: Caches results so repeated calls are instant. lru_cache memoizes results, turning exponential recursion into near-instant lookups.

What does timeit.timeit('expr', number=1000) return?

  • The result of the expression
  • A profile report object
  • The total time (a float) to run it that many times
  • The number of calls

Answer: The total time (a float) to run it that many times. timeit returns the total elapsed time as a float for the given number of executions.

Why use a generator expression (x*x for x in nums) over a list for large data?

  • It is always faster to build
  • It saves memory by yielding items lazily
  • It sorts the data
  • It runs in parallel

Answer: It saves memory by yielding items lazily. Generators produce items one at a time instead of building the whole list in memory.

What is the recommended fix for slow string concatenation in a loop?

  • Use s = s + x each time
  • Use a global string
  • Use print() to build it
  • Use ''.join(list_of_strings)

Answer: Use ''.join(list_of_strings). ''.join builds the result in one pass; repeated + creates a new string object every iteration.

What does the 80/20 rule say about performance?

  • 80% of code runs in 20% of the time
  • Roughly 20% of code accounts for ~80% of runtime
  • Optimise 80% of functions
  • 20% of bugs cause 80% of crashes

Answer: Roughly 20% of code accounts for ~80% of runtime. A small fraction of code dominates runtime, so optimising the wrong 80% gives no improvement.

Which tool tracks memory usage line-by-line?

  • cProfile
  • timeit
  • memory_profiler
  • pstats

Answer: memory_profiler. memory_profiler shows memory growth per line; cProfile and timeit measure time, not memory.

Continue this course