Concurrency: Threads vs Processes
Reviewed & published by Brayan K
Master Python's concurrency models and learn when to use threads, processes, or AsyncIO for maximum performance.
Part of the free Python course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
What You'll Learn in This Lesson
- • The difference between concurrency and parallelism — and why it matters
- • How Python threads work and when the GIL limits you
- • When to use threading vs multiprocessing vs asyncio
- • How to share data safely between threads using locks and queues
- • How to spawn and manage worker processes for CPU-bound tasks
- • Real-world patterns: scrapers, batch processors, concurrent file downloads
🔥 1. Why Concurrency Exists
Computers run tasks in parallel using:
- CPU cores → true parallelism (multiple chefs)
- Scheduling → switching between tasks (one chef, many pots)
- I/O waits → idle time you can use productively (waiting for water to boil)
| Type of Work | What It Means | Best Solution | Real Example |
|---|---|---|---|
| CPU-bound | Heavy calculations that keep the CPU busy | multiprocessing | Image processing, ML training |
| I/O-bound | Waiting for external resources | threads or asyncio | API calls, file downloads |
⚙️ 2. Threads in Python
A thread is a lightweight unit of execution within a single process.
| What Threads Share | Why It Matters |
|---|---|
| Memory | Fast communication, but risk of conflicts |
| Variables | Easy data sharing, but need locks for safety |
| File handles | Can work with same files simultaneously |
| Python interpreter | Limited by GIL for CPU work |
- network requests (waiting for servers)
- reading/writing files (waiting for disk)
- user interface responsiveness (don't freeze the UI)
- downloading many URLs (lots of waiting)
- heavy computation (GIL blocks parallelism)
- CPU-bound workload (use processes instead)
import threading
import time
def task(name):
# Each thread announces when it starts
print(f"{name} starting")
# Simulate I/O wait (like downloading a file)
# During this sleep, other threads can run!
time.sleep(0.5)
print(f"{name} done")
# Create 3 thread objects
# Each thread will run the 'task' function with a different name
threads = [
threading.Thread(target=task, args=(f"Thread-{i}",))
for i in range(3)
]
# Start all threads (they begin running)
for t in threads:
t.start()
# Wait for all threads to complete before continuing
# Without join(), the main program might exit before threads finish
for t in threads:
t.join()
print("All threads finished")
# Notice: All 3 threads ran their sleep at the same time!
# Total time: ~0.5 seconds, NOT 1.5 seconds🧠 3. Understanding the GIL (Global Interpreter Lock)
The GIL (Global Interpreter Lock) is a mutex that protects access to Python objects.
| Situation | GIL Effect | Result |
|---|---|---|
| Running Python code | GIL is held | Other threads wait |
| Waiting for network/file | GIL is released | Other threads can run |
| Using C extensions (NumPy) | Often released | True parallelism possible |
❌ Threads cannot speed up CPU-heavy tasks
(e.g., image processing, hashing, compression)
✔ Threads can hugely speed up I/O tasks
(e.g., APIs, web scraping, DB queries)
⚡ 4. Processes in Python
A process is a full Python interpreter with its own memory.
| Aspect | Threads | Processes |
|---|---|---|
| Memory | Shared | Separate (isolated) |
| GIL | Shared (limits CPU work) | Each has its own |
| Startup time | Fast (microseconds) | Slow (milliseconds) |
| Data sharing | Easy (same memory) | Requires serialization |
- True parallelism (uses multiple CPU cores)
- Great for CPU-bound work
- No GIL problems — each process has its own
- More memory used (each process needs its own)
- Slower to start (spawning a new Python interpreter)
- Harder to share data (must serialize/deserialize)
import multiprocessing
import time
def cpu_task(name):
print(f"{name} starting")
# CPU-intensive work: calculate sum of squares
# This keeps the CPU busy (not waiting for I/O)
total = sum(i * i for i in range(10_000_000))
print(f"{name} done: {total}")
# IMPORTANT: This guard is required on Windows!
# Without it, the script would spawn infinite processes
if __name__ == "__main__":
# Create 2 process objects
processes = [
multiprocessing.Process(target=cpu_task, args=(f"Process-{i}",))
for i in range(2)
]
# Start all processes
for p in processes:
p.start()
# Wait for all to complete
for p in processes:
p.join()
print("All processes finished")
# On a 2+ core machine, both processes run TRULY simultaneously!📖 Worked Example: Threads on I/O-Bound Work
This is the single most useful thing threads do for you: waiting in parallel. The program below downloads four pages twice — once one after the other, once with a pool of four threads — and times both. time.sleep() stands in for the network, because a sleeping thread releases the GIL exactly the way a waiting socket does.
ThreadPoolExecutor is the modern way to use threads: you hand it a function and a list, it runs them across its workers, and pool.map() gives the results back in the original order — so threading does not scramble your data.
import time
from concurrent.futures import ThreadPoolExecutor
def download(page):
"""Pretend to fetch a page. time.sleep stands in for network waiting."""
time.sleep(0.5) # the thread releases the GIL while it waits
return f"{page}: 1200 bytes"
pages = ["home", "about", "pricing", "contact"]
# --- Run 1: one at a time. Four half-second waits = two seconds. ---
start = time.perf_counter()
sequential = [download(p) for p in pages]
seq_time = time.perf_counter() - start
# --- Run 2: four threads, so the four waits overlap. ---
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=4) as pool: # 'with' shuts the pool down for you
threaded = list(pool.map(download, pages)) # map keeps the input order
par_time = time.perf_counter() - start
for line in threaded:
print(line)
print("Sequential total:", round(seq_time, 1), "seconds")
print("Threaded total: ", round(par_time, 1), "seconds")
print("About", round(seq_time / par_time), "times faster")
print("Same results either way?", sequential == threaded)
# ✅ Expected output (timings may differ by a few hundredths on your machine):
# home: 1200 bytes
# about: 1200 bytes
# pricing: 1200 bytes
# contact: 1200 bytes
# Sequential total: 2.0 seconds
# Threaded total: 0.5 seconds
# About 4 times faster
# Same results either way? True🔄 5. Side-by-Side Comparison
| Feature | Threads | Processes |
|---|---|---|
| Speed for I/O | ⭐⭐⭐⭐⭐ | ⭐⭐⭐ |
| Speed for CPU | ⭐⭐ | ⭐⭐⭐⭐⭐ |
| Memory usage | Low | High |
| Startup time | Fast | Slow |
| Shares memory? | Yes | No |
| Avoids GIL? | ❌ | ✔ |
| Best for? | I/O tasks | CPU tasks |
🧪 6. Real-World Examples
✔ Threads Example: Web Scraping
import threading
import requests
import time
def fetch_url(url):
response = requests.get(url)
print(f"Fetched {url}: {len(response.content)} bytes")
urls = [
"https://api.github.com/users/github",
"https://api.github.com/users/microsoft",
"https://api.github.com/users/google",
"https://api.github.com/users/facebook"
]
start = time.time()
threads = [threading.Thread(target=fetch_url, args=(url,)) for url in urls]
for t in threads:
t.start()
for t in threads:
t.join()
print(f"Completed in {time.time() - start:.2f} seconds")Threads shine because requests are I/O-bound.
✔ Processes Example: Image Processing
import multiprocessing
from PIL import Image
import time
def process_image(filename):
# CPU-heavy: resize, filter, transform
img = Image.open(filename)
img = img.resize((800, 600))
img = img.convert('L') # grayscale
img.save(f"processed_{filename}")
print(f"Processed {filename}")
if __name__ == "__main__":
images = ["img1.jpg", "img2.jpg", "img3.jpg", "img4.jpg"]
start = time.time()
processes = [
multiprocessing.Process(target=process_image, args=(img,))
for img in images
]
for p in processes:
p.start()
for p in processes:
p.join()
print(f"Completed in {time.time() - start:.2f} seconds")CPU work → multiprocessing is ideal.
🧱 7. Mixing Threads & Processes (Hybrid Model)
Many real systems use both:
- Processes for CPU-heavy pipelines
- Threads for network/database operations
- AsyncIO for massive lightweight tasks
| Component | Role in Hybrid System | Why? |
|---|---|---|
| AsyncIO | Orchestra conductor | Lightweight coordination of thousands of tasks |
| Threads | I/O specialists | Handle blocking I/O without stopping AsyncIO |
| Processes | Heavy lifters | CPU work on multiple cores simultaneously |
Example: A crawler using:
- Threads → fetch 1,000 URLs
- Processes → process each page's text
- AsyncIO → coordinate flows
This is how production-grade scrapers, ML preprocessors, and automation bots work.
🔥 8. When Should You Use What?
- ✔ Web scraping
- ✔ Automation
- ✔ Network servers
- ✔ Waiting on APIs
- ✔ ML preprocessing
- ✔ Math-heavy operations
- ✔ Hashing, encryption
- ✔ Image/video processing
- ✔ Massive concurrency
- ✔ Lightweight operations
- ✔ Network-first tasks
- ✔ High scalability needed
🧠 9. How Python Schedules Threads Internally (Advanced)
Python uses cooperative + preemptive scheduling for threads.
Here's what really happens:
✔ The OS schedules threads
The operating system decides when each thread runs, based on CPU availability.
✔ The GIL schedules Python bytecode
Inside Python, only ONE thread can execute Python bytecode at once.
- artificial bottlenecks for CPU tasks
- no bottlenecks for I/O (because threads release the GIL while waiting)
How the GIL behaves:
- A thread starts running Python code
- When it hits I/O (network, file), it releases the GIL
- Another thread can run
- When I/O completes, the thread resumes and reacquires GIL
- ✔ 100 threads downloading files works great
- ❌ 100 threads crunching numbers does NOT
⚡ 10. The True Strength of Threads: I/O Parallelism
Let's say you need to download 10,000 images.
Sequential time = 10,000 × (0.4 seconds each) = ~4000 seconds (1.1 hours)
Threaded time (200 threads) = ~20 seconds total
Because threads wait most of the time, so Python overlaps waits.
- HTTP requests
- reading/writing files
- database queries
- waiting for user input
- network sockets
Threads shine because I/O releases the GIL.
🔥 11. The True Strength of Processes: CPU Parallelism
- ML preprocessing
- physics simulation
- number crunching
- image/video filtering
- audio processing
If you run these in threads → NO speed improvement.
If you run these in processes → 4× faster on 4 cores, 12× on 12 cores, etc.
Processes use true multi-core hardware.
🧬 12-18. Advanced Concurrency Topics
The next sections cover professional-level concurrency patterns:
12. Data Sharing Between Threads
Lock, RLock, Event, Semaphore, Queue — Safe shared memory primitives
13. Data Sharing Between Processes
multiprocessing.Queue, Pipe, Manager, shared memory arrays
14. ThreadPoolExecutor vs ProcessPoolExecutor
concurrent.futures abstraction for both models
15. Hybrid Concurrency Pattern
Combining AsyncIO + Threads + Processes like YouTube/Instagram
16. Avoiding Common Bugs
Race conditions, deadlocks, blocking calls, serialization issues
CPU-heavy: 8× with processes; I/O-heavy: 50× with threads
✔ Threads for I/O | ✔ Processes for CPU | ✔ AsyncIO for massive concurrency
🧨 19. Race Conditions — The Silent Killer
A race condition happens when:
- Two or more threads access shared data
- AND at least one thread modifies it
- AND execution order determines the final result
| Symptom | What's Happening | Example |
|---|---|---|
| Inconsistent results | Different output each run | Counter shows 987,432 instead of 1,000,000 |
| Lost updates | Changes disappear | Two users edit same record, one is lost |
| Works sometimes | Timing-dependent bugs | Passes tests locally, fails in production |
import threading
counter = 0
def increment():
global counter
for _ in range(100_000):
counter += 1
threads = [threading.Thread(target=increment) for _ in range(10)]
[t.start() for t in threads]
[t.join() for t in threads]
print(counter) # not guaranteed to be 1000000 — see the note belowcounter += 1 is NOT atomic. It's 3 instructions:
- load counter
- store result
Two threads collide → corrupted values.
A version where the race really does show up:
Here the read and the write are separate statements with a yield point between them — which is what happens naturally as soon as your critical section does anything more than one arithmetic step.
import threading
import time
counter = 0
def increment():
global counter
for _ in range(2000):
current = counter # READ the shared value
time.sleep(0) # yield: another thread may run right here
counter = current + 1 # WRITE back — overwriting anyone who ran in between
threads = [threading.Thread(target=increment) for _ in range(5)]
for t in threads:
t.start()
for t in threads:
t.join()
print("Expected:", 5 * 2000)
print("Actual: ", counter)
print("Lost updates:", 5 * 2000 - counter)
# There is no fixed expected output here — that IS the lesson. Every run
# differs. One real run on Python 3.11 printed:
# Expected: 10000
# Actual: 2005
# Lost updates: 7995
# Six runs on the same machine gave 2005, 2001, 2008, 2002, 2006 and 2000.
# Roughly 80% of the work vanished, and never by the same amount twice.🧱 20. Fixing Race Conditions With Locks
Locks ensure only ONE thread accesses critical code at once.
| Lock Method | What It Does | When to Use |
|---|---|---|
| lock.acquire() | Grab the lock (waits if taken) | Manual control needed |
| lock.release() | Release the lock | After acquire() |
| with lock: | Auto acquire + release | ✅ Always prefer this! |
import threading
lock = threading.Lock()
counter = 0
def increment():
global counter
for _ in range(100_000):
with lock:
counter += 1
threads = [threading.Thread(target=increment) for _ in range(10)]
[t.start() for t in threads]
[t.join() for t in threads]
print(counter) # Now correctly equals 1,000,000- ✔ Deterministic
- ✔ Correct final value
But… ❗ Locks introduce blocking, which could slow threads.
🎯 Your Turn: Make the Till Add Up
Five checkout threads each ring up 1,000 sales of £10. The shared total is read, changed, then written back — the exact pattern that loses updates. Everything is written for you except the three pieces of the lock. Fill in each ___ and run it.
import threading
# 🎯 YOUR TURN — replace the three ___ blanks
lock = threading.___() # 👉 which threading class hands out exclusive access?
total = 0
def add_sales(amount, times):
global total
for _ in range(times):
with ___: # 👉 hold the lock while touching shared data
current = total # READ
current += amount # CHANGE
total = current # WRITE
# Five checkout threads, 1000 sales of £10 each
threads = [threading.Thread(target=add_sales, args=(10, 1000)) for _ in range(___)] # 👉 how many threads?
for t in threads:
t.start()
for t in threads:
t.join() # wait for every thread before reading the result
print("Threads used:", len(threads))
print("Total:", total)
print("Correct?", total == 50000)
# ✅ Expected output:
# Threads used: 5
# Total: 50000
# Correct? True1) threading.Lock() — the basic mutual-exclusion lock.
2) with lock: — acquires on the way in and releases on the way out, even if the block raises.
3) 5 — five threads × 1000 sales × £10 = 50000.
If you get AttributeError: module 'threading' has no attribute 'lock', remember the class name is capitalised: Lock.
🏆 Mini-Challenge: Concurrent Health Check
Outline only this time. You monitor five servers. Checking one takes 0.3 seconds of pure waiting, so doing them one after the other takes 1.5 seconds — but they should all be checked at once, in about 0.3 seconds total. Decide for yourself whether this is a job for threads or processes before you write a line.
import time
from concurrent.futures import ThreadPoolExecutor
# 🎯 MINI-CHALLENGE: check five servers at once
# 1. Write check(host): sleep 0.3 seconds, then return f"{host}: OK"
# 2. hosts = ["alpha", "beta", "gamma", "delta", "epsilon"]
# 3. Record the start time with time.perf_counter()
# 4. Open a ThreadPoolExecutor with 5 workers and use pool.map(check, hosts),
# wrapping it in list(...) so every result is collected before you time it
# 5. Print each result line, then print how long the whole thing took:
# f"Checked {len(hosts)} hosts in {round(elapsed, 1)} seconds"
#
# Why threads and not processes? Nothing here does any computing — it only
# waits. Processes would work but cost far more to start.
#
# ✅ Expected output:
# alpha: OK
# beta: OK
# gamma: OK
# delta: OK
# epsilon: OK
# Checked 5 hosts in 0.3 seconds
# your code hereimport time
from concurrent.futures import ThreadPoolExecutor
def check(host):
time.sleep(0.3) # stands in for a network round-trip
return f"{host}: OK"
hosts = ["alpha", "beta", "gamma", "delta", "epsilon"]
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=5) as pool:
results = list(pool.map(check, hosts))
elapsed = time.perf_counter() - start
for line in results:
print(line)
print(f"Checked {len(hosts)} hosts in {round(elapsed, 1)} seconds")Drop max_workers to 2 and the same run takes about 0.9 seconds — three batches instead of one. Pool size is the dial you tune.
🔒 21-32. Expert-Level Concurrency Patterns
21. RLock — Reentrant Locks
Same thread can acquire lock multiple times
When threads freeze forever waiting on each other
Controlling access to limited resources
26. Shared Memory Between Processes
Value, Array, Manager, shared_memory
What breaks when passing objects to processes
chunksize, initializer, preloading data
29. Real Architecture: High-Performance Scraper
AsyncIO → ThreadPool → ProcessPool pipeline
30. Real Architecture: ML Preprocessing
ProcessPool for CPU, ThreadPool for I/O, AsyncIO for APIs
How FastAPI, Scrapy, PyTorch use concurrency
32. The Future: No-GIL Python
Python 3.13+ will enable true parallel threads for CPU work
🎓 Final Summary
You now understand advanced concurrency:
✔ Threads vs Processes — when to use each
✔ The GIL and its impact
✔ Race conditions, Locks, RLocks
✔ Events & Conditions for coordination
✔ Shared memory across processes
You're operating at professional backend engineer level now.
📋 Quick Reference — Concurrency
| Tool | Best for |
|---|---|
| threading.Thread | I/O-bound tasks, network calls |
| multiprocessing.Process | CPU-bound tasks (bypasses GIL) |
| threading.Lock() | Protect shared state between threads |
| queue.Queue() | Thread-safe data passing |
| multiprocessing.Queue() | Process-safe data passing |
🎉 Great work! You've completed this lesson.
You now understand the GIL, when to use threads vs processes, and how to safely share data between concurrent workers.
Practice quiz
What is the difference between concurrency and parallelism?
- They are the same thing
- Parallelism is slower than concurrency
- Concurrency switches between tasks; parallelism does multiple things at the exact same time
- Concurrency requires multiple CPU cores
Answer: Concurrency switches between tasks; parallelism does multiple things at the exact same time. Concurrency interleaves tasks; parallelism runs them truly simultaneously on multiple cores.
What is the GIL (Global Interpreter Lock)?
- A mutex that lets only one thread run Python bytecode at a time
- A networking protocol
- A way to lock files
- A garbage collector
Answer: A mutex that lets only one thread run Python bytecode at a time. The GIL is a mutex ensuring only one thread executes Python bytecode at a time.
For CPU-bound work, which approach actually achieves true parallelism in Python?
- threading
- asyncio
- None — Python can't parallelize
- multiprocessing
Answer: multiprocessing. Each process has its own interpreter and GIL, so multiprocessing gives true CPU parallelism.
Why do threads work well for I/O-bound tasks despite the GIL?
- The GIL is disabled for threads
- A thread releases the GIL while waiting on I/O, letting another thread run
- I/O tasks don't use the GIL at all
- Threads create separate interpreters
Answer: A thread releases the GIL while waiting on I/O, letting another thread run. When a thread waits on I/O it releases the GIL, so other threads can make progress.
Do threads or processes share memory by default?
- Threads share memory; processes have separate memory
- Processes share memory; threads do not
- Both share memory
- Neither shares memory
Answer: Threads share memory; processes have separate memory. Threads share the same memory; processes are isolated and need serialization to share data.
Why is counter += 1 unsafe across multiple threads without a lock?
- It is always atomic and safe
- Integers can't be shared
- It is not atomic — load, add, and store can interleave
- The GIL prevents all sharing
Answer: It is not atomic — load, add, and store can interleave. counter += 1 is three steps (load, add, store); interleaving threads corrupt the value — a race condition.
What is the recommended way to use a threading.Lock around shared data?
- lock.acquire() and forget to release
- with lock:
- Set lock = True
- No lock is needed
Answer: with lock:. with lock: auto-acquires and auto-releases, which is the safest pattern.
Compared with threads, how do processes generally start up?
- Faster, in microseconds
- Instantly with zero cost
- At the same speed as threads
- Slower, in milliseconds (a new interpreter must spawn)
Answer: Slower, in milliseconds (a new interpreter must spawn). Processes are slower to start since each spawns its own Python interpreter and memory.
Which model is best for handling massive numbers of lightweight network tasks?
- multiprocessing
- asyncio
- One thread per task
- Pure sequential code
Answer: asyncio. asyncio scales to massive lightweight, network-first concurrency on a single thread.
Running heavy number-crunching in 100 threads gives what result in CPython?
- A ~100x speedup
- A guaranteed crash
- No real speedup — the GIL serializes CPU-bound bytecode
- True parallelism across cores
Answer: No real speedup — the GIL serializes CPU-bound bytecode. The GIL serializes Python bytecode, so CPU-bound threads get no real speedup — use processes.
Continue this course
- Previous: AsyncIO: Event Loop, Tasks & Futures
- Next: Parallelism with concurrent.futures — Run CPU-bound tasks in parallel with ProcessPoolExecutor
- Quick reference: Python cheat sheet