Files, Streams & Large Datasets
Reviewed & published by Brayan K
Master efficient file handling, streaming patterns, and processing massive datasets that don't fit in memory.
Part of the free Python course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
📊 What You'll Learn
This lesson teaches professional data handling techniques used in:
- Machine learning systems & data pipelines
- Analytics & ETL jobs
- Log processing & monitoring systems
- Web scrapers & API integrations
- Large-scale data transformations
You'll learn how to:
- Stream files without loading them into memory
- Process multi-gigabyte datasets efficiently
- Build composable data pipelines
- Handle compressed and binary data
- Implement parallel file processing
📥 Python Download & Setup
Download Python from: python.org/downloads
Latest version recommended (3.11+)
Part 1: File I/O Fundamentals & Streaming Basics
1. Python's File I/O Model (Text vs Binary, Buffered vs Unbuffered)
- • Raw I/O = Receiving individual items one by one (slow)
- • Buffered I/O = Receiving full pallets at a time (efficient)
- • Text I/O = Unpacking and translating foreign labels into English
When you open a file, Python uses a layered I/O system:
| Layer | What It Does | Speed |
|---|---|---|
| Raw I/O | Direct interaction with OS file descriptors | Slowest |
| Buffered I/O | Uses memory buffers (reads/writes in chunks) | Fast |
| Text I/O | Decodes bytes → Unicode strings | Automatic encoding |
with open("data.txt", "r", encoding="utf-8") as f:
content = f.read()
print(content)Modes:
- "w" → write (overwrite)
- "a" → append
- "b" → binary mode
- "rb" → read binary
- "wb" → write binary
Buffering matters: Python reads data in chunks internally to reduce syscalls, making it ideal for large datasets.
2. Reading Files the Right Way (Avoid .read() on Large Files)
data = open("log.txt").read()
print(data[:100])This loads the entire file into memory, which is disastrous for large logs.
with open("log.txt") as f:
for line in f:
# Process each line without loading full file
print(line.strip())| Approach | Memory Usage | For 10GB File |
|---|---|---|
| f.read() | Entire file in RAM | 10GB+ RAM needed 💥 |
| for line in f: | One line at a time | ~KB of RAM ✓ |
Advantages:
- ✔ No RAM explosion
- ✔ Pause/resume possible
- ✔ Efficient buffering
- ✔ Ideal for multi-GB files
🟢 Worked Example: write a file, then stream it back
The snippets above open files like data.txt and log.txt that have to already exist — run one of them on a machine without that file and you get FileNotFoundError. So this example creates its own file first, then reads it back four different ways, then deletes it. It is completely self-contained: run it anywhere and it works.
# 🟢 WORKED EXAMPLE — create a file, then stream it back, all in one script
from pathlib import Path
LOG = Path("visits.log") # Path is the modern way to name a file; no string juggling
lines = [
"2026-05-01 GET /home 200",
"2026-05-01 GET /pricing 200",
"2026-05-01 GET /admin 403",
"2026-05-01 GET /home 200",
"2026-05-01 GET /checkout 500",
]
# ---------- 1. Write it ----------
# "w" creates the file, or empties it if it already exists. The with block closes
# the file for you the moment you leave it — even if an error is raised inside.
with open(LOG, "w", encoding="utf-8") as f:
for line in lines:
f.write(line + "\n") # write() adds NO newline of its own; that is your job
print("bytes on disk:", LOG.stat().st_size)
# ---------- 2. Read the lot in one gulp ----------
# Fine for 133 bytes. On a 10 GB log this is the line that kills the process.
with open(LOG, encoding="utf-8") as f: # no mode given, so "r" (read) is assumed
whole = f.read()
print("characters read:", len(whole))
print("first line:", whole.splitlines()[0])
# ---------- 3. Stream it: one line in memory at a time ----------
# Looping over the file object itself is the memory-safe version of the above.
statuses = {}
with open(LOG, encoding="utf-8") as f:
for line in f:
code = line.strip().split()[-1] # strip the newline, take the last field
statuses[code] = statuses.get(code, 0) + 1 # tally it up
print("status counts:", statuses)
# ---------- 4. Append without destroying what is there ----------
# "a" starts writing at the end. Using "w" here would wipe the whole file.
with open(LOG, "a", encoding="utf-8") as f:
f.write("2026-05-01 GET /pricing 200\n")
with open(LOG, encoding="utf-8") as f:
print("line count now:", sum(1 for _ in f)) # counts lines without storing any
# ---------- 5. Tidy up after yourself ----------
LOG.unlink() # delete the file
print("file still there?", LOG.exists())
# ✅ Expected output:
# bytes on disk: 133
# characters read: 133
# first line: 2026-05-01 GET /home 200
# status counts: {'200': 3, '403': 1, '500': 1}
# line count now: 6
# file still there? FalseTwo things worth pausing on. The byte count and the character count match here only because every character is plain ASCII — put a £ or an emoji in that file and UTF-8 spends more than one byte on it, so the numbers separate. And sum(1 for _ in f) counts six lines while holding exactly one in memory: that is the whole streaming idea in a single line.
🎯 Your Turn: write and re-read a to-do file
Three blanks, each one a detail that beginners get wrong on their first file. Fill them in and run it.
# 🎯 YOUR TURN — fill in the blanks marked with ___
from pathlib import Path
NOTES = Path("notes.txt")
tasks = ["buy milk", "call dentist", "finish lesson"]
# 1) Open the file for writing — the mode that creates it fresh.
with open(NOTES, "___", encoding="utf-8") as f: # 👉 "r", "w" or "a"?
for task in tasks:
# 2) write() never adds a line break. Put one on the end yourself.
f.write(task + "___") # 👉 replace ___ with \n (backslash n)
# Read it back, numbering each line as it streams past.
with open(NOTES, encoding="utf-8") as f:
for number, line in enumerate(f, start=1): # start=1 so humans count from 1
# 3) Each line still carries its trailing newline. Remove it.
print(f"{number}. {line.___()}") # 👉 which string method trims whitespace?
with open(NOTES, encoding="utf-8") as f:
print("Total lines:", sum(1 for _ in f))
NOTES.unlink() # clean up so re-running starts from scratch
# ✅ Expected output:
# 1. buy milk
# 2. call dentist
# 3. finish lesson
# Total lines: 3If your numbers come out with blank lines between them, blank 3 is the culprit — the newline from the file plus the newline from print() gives you two.
3. Using Buffered Streams Explicitly
For very large binary files (videos, models, audio):
from io import BufferedReader
with open("video.mp4", "rb") as f:
stream = BufferedReader(f)
chunk = stream.read(4096)
print(f"Read {len(chunk)} bytes")4. Writing Files Efficiently (Avoid Tiny Writes)
Frequent small writes slow everything down.
from io import StringIO
items = ["apple", "banana", "cherry", "date"]
buffer = StringIO()
for item in items:
buffer.write(item + "\n")
with open("output.txt", "w") as f:
f.write(buffer.getvalue())
print("Written to file!")
# ✅ Expected output:
# Written to file!5. Working With CSV Files at Scale
Python's csv module supports streaming:
import csv
with open("data.csv") as f:
reader = csv.reader(f)
for row in reader:
# Process each row
print(row)Never convert a large CSV into a list:
🏁 Mini-Challenge: a CSV round trip
Brief only this time. Write three sales rows to a CSV file, then stream them back and work out the takings — without ever building a list of every row.
Two things to look up rather than guess: csv.DictWriter needs a fieldnames list and a call to writeheader(), and every value that comes back out of csv.DictReader is a string, so you must convert with int() before doing any arithmetic.
# 🏁 MINI-CHALLENGE — write it yourself
import csv
from pathlib import Path
SALES = Path("sales.csv")
rows = [
{"item": "tea", "units": 3, "pence": 120},
{"item": "coffee", "units": 2, "pence": 275},
{"item": "cake", "units": 1, "pence": 340},
]
# 1. Write rows to SALES with csv.DictWriter.
# open(SALES, "w", newline="", encoding="utf-8") <- newline="" is required on
# Windows or you get a blank line between every row.
# fieldnames=["item", "units", "pence"], then writeheader(), then writerows(rows).
# 2. Print the raw file so you can see what a CSV actually is:
# print(SALES.read_text(encoding="utf-8"))
# 3. Stream it back with csv.DictReader. For each row:
# line_total = int(row["units"]) * int(row["pence"])
# add it to a running total and print item: units x pence = line_total p
# 4. Print the grand total in pounds, 2 decimal places.
# 5. SALES.unlink() to clean up.
# your code here
# ✅ Expected output (the blank line after the raw CSV is real — the file
# ends in a newline and print() adds another):
# item,units,pence
# tea,3,120
# coffee,2,275
# cake,1,340
#
# tea: 3 x 120p = 360p
# coffee: 2 x 275p = 550p
# cake: 1 x 340p = 340p
# Total: £12.506. Handling JSON Without Crashing RAM
import json
# For line-delimited JSON (NDJSON)
# Each line is a separate JSON object
sample_ndjson = '''{"name": "Alice", "age": 30}
{"name": "Bob", "age": 25}
{"name": "Charlie", "age": 35}'''
for line in sample_ndjson.strip().split("\n"):
item = json.loads(line)
print(f"Processing: {item['name']}")
# ✅ Expected output:
# Processing: Alice
# Processing: Bob
# Processing: Charlie7. Working With Very Large Binary Files
# Simulate chunked reading
data = b"X" * 5000 # 5KB of data
chunk_size = 1024
offset = 0
while offset < len(data):
chunk = data[offset:offset + chunk_size]
print(f"Processing chunk at offset {offset}, size {len(chunk)}")
offset += chunk_size
# ✅ Expected output:
# Processing chunk at offset 0, size 1024
# Processing chunk at offset 1024, size 1024
# Processing chunk at offset 2048, size 1024
# Processing chunk at offset 3072, size 1024
# Processing chunk at offset 4096, size 9048. Memory Mapping (mmap) — Fast Random Access to Huge Files
import mmap
# Simulated memory-mapped access pattern
# In real use: mm = mmap.mmap(f.fileno(), 0)
print("Memory mapping allows:")
print("- Instant random access to any position")
print("- No RAM explosion for huge files")
print("- Extremely fast reads")
print("- Ideal for ML datasets and databases")
# Example pattern:
# with open("huge.dat", "rb") as f:
# mm = mmap.mmap(f.fileno(), 0, access=mmap.ACCESS_READ)
# print(mm[1000:1010]) # read 10 bytes at offset 1000
# ✅ Expected output:
# Memory mapping allows:
# - Instant random access to any position
# - No RAM explosion for huge files
# - Extremely fast reads
# - Ideal for ML datasets and databases| Method | Random Access Speed | Memory Usage |
|---|---|---|
| f.seek() + f.read() | Medium | Low (read size) |
| f.read() entire file | Fast (after load) | File size 💥 |
| mmap | Instant | Near zero |
9. Chunk Processing for Massive Datasets
def read_in_chunks(data, size=1024):
"""Generator that yields chunks of data."""
offset = 0
while offset < len(data):
yield data[offset:offset + size]
offset += size
# Demo with sample data
sample = "X" * 5000
for i, chunk in enumerate(read_in_chunks(sample, 1000)):
print(f"Chunk {i}: {len(chunk)} chars")
# ✅ Expected output:
# Chunk 0: 1000 chars
# Chunk 1: 1000 chars
# Chunk 2: 1000 chars
# Chunk 3: 1000 chars
# Chunk 4: 1000 chars10. Using Iterators & Generators in Data Pipelines
def read_lines(lines):
for line in lines:
yield line
def filter_errors(lines):
for line in lines:
if "ERROR" in line:
yield line
# Sample log data
log_lines = [
"INFO: Starting service",
"ERROR: Connection failed",
"INFO: Retrying...",
"ERROR: Timeout occurred",
"INFO: Service stopped"
]
for line in filter_errors(read_lines(log_lines)):
print(line)
# ✅ Expected output:
# ERROR: Connection failed
# ERROR: Timeout occurred11. Working With Compressed Data (ZIP, GZIP)
import gzip
import io
# Create sample gzipped data
sample_text = "Hello, compressed world!\nLine 2\nLine 3"
compressed = gzip.compress(sample_text.encode())
# Read it back (streaming pattern)
with gzip.open(io.BytesIO(compressed), "rt") as f:
for line in f:
print(line.strip())
# ✅ Expected output:
# Hello, compressed world!
# Line 2
# Line 312. Handling Data That Doesn't Fit In RAM
# Pattern for counting lines in a 10GB file
# Memory usage stays near 0MB!
lines = [f"Line {i}" for i in range(100000)]
total = 0
for _ in lines:
total += 1
print(f"Total lines: {total:,}")
# ✅ Expected output:
# Total lines: 100,00013. Avoiding the Most Common File-I/O Mistakes
- ❌ Reading entire large files into memory
- ❌ Using .read() instead of iteration
- ❌ Building huge lists from file rows
- ❌ Tiny writes inside loops
- ❌ Forgetting to close files
- ❌ Failing to stream compressed data
- ❌ Loading massive JSON using json.load()
- ❌ Storing temporary files unnecessarily
14. Best Practices Recap
- ✔ Use with for everything
- ✔ Stream line-by-line
- ✔ Use chunked reads for large binaries
- ✔ Use ijson for large JSON
- ✔ Use csv reader instead of list(reader)
- ✔ Use memory mapping for random access
- ✔ Use generators to build pipelines
- ✔ Use compression modules without extracting
Part 2: Advanced Streaming & Dataset Engineering
1. Understanding Buffered I/O Depth
from io import FileIO, BufferedReader
# Python's I/O layers:
# 1. Raw I/O (FileIO) - OS file descriptors
# 2. Buffered I/O - memory buffers
# 3. Text I/O - decoding/encoding
print("I/O Layer Architecture:")
print(" Raw I/O → direct OS interaction")
print(" Buffered → efficient chunking")
print(" Text I/O → Unicode handling")
# Larger buffer sizes improve throughput
# with open("large.bin", "rb") as f:
# reader = BufferedReader(f, buffer_size=1024*1024)
# ✅ Expected output:
# I/O Layer Architecture:
# Raw I/O → direct OS interaction
# Buffered → efficient chunking
# Text I/O → Unicode handling2. Random Access With Seek
import io
# Create sample file-like object
data = b"HEADER_DATA_HERE" + b"X" * 100 + b"TARGET"
file = io.BytesIO(data)
# Jump to position
file.seek(116) # Skip to "TARGET"
print(f"Read at position 116: {file.read(6)}")
# Go back to start
file.seek(0)
print(f"Header: {file.read(16)}")
# ✅ Expected output:
# Read at position 116: b'TARGET'
# Header: b'HEADER_DATA_HERE'3. Streaming API Data
# Pattern for streaming API data
# Prevents loading giant payloads into RAM
# import requests
# with requests.get(url, stream=True) as r:
# for chunk in r.iter_content(1024):
# process(chunk)
print("Streaming API pattern:")
print("1. Set stream=True in request")
print("2. Use iter_content() or iter_lines()")
print("3. Process chunks incrementally")
print("")
print("Use cases:")
print(" - Live logs")
print(" - Large JSON exports")
print(" - Video streaming")
print(" - Binary downloads")
# ✅ Expected output:
# Streaming API pattern:
# 1. Set stream=True in request
# 2. Use iter_content() or iter_lines()
# 3. Process chunks incrementally
#
# Use cases:
# - Live logs
# - Large JSON exports
# - Video streaming
# - Binary downloads4. Chunked CSV with Pandas
# Pattern for large CSVs with pandas
# import pandas as pd
# for chunk in pd.read_csv("big.csv", chunksize=100_000):
# process(chunk)
print("Pandas chunked reading:")
print(" - Only small chunks in memory")
print(" - Fast vectorized operations")
print(" - Ideal for multi-GB datasets")
print("")
print("Parameters:")
print(" chunksize=100000 # rows per chunk")
print(" usecols=[...] # only needed columns")
print(" dtype={...} # optimize memory")
# ✅ Expected output:
# Pandas chunked reading:
# - Only small chunks in memory
# - Fast vectorized operations
# - Ideal for multi-GB datasets
#
# Parameters:
# chunksize=100000 # rows per chunk
# usecols=[...] # only needed columns
# dtype={...} # optimize memory5. Producer/Consumer with Backpressure
from queue import Queue
from threading import Thread
import time
q = Queue(maxsize=5) # Limit queue size
def producer():
for i in range(10):
print(f"Producing {i}")
q.put(i)
time.sleep(0.1)
q.put(None) # Signal done
def consumer():
while True:
item = q.get()
if item is None:
break
print(f" Consuming {item}")
time.sleep(0.2)
q.task_done()
# Run pipeline
t1 = Thread(target=producer)
t2 = Thread(target=consumer)
t1.start()
t2.start()
t1.join()
t2.join()
print("Done!")6. Live Log Tailing
import time
def follow(lines, delay=0.1):
"""Simulate tail -f behavior."""
for line in lines:
yield line
time.sleep(delay)
# Demo with sample data
log_entries = [
"[INFO] Server started",
"[DEBUG] Connection opened",
"[ERROR] Timeout on request",
"[INFO] Retry successful",
]
print("Tailing log file...")
for line in follow(log_entries):
print(line)
# ✅ Expected output:
# Tailing log file...
# [INFO] Server started
# [DEBUG] Connection opened
# [ERROR] Timeout on request
# [INFO] Retry successful7. Temporary Files
import tempfile
import os
# Create temporary file
with tempfile.NamedTemporaryFile(mode='w', delete=False, suffix='.txt') as tmp:
tmp.write("Temporary data")
tmp_path = tmp.name
print(f"Created: {tmp_path}")
# Read it back
with open(tmp_path) as f:
print(f"Content: {f.read()}")
# Clean up
os.unlink(tmp_path)
print("Cleaned up!")8. Multiprocessing for Parallel Files
from multiprocessing import Pool
def process_file(filename):
"""Process a single file."""
# Simulate processing
return f"Processed: {filename}"
# Files to process
files = ["data1.csv", "data2.csv", "data3.csv"]
# Process in parallel
with Pool(processes=3) as pool:
results = pool.map(process_file, files)
for result in results:
print(result)
# ✅ Expected output:
# Processed: data1.csv
# Processed: data2.csv
# Processed: data3.csv9. Binary Struct Parsing
import struct
# Define record format: int, float, float
fmt = "I f f" # unsigned int, 2 floats
size = struct.calcsize(fmt)
# Create sample binary data
records = [
struct.pack(fmt, 1, 23.5, 65.2),
struct.pack(fmt, 2, 24.1, 68.0),
struct.pack(fmt, 3, 22.8, 70.5),
]
data = b"".join(records)
# Parse records
print(f"Record size: {size} bytes")
for i in range(0, len(data), size):
sensor_id, temp, humidity = struct.unpack(fmt, data[i:i+size])
print(f"Sensor {sensor_id}: temp={temp:.1f}, humidity={humidity:.1f}")
# ✅ Expected output:
# Record size: 12 bytes
# Sensor 1: temp=23.5, humidity=65.2
# Sensor 2: temp=24.1, humidity=68.0
# Sensor 3: temp=22.8, humidity=70.510. Composable Pipeline Functions
def read_data(items):
for item in items:
yield item
def strip_whitespace(items):
for item in items:
yield item.strip()
def filter_empty(items):
for item in items:
if item:
yield item
def uppercase(items):
for item in items:
yield item.upper()
# Compose pipeline
data = [" hello ", "", " world ", " python ", ""]
pipeline = uppercase(filter_empty(strip_whitespace(read_data(data))))
print("Pipeline output:")
for item in pipeline:
print(f" {item}")
# ✅ Expected output:
# Pipeline output:
# HELLO
# WORLD
# PYTHON11. Avoiding Common Streaming Pitfalls
- ❌ Forgetting .close() outside a with
- ❌ Inefficient small reads
- ❌ Building giant Python lists
- ❌ Using pandas on files too large
- ❌ Decompressing entire files unnecessarily
- ❌ Excessive text splits / regex in hot loops
- ❌ Ignoring buffer sizes
- ❌ Reading binary as text
Part 3: High-Volume File & Data Processing
1. Streaming Compressed Files
import gzip
import io
# Create compressed data
original = "\n".join([f"Log line {i}" for i in range(100)])
compressed = gzip.compress(original.encode())
print(f"Original: {len(original)} bytes")
print(f"Compressed: {len(compressed)} bytes")
# Stream it
with gzip.open(io.BytesIO(compressed), "rt") as f:
count = 0
for line in f:
count += 1
print(f"Streamed {count} lines")
# ✅ Expected output:
# Original: 1189 bytes
# Compressed: 222 bytes
# Streamed 100 lines2. Zero-Copy with Memoryview
# Memoryview enables zero-copy slicing
data = bytearray(b"HEADER" + b"X" * 1000 + b"FOOTER")
# Create memoryview (no copy!)
mv = memoryview(data)
# Slice without allocating new memory
header = mv[:6]
footer = mv[-6:]
print(f"Header: {bytes(header)}")
print(f"Footer: {bytes(footer)}")
print(f"Total bytes: {len(data)}")
print("No memory was copied!")
# ✅ Expected output:
# Header: b'HEADER'
# Footer: b'FOOTER'
# Total bytes: 1012
# No memory was copied!3. Async File I/O
import asyncio
async def process_data():
"""Simulate async file processing."""
# With aiofiles:
# async with aiofiles.open("file.txt") as f:
# async for line in f:
# await process(line)
items = ["item1", "item2", "item3"]
for item in items:
await asyncio.sleep(0.1)
print(f"Processed: {item}")
asyncio.run(process_data())
print("Async processing complete!")
# ✅ Expected output:
# Processed: item1
# Processed: item2
# Processed: item3
# Async processing complete!4. Multiprocessing with imap
from multiprocessing import Pool
def transform(x):
"""CPU-intensive transformation."""
return x ** 2
data = range(10)
# imap_unordered streams results as they complete
with Pool(4) as p:
results = list(p.imap_unordered(transform, data, chunksize=2))
print("Results:", sorted(results))
# ✅ Expected output:
# Results: [0, 1, 4, 9, 16, 25, 36, 49, 64, 81]5. Rolling Window Processing
from collections import deque
def rolling_average(data, window_size):
"""Calculate rolling average."""
window = deque(maxlen=window_size)
for value in data:
window.append(value)
if len(window) == window_size:
avg = sum(window) / window_size
yield avg
# Demo
prices = [100, 102, 98, 103, 105, 101, 99, 104]
print("5-period rolling average:")
for avg in rolling_average(prices, 5):
print(f" {avg:.2f}")
# ✅ Expected output:
# 5-period rolling average:
# 101.60
# 101.80
# 101.20
# 102.406. Endless Streaming Pipelines
def generate():
"""Generate endless data."""
for i in range(1000):
yield f"record_{i}"
def clean(records):
for r in records:
yield r.strip()
def enrich(records):
for r in records:
yield {"id": r, "processed": True}
def take(records, n):
for i, r in enumerate(records):
if i >= n:
break
yield r
# Build pipeline (no memory growth!)
pipeline = take(enrich(clean(generate())), 5)
print("First 5 records:")
for record in pipeline:
print(f" {record}")
# ✅ Expected output:
# First 5 records:
# {'id': 'record_0', 'processed': True}
# {'id': 'record_1', 'processed': True}
# {'id': 'record_2', 'processed': True}
# {'id': 'record_3', 'processed': True}
# {'id': 'record_4', 'processed': True}7. Atomic File Writes
import tempfile
import os
def atomic_write(path, data):
"""Write atomically to prevent corruption."""
# Write to temp file first
dir_name = os.path.dirname(path) or "."
fd, tmp_path = tempfile.mkstemp(dir=dir_name)
try:
with os.fdopen(fd, 'w') as f:
f.write(data)
# Atomic rename
os.replace(tmp_path, path)
print(f"Atomically wrote to {path}")
except:
os.unlink(tmp_path)
raise
# Demo
atomic_write("/tmp/test_atomic.txt", "Important data!")8. Common Pitfalls in Large-Dataset Work
- ❌ Naïve .read() on giant files
- ❌ Unbounded queues
- ❌ Excessive process spawning
- ❌ Forgetting to chunk operations
- ❌ Mixing CPU & I/O tasks in same thread
- ❌ Converting everything to pandas DataFrame
- ❌ Relying on recursion for streaming
🎓 Final Summary
You've now mastered professional file handling and large-scale data processing in Python.
- Stream massive files without loading them into memory
- Build efficient data pipelines with generators
- Process compressed and encrypted data on the fly
- Handle binary formats and structured records
- Use memory mapping for fast random access
- Implement multiprocessing for parallel file processing
- Work with columnar formats like Parquet
- Build real-time log monitoring systems
- Handle terabyte-scale datasets efficiently
📋 Quick Reference — Files & Streams
| Syntax | What it does |
|---|---|
| with open(path, 'rb') as f: | Open file in binary read mode |
| pathlib.Path(path) | Modern path manipulation |
| for chunk in iter(f.read, b''): | Read large file in chunks |
| csv.DictReader(f) | Read CSV as list of dicts |
| io.StringIO(data) | In-memory file-like object |
🎉 Great work! You've completed this lesson.
You can now handle files of any size efficiently, parse CSV/JSON/binary formats, and stream large datasets without memory issues.
Practice quiz
Why avoid f.read() on a very large file?
- It is slower to type
- It corrupts the file
- It loads the entire file into memory, which can exhaust RAM
- It only reads the first line
Answer: It loads the entire file into memory, which can exhaust RAM. f.read() pulls the whole file into RAM at once — disastrous for multi-GB files. Stream line by line instead.
What is the memory-efficient way to process a huge text file line by line?
- for line in f:
- data = f.read().split()
- rows = list(f)
- f.readlines()
Answer: for line in f:. Iterating 'for line in f:' reads one line at a time, keeping memory usage tiny regardless of file size.
Why is 'with open(path) as f:' preferred for file handling?
- It is faster to parse
- It enables binary mode
- It compresses the file
- It automatically closes the file even if an error occurs
Answer: It automatically closes the file even if an error occurs. The with statement (context manager) guarantees the file is closed when the block exits, even on exceptions.
What does the file mode 'rb' mean?
- Read backwards
- Read binary
- Replace bytes
- Read buffered
Answer: Read binary. 'r' is read and 'b' is binary, so 'rb' opens a file for reading in binary mode — used for non-text data like videos.
Why should you NOT do 'rows = list(reader)' on a large CSV?
- It loads every row into memory at once
- list is not callable
- csv.reader has no list support
- It sorts the rows
Answer: It loads every row into memory at once. Wrapping the reader in list() materializes every row in RAM. Iterate the reader row by row to stay memory-efficient.
What advantage does mmap (memory mapping) give for huge files?
- It compresses the file
- It encrypts the data
- Instant random access to any position with near-zero RAM usage
- It sorts the bytes
Answer: Instant random access to any position with near-zero RAM usage. Memory mapping lets the OS provide fast random access to any offset without loading the whole file into RAM.
In the lesson, what does f.seek(116) do to a file object before f.read(6)?
- Reads 116 bytes
- Moves the read position to byte offset 116
- Deletes 116 bytes
- Skips 6 lines
Answer: Moves the read position to byte offset 116. seek(116) moves the file pointer to byte offset 116, so the following read(6) starts from there (reading 'TARGET' in the example).
Why use a bounded queue (Queue(maxsize=5)) in a producer/consumer pipeline?
- To sort items
- To run faster
- To deduplicate items
- To apply backpressure and limit memory use
Answer: To apply backpressure and limit memory use. A bounded queue blocks the producer when full, applying backpressure so memory stays under control.
What does a deque(maxlen=N) do as you append beyond N items?
- Raises an error
- Automatically drops items from the opposite end, keeping the last N
- Ignores new items
- Doubles in size
Answer: Automatically drops items from the opposite end, keeping the last N. A deque with maxlen automatically discards the oldest item when full — perfect for rolling/sliding-window processing.
What is the benefit of an atomic write (write to temp file, then os.replace)?
- Faster writes
- It compresses the output
- It prevents file corruption by swapping in the complete file in one step
- It appends instead of overwriting
Answer: It prevents file corruption by swapping in the complete file in one step. Writing to a temp file and atomically renaming means readers never see a half-written file — the replace is all-or-nothing.
Continue this course
- Previous: Packaging & Publishing Python Libraries to PyPI
- Next: Advanced Collections, itertools & functools — Use Counter, defaultdict, chain, groupby, partial and more
- Quick reference: Python cheat sheet