Files, Streams & Large Datasets

Reviewed & published by Brayan K

Master efficient file handling, streaming patterns, and processing massive datasets that don't fit in memory.

Part of the free Python course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

📊 What You'll Learn

This lesson teaches professional data handling techniques used in:

You'll learn how to:

📥 Python Download & Setup

Download Python from: python.org/downloads

Latest version recommended (3.11+)

Part 1: File I/O Fundamentals & Streaming Basics

1. Python's File I/O Model (Text vs Binary, Buffered vs Unbuffered)

When you open a file, Python uses a layered I/O system:

LayerWhat It DoesSpeed
Raw I/ODirect interaction with OS file descriptorsSlowest
Buffered I/OUses memory buffers (reads/writes in chunks)Fast
Text I/ODecodes bytes → Unicode stringsAutomatic encoding
with open("data.txt", "r", encoding="utf-8") as f:
    content = f.read()
    print(content)

Modes:

Buffering matters: Python reads data in chunks internally to reduce syscalls, making it ideal for large datasets.

2. Reading Files the Right Way (Avoid .read() on Large Files)

data = open("log.txt").read()
print(data[:100])

This loads the entire file into memory, which is disastrous for large logs.

with open("log.txt") as f:
    for line in f:
        # Process each line without loading full file
        print(line.strip())
ApproachMemory UsageFor 10GB File
f.read()Entire file in RAM10GB+ RAM needed 💥
for line in f:One line at a time~KB of RAM ✓

Advantages:

🟢 Worked Example: write a file, then stream it back

The snippets above open files like data.txt and log.txt that have to already exist — run one of them on a machine without that file and you get FileNotFoundError. So this example creates its own file first, then reads it back four different ways, then deletes it. It is completely self-contained: run it anywhere and it works.

# 🟢 WORKED EXAMPLE — create a file, then stream it back, all in one script
from pathlib import Path

LOG = Path("visits.log")     # Path is the modern way to name a file; no string juggling

lines = [
    "2026-05-01 GET /home 200",
    "2026-05-01 GET /pricing 200",
    "2026-05-01 GET /admin 403",
    "2026-05-01 GET /home 200",
    "2026-05-01 GET /checkout 500",
]

# ---------- 1. Write it ----------
# "w" creates the file, or empties it if it already exists. The with block closes
# the file for you the moment you leave it — even if an error is raised inside.
with open(LOG, "w", encoding="utf-8") as f:
    for line in lines:
        f.write(line + "\n")     # write() adds NO newline of its own; that is your job

print("bytes on disk:", LOG.stat().st_size)

# ---------- 2. Read the lot in one gulp ----------
# Fine for 133 bytes. On a 10 GB log this is the line that kills the process.
with open(LOG, encoding="utf-8") as f:      # no mode given, so "r" (read) is assumed
    whole = f.read()
print("characters read:", len(whole))
print("first line:", whole.splitlines()[0])

# ---------- 3. Stream it: one line in memory at a time ----------
# Looping over the file object itself is the memory-safe version of the above.
statuses = {}
with open(LOG, encoding="utf-8") as f:
    for line in f:
        code = line.strip().split()[-1]              # strip the newline, take the last field
        statuses[code] = statuses.get(code, 0) + 1   # tally it up
print("status counts:", statuses)

# ---------- 4. Append without destroying what is there ----------
# "a" starts writing at the end. Using "w" here would wipe the whole file.
with open(LOG, "a", encoding="utf-8") as f:
    f.write("2026-05-01 GET /pricing 200\n")

with open(LOG, encoding="utf-8") as f:
    print("line count now:", sum(1 for _ in f))      # counts lines without storing any

# ---------- 5. Tidy up after yourself ----------
LOG.unlink()                                          # delete the file
print("file still there?", LOG.exists())

# ✅ Expected output:
# bytes on disk: 133
# characters read: 133
# first line: 2026-05-01 GET /home 200
# status counts: {'200': 3, '403': 1, '500': 1}
# line count now: 6
# file still there? False

Two things worth pausing on. The byte count and the character count match here only because every character is plain ASCII — put a £ or an emoji in that file and UTF-8 spends more than one byte on it, so the numbers separate. And sum(1 for _ in f) counts six lines while holding exactly one in memory: that is the whole streaming idea in a single line.

🎯 Your Turn: write and re-read a to-do file

Three blanks, each one a detail that beginners get wrong on their first file. Fill them in and run it.

# 🎯 YOUR TURN — fill in the blanks marked with ___
from pathlib import Path

NOTES = Path("notes.txt")
tasks = ["buy milk", "call dentist", "finish lesson"]

# 1) Open the file for writing — the mode that creates it fresh.
with open(NOTES, "___", encoding="utf-8") as f:    # 👉 "r", "w" or "a"?
    for task in tasks:
        # 2) write() never adds a line break. Put one on the end yourself.
        f.write(task + "___")                      # 👉 replace ___ with \n (backslash n)

# Read it back, numbering each line as it streams past.
with open(NOTES, encoding="utf-8") as f:
    for number, line in enumerate(f, start=1):     # start=1 so humans count from 1
        # 3) Each line still carries its trailing newline. Remove it.
        print(f"{number}. {line.___()}")           # 👉 which string method trims whitespace?

with open(NOTES, encoding="utf-8") as f:
    print("Total lines:", sum(1 for _ in f))

NOTES.unlink()          # clean up so re-running starts from scratch

# ✅ Expected output:
# 1. buy milk
# 2. call dentist
# 3. finish lesson
# Total lines: 3

If your numbers come out with blank lines between them, blank 3 is the culprit — the newline from the file plus the newline from print() gives you two.

3. Using Buffered Streams Explicitly

For very large binary files (videos, models, audio):

from io import BufferedReader

with open("video.mp4", "rb") as f:
    stream = BufferedReader(f)
    chunk = stream.read(4096)
    print(f"Read {len(chunk)} bytes")

4. Writing Files Efficiently (Avoid Tiny Writes)

Frequent small writes slow everything down.

from io import StringIO

items = ["apple", "banana", "cherry", "date"]

buffer = StringIO()
for item in items:
    buffer.write(item + "\n")

with open("output.txt", "w") as f:
    f.write(buffer.getvalue())

print("Written to file!")

# ✅ Expected output:
# Written to file!

5. Working With CSV Files at Scale

Python's csv module supports streaming:

import csv

with open("data.csv") as f:
    reader = csv.reader(f)
    for row in reader:
        # Process each row
        print(row)

Never convert a large CSV into a list:

🏁 Mini-Challenge: a CSV round trip

Brief only this time. Write three sales rows to a CSV file, then stream them back and work out the takings — without ever building a list of every row.

Two things to look up rather than guess: csv.DictWriter needs a fieldnames list and a call to writeheader(), and every value that comes back out of csv.DictReader is a string, so you must convert with int() before doing any arithmetic.

# 🏁 MINI-CHALLENGE — write it yourself
import csv
from pathlib import Path

SALES = Path("sales.csv")
rows = [
    {"item": "tea", "units": 3, "pence": 120},
    {"item": "coffee", "units": 2, "pence": 275},
    {"item": "cake", "units": 1, "pence": 340},
]

# 1. Write rows to SALES with csv.DictWriter.
#    open(SALES, "w", newline="", encoding="utf-8")  <- newline="" is required on
#    Windows or you get a blank line between every row.
#    fieldnames=["item", "units", "pence"], then writeheader(), then writerows(rows).

# 2. Print the raw file so you can see what a CSV actually is:
#    print(SALES.read_text(encoding="utf-8"))

# 3. Stream it back with csv.DictReader. For each row:
#    line_total = int(row["units"]) * int(row["pence"])
#    add it to a running total and print  item: units x pence = line_total p

# 4. Print the grand total in pounds, 2 decimal places.

# 5. SALES.unlink() to clean up.

# your code here


# ✅ Expected output (the blank line after the raw CSV is real — the file
#    ends in a newline and print() adds another):
# item,units,pence
# tea,3,120
# coffee,2,275
# cake,1,340
#
# tea: 3 x 120p = 360p
# coffee: 2 x 275p = 550p
# cake: 1 x 340p = 340p
# Total: £12.50

6. Handling JSON Without Crashing RAM

import json

# For line-delimited JSON (NDJSON)
# Each line is a separate JSON object
sample_ndjson = '''{"name": "Alice", "age": 30}
{"name": "Bob", "age": 25}
{"name": "Charlie", "age": 35}'''

for line in sample_ndjson.strip().split("\n"):
    item = json.loads(line)
    print(f"Processing: {item['name']}")

# ✅ Expected output:
# Processing: Alice
# Processing: Bob
# Processing: Charlie

7. Working With Very Large Binary Files

# Simulate chunked reading
data = b"X" * 5000  # 5KB of data

chunk_size = 1024
offset = 0

while offset < len(data):
    chunk = data[offset:offset + chunk_size]
    print(f"Processing chunk at offset {offset}, size {len(chunk)}")
    offset += chunk_size

# ✅ Expected output:
# Processing chunk at offset 0, size 1024
# Processing chunk at offset 1024, size 1024
# Processing chunk at offset 2048, size 1024
# Processing chunk at offset 3072, size 1024
# Processing chunk at offset 4096, size 904

8. Memory Mapping (mmap) — Fast Random Access to Huge Files

import mmap

# Simulated memory-mapped access pattern
# In real use: mm = mmap.mmap(f.fileno(), 0)
print("Memory mapping allows:")
print("- Instant random access to any position")
print("- No RAM explosion for huge files")
print("- Extremely fast reads")
print("- Ideal for ML datasets and databases")

# Example pattern:
# with open("huge.dat", "rb") as f:
#     mm = mmap.mmap(f.fileno(), 0, access=mmap.ACCESS_READ)
#     print(mm[1000:1010])  # read 10 bytes at offset 1000

# ✅ Expected output:
# Memory mapping allows:
# - Instant random access to any position
# - No RAM explosion for huge files
# - Extremely fast reads
# - Ideal for ML datasets and databases
MethodRandom Access SpeedMemory Usage
f.seek() + f.read()MediumLow (read size)
f.read() entire fileFast (after load)File size 💥
mmapInstantNear zero

9. Chunk Processing for Massive Datasets

def read_in_chunks(data, size=1024):
    """Generator that yields chunks of data."""
    offset = 0
    while offset < len(data):
        yield data[offset:offset + size]
        offset += size

# Demo with sample data
sample = "X" * 5000

for i, chunk in enumerate(read_in_chunks(sample, 1000)):
    print(f"Chunk {i}: {len(chunk)} chars")

# ✅ Expected output:
# Chunk 0: 1000 chars
# Chunk 1: 1000 chars
# Chunk 2: 1000 chars
# Chunk 3: 1000 chars
# Chunk 4: 1000 chars

10. Using Iterators & Generators in Data Pipelines

def read_lines(lines):
    for line in lines:
        yield line

def filter_errors(lines):
    for line in lines:
        if "ERROR" in line:
            yield line

# Sample log data
log_lines = [
    "INFO: Starting service",
    "ERROR: Connection failed",
    "INFO: Retrying...",
    "ERROR: Timeout occurred",
    "INFO: Service stopped"
]

for line in filter_errors(read_lines(log_lines)):
    print(line)

# ✅ Expected output:
# ERROR: Connection failed
# ERROR: Timeout occurred

11. Working With Compressed Data (ZIP, GZIP)

import gzip
import io

# Create sample gzipped data
sample_text = "Hello, compressed world!\nLine 2\nLine 3"
compressed = gzip.compress(sample_text.encode())

# Read it back (streaming pattern)
with gzip.open(io.BytesIO(compressed), "rt") as f:
    for line in f:
        print(line.strip())

# ✅ Expected output:
# Hello, compressed world!
# Line 2
# Line 3

12. Handling Data That Doesn't Fit In RAM

# Pattern for counting lines in a 10GB file
# Memory usage stays near 0MB!

lines = [f"Line {i}" for i in range(100000)]

total = 0
for _ in lines:
    total += 1

print(f"Total lines: {total:,}")

# ✅ Expected output:
# Total lines: 100,000

13. Avoiding the Most Common File-I/O Mistakes

14. Best Practices Recap

Part 2: Advanced Streaming & Dataset Engineering

1. Understanding Buffered I/O Depth

from io import FileIO, BufferedReader

# Python's I/O layers:
# 1. Raw I/O (FileIO) - OS file descriptors
# 2. Buffered I/O - memory buffers
# 3. Text I/O - decoding/encoding

print("I/O Layer Architecture:")
print("  Raw I/O    → direct OS interaction")
print("  Buffered   → efficient chunking")
print("  Text I/O   → Unicode handling")

# Larger buffer sizes improve throughput
# with open("large.bin", "rb") as f:
#     reader = BufferedReader(f, buffer_size=1024*1024)

# ✅ Expected output:
# I/O Layer Architecture:
#   Raw I/O    → direct OS interaction
#   Buffered   → efficient chunking
#   Text I/O   → Unicode handling

2. Random Access With Seek

import io

# Create sample file-like object
data = b"HEADER_DATA_HERE" + b"X" * 100 + b"TARGET"
file = io.BytesIO(data)

# Jump to position
file.seek(116)  # Skip to "TARGET"
print(f"Read at position 116: {file.read(6)}")

# Go back to start
file.seek(0)
print(f"Header: {file.read(16)}")

# ✅ Expected output:
# Read at position 116: b'TARGET'
# Header: b'HEADER_DATA_HERE'

3. Streaming API Data

# Pattern for streaming API data
# Prevents loading giant payloads into RAM

# import requests
# with requests.get(url, stream=True) as r:
#     for chunk in r.iter_content(1024):
#         process(chunk)

print("Streaming API pattern:")
print("1. Set stream=True in request")
print("2. Use iter_content() or iter_lines()")
print("3. Process chunks incrementally")
print("")
print("Use cases:")
print("  - Live logs")
print("  - Large JSON exports")
print("  - Video streaming")
print("  - Binary downloads")

# ✅ Expected output:
# Streaming API pattern:
# 1. Set stream=True in request
# 2. Use iter_content() or iter_lines()
# 3. Process chunks incrementally
#
# Use cases:
#   - Live logs
#   - Large JSON exports
#   - Video streaming
#   - Binary downloads

4. Chunked CSV with Pandas

# Pattern for large CSVs with pandas
# import pandas as pd
# for chunk in pd.read_csv("big.csv", chunksize=100_000):
#     process(chunk)

print("Pandas chunked reading:")
print("  - Only small chunks in memory")
print("  - Fast vectorized operations")
print("  - Ideal for multi-GB datasets")
print("")
print("Parameters:")
print("  chunksize=100000  # rows per chunk")
print("  usecols=[...]     # only needed columns")
print("  dtype={...}       # optimize memory")

# ✅ Expected output:
# Pandas chunked reading:
#   - Only small chunks in memory
#   - Fast vectorized operations
#   - Ideal for multi-GB datasets
#
# Parameters:
#   chunksize=100000  # rows per chunk
#   usecols=[...]     # only needed columns
#   dtype={...}       # optimize memory

5. Producer/Consumer with Backpressure

from queue import Queue
from threading import Thread
import time

q = Queue(maxsize=5)  # Limit queue size

def producer():
    for i in range(10):
        print(f"Producing {i}")
        q.put(i)
        time.sleep(0.1)
    q.put(None)  # Signal done

def consumer():
    while True:
        item = q.get()
        if item is None:
            break
        print(f"  Consuming {item}")
        time.sleep(0.2)
        q.task_done()

# Run pipeline
t1 = Thread(target=producer)
t2 = Thread(target=consumer)
t1.start()
t2.start()
t1.join()
t2.join()
print("Done!")

6. Live Log Tailing

import time

def follow(lines, delay=0.1):
    """Simulate tail -f behavior."""
    for line in lines:
        yield line
        time.sleep(delay)

# Demo with sample data
log_entries = [
    "[INFO] Server started",
    "[DEBUG] Connection opened",
    "[ERROR] Timeout on request",
    "[INFO] Retry successful",
]

print("Tailing log file...")
for line in follow(log_entries):
    print(line)

# ✅ Expected output:
# Tailing log file...
# [INFO] Server started
# [DEBUG] Connection opened
# [ERROR] Timeout on request
# [INFO] Retry successful

7. Temporary Files

import tempfile
import os

# Create temporary file
with tempfile.NamedTemporaryFile(mode='w', delete=False, suffix='.txt') as tmp:
    tmp.write("Temporary data")
    tmp_path = tmp.name
    print(f"Created: {tmp_path}")

# Read it back
with open(tmp_path) as f:
    print(f"Content: {f.read()}")

# Clean up
os.unlink(tmp_path)
print("Cleaned up!")

8. Multiprocessing for Parallel Files

from multiprocessing import Pool

def process_file(filename):
    """Process a single file."""
    # Simulate processing
    return f"Processed: {filename}"

# Files to process
files = ["data1.csv", "data2.csv", "data3.csv"]

# Process in parallel
with Pool(processes=3) as pool:
    results = pool.map(process_file, files)

for result in results:
    print(result)

# ✅ Expected output:
# Processed: data1.csv
# Processed: data2.csv
# Processed: data3.csv

9. Binary Struct Parsing

import struct

# Define record format: int, float, float
fmt = "I f f"  # unsigned int, 2 floats
size = struct.calcsize(fmt)

# Create sample binary data
records = [
    struct.pack(fmt, 1, 23.5, 65.2),
    struct.pack(fmt, 2, 24.1, 68.0),
    struct.pack(fmt, 3, 22.8, 70.5),
]
data = b"".join(records)

# Parse records
print(f"Record size: {size} bytes")
for i in range(0, len(data), size):
    sensor_id, temp, humidity = struct.unpack(fmt, data[i:i+size])
    print(f"Sensor {sensor_id}: temp={temp:.1f}, humidity={humidity:.1f}")

# ✅ Expected output:
# Record size: 12 bytes
# Sensor 1: temp=23.5, humidity=65.2
# Sensor 2: temp=24.1, humidity=68.0
# Sensor 3: temp=22.8, humidity=70.5

10. Composable Pipeline Functions

def read_data(items):
    for item in items:
        yield item

def strip_whitespace(items):
    for item in items:
        yield item.strip()

def filter_empty(items):
    for item in items:
        if item:
            yield item

def uppercase(items):
    for item in items:
        yield item.upper()

# Compose pipeline
data = ["  hello  ", "", "  world  ", "  python  ", ""]
pipeline = uppercase(filter_empty(strip_whitespace(read_data(data))))

print("Pipeline output:")
for item in pipeline:
    print(f"  {item}")

# ✅ Expected output:
# Pipeline output:
#   HELLO
#   WORLD
#   PYTHON

11. Avoiding Common Streaming Pitfalls

Part 3: High-Volume File & Data Processing

1. Streaming Compressed Files

import gzip
import io

# Create compressed data
original = "\n".join([f"Log line {i}" for i in range(100)])
compressed = gzip.compress(original.encode())

print(f"Original: {len(original)} bytes")
print(f"Compressed: {len(compressed)} bytes")

# Stream it
with gzip.open(io.BytesIO(compressed), "rt") as f:
    count = 0
    for line in f:
        count += 1
    print(f"Streamed {count} lines")

# ✅ Expected output:
# Original: 1189 bytes
# Compressed: 222 bytes
# Streamed 100 lines

2. Zero-Copy with Memoryview

# Memoryview enables zero-copy slicing
data = bytearray(b"HEADER" + b"X" * 1000 + b"FOOTER")

# Create memoryview (no copy!)
mv = memoryview(data)

# Slice without allocating new memory
header = mv[:6]
footer = mv[-6:]

print(f"Header: {bytes(header)}")
print(f"Footer: {bytes(footer)}")
print(f"Total bytes: {len(data)}")
print("No memory was copied!")

# ✅ Expected output:
# Header: b'HEADER'
# Footer: b'FOOTER'
# Total bytes: 1012
# No memory was copied!

3. Async File I/O

import asyncio

async def process_data():
    """Simulate async file processing."""
    # With aiofiles:
    # async with aiofiles.open("file.txt") as f:
    #     async for line in f:
    #         await process(line)
    
    items = ["item1", "item2", "item3"]
    for item in items:
        await asyncio.sleep(0.1)
        print(f"Processed: {item}")

asyncio.run(process_data())
print("Async processing complete!")

# ✅ Expected output:
# Processed: item1
# Processed: item2
# Processed: item3
# Async processing complete!

4. Multiprocessing with imap

from multiprocessing import Pool

def transform(x):
    """CPU-intensive transformation."""
    return x ** 2

data = range(10)

# imap_unordered streams results as they complete
with Pool(4) as p:
    results = list(p.imap_unordered(transform, data, chunksize=2))

print("Results:", sorted(results))

# ✅ Expected output:
# Results: [0, 1, 4, 9, 16, 25, 36, 49, 64, 81]

5. Rolling Window Processing

from collections import deque

def rolling_average(data, window_size):
    """Calculate rolling average."""
    window = deque(maxlen=window_size)
    
    for value in data:
        window.append(value)
        if len(window) == window_size:
            avg = sum(window) / window_size
            yield avg

# Demo
prices = [100, 102, 98, 103, 105, 101, 99, 104]

print("5-period rolling average:")
for avg in rolling_average(prices, 5):
    print(f"  {avg:.2f}")

# ✅ Expected output:
# 5-period rolling average:
#   101.60
#   101.80
#   101.20
#   102.40

6. Endless Streaming Pipelines

def generate():
    """Generate endless data."""
    for i in range(1000):
        yield f"record_{i}"

def clean(records):
    for r in records:
        yield r.strip()

def enrich(records):
    for r in records:
        yield {"id": r, "processed": True}

def take(records, n):
    for i, r in enumerate(records):
        if i >= n:
            break
        yield r

# Build pipeline (no memory growth!)
pipeline = take(enrich(clean(generate())), 5)

print("First 5 records:")
for record in pipeline:
    print(f"  {record}")

# ✅ Expected output:
# First 5 records:
#   {'id': 'record_0', 'processed': True}
#   {'id': 'record_1', 'processed': True}
#   {'id': 'record_2', 'processed': True}
#   {'id': 'record_3', 'processed': True}
#   {'id': 'record_4', 'processed': True}

7. Atomic File Writes

import tempfile
import os

def atomic_write(path, data):
    """Write atomically to prevent corruption."""
    # Write to temp file first
    dir_name = os.path.dirname(path) or "."
    fd, tmp_path = tempfile.mkstemp(dir=dir_name)
    
    try:
        with os.fdopen(fd, 'w') as f:
            f.write(data)
        # Atomic rename
        os.replace(tmp_path, path)
        print(f"Atomically wrote to {path}")
    except:
        os.unlink(tmp_path)
        raise

# Demo
atomic_write("/tmp/test_atomic.txt", "Important data!")

8. Common Pitfalls in Large-Dataset Work

🎓 Final Summary

You've now mastered professional file handling and large-scale data processing in Python.

📋 Quick Reference — Files & Streams

SyntaxWhat it does
with open(path, 'rb') as f:Open file in binary read mode
pathlib.Path(path)Modern path manipulation
for chunk in iter(f.read, b''):Read large file in chunks
csv.DictReader(f)Read CSV as list of dicts
io.StringIO(data)In-memory file-like object

🎉 Great work! You've completed this lesson.

You can now handle files of any size efficiently, parse CSV/JSON/binary formats, and stream large datasets without memory issues.

Practice quiz

Why avoid f.read() on a very large file?

  • It is slower to type
  • It corrupts the file
  • It loads the entire file into memory, which can exhaust RAM
  • It only reads the first line

Answer: It loads the entire file into memory, which can exhaust RAM. f.read() pulls the whole file into RAM at once — disastrous for multi-GB files. Stream line by line instead.

What is the memory-efficient way to process a huge text file line by line?

  • for line in f:
  • data = f.read().split()
  • rows = list(f)
  • f.readlines()

Answer: for line in f:. Iterating 'for line in f:' reads one line at a time, keeping memory usage tiny regardless of file size.

Why is 'with open(path) as f:' preferred for file handling?

  • It is faster to parse
  • It enables binary mode
  • It compresses the file
  • It automatically closes the file even if an error occurs

Answer: It automatically closes the file even if an error occurs. The with statement (context manager) guarantees the file is closed when the block exits, even on exceptions.

What does the file mode 'rb' mean?

  • Read backwards
  • Read binary
  • Replace bytes
  • Read buffered

Answer: Read binary. 'r' is read and 'b' is binary, so 'rb' opens a file for reading in binary mode — used for non-text data like videos.

Why should you NOT do 'rows = list(reader)' on a large CSV?

  • It loads every row into memory at once
  • list is not callable
  • csv.reader has no list support
  • It sorts the rows

Answer: It loads every row into memory at once. Wrapping the reader in list() materializes every row in RAM. Iterate the reader row by row to stay memory-efficient.

What advantage does mmap (memory mapping) give for huge files?

  • It compresses the file
  • It encrypts the data
  • Instant random access to any position with near-zero RAM usage
  • It sorts the bytes

Answer: Instant random access to any position with near-zero RAM usage. Memory mapping lets the OS provide fast random access to any offset without loading the whole file into RAM.

In the lesson, what does f.seek(116) do to a file object before f.read(6)?

  • Reads 116 bytes
  • Moves the read position to byte offset 116
  • Deletes 116 bytes
  • Skips 6 lines

Answer: Moves the read position to byte offset 116. seek(116) moves the file pointer to byte offset 116, so the following read(6) starts from there (reading 'TARGET' in the example).

Why use a bounded queue (Queue(maxsize=5)) in a producer/consumer pipeline?

  • To sort items
  • To run faster
  • To deduplicate items
  • To apply backpressure and limit memory use

Answer: To apply backpressure and limit memory use. A bounded queue blocks the producer when full, applying backpressure so memory stays under control.

What does a deque(maxlen=N) do as you append beyond N items?

  • Raises an error
  • Automatically drops items from the opposite end, keeping the last N
  • Ignores new items
  • Doubles in size

Answer: Automatically drops items from the opposite end, keeping the last N. A deque with maxlen automatically discards the oldest item when full — perfect for rolling/sliding-window processing.

What is the benefit of an atomic write (write to temp file, then os.replace)?

  • Faster writes
  • It compresses the output
  • It prevents file corruption by swapping in the complete file in one step
  • It appends instead of overwriting

Answer: It prevents file corruption by swapping in the complete file in one step. Writing to a temp file and atomically renaming means readers never see a half-written file — the replace is all-or-nothing.

Continue this course