Advanced Collections, itertools & functools

Reviewed & published by Brayan K

Master Python's most powerful standard library modules for building high-performance, memory-efficient code.

Part of the free Python course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

🧰 What You'll Learn

This lesson covers incredibly powerful modules that advanced developers rely on every day:

Master the tools used in:

📥 Python Download & Setup

Download Python from: python.org/downloads

Latest version recommended (3.11+)

Part 1: Core Modules — collections, itertools, functools

1. collections — High-Performance Container Tools

The collections module provides optimized data structures that outperform normal lists/dicts in many use cases.

1.1 Counter — Counting Made Easy

Best for: word frequency, counting events, histograms, text processing

from collections import Counter

c = Counter("mississippi")
print(c)

# Common operations
print(c.most_common(3))
c.update("moreletters")
print(c)
c.subtract("abc")
print(c)

# ✅ Expected output:
# Counter({'i': 4, 's': 4, 'p': 2, 'm': 1})
# [('i', 4), ('s', 4), ('p', 2)]
# Counter({'s': 5, 'i': 4, 'e': 3, 'm': 2, 'p': 2, 'r': 2, 't': 2, 'o': 1, 'l': 1})
# Counter({'s': 5, 'i': 4, 'e': 3, 'm': 2, 'p': 2, 'r': 2, 't': 2, 'o': 1, 'l': 1, 'a': -1, 'b': -1, 'c': -1})

1.2 defaultdict — Automatic Missing Values

Instead of manually checking if keys exist:

from collections import defaultdict

# Instead of:
# d = {}
# if key not in d:
#     d[key] = []
# d[key].append(value)

# Use defaultdict:
groups = defaultdict(list)
groups["a"].append(1)
groups["a"].append(2)
groups["b"].append(3)

print(dict(groups))

# ✅ Expected output:
# {'a': [1, 2], 'b': [3]}

1.3 deque — Fast Queue / Stack / Sliding Window

from collections import deque

dq = deque(maxlen=5)
for i in range(10):
    dq.append(i)
    print(f"After adding {i}: {list(dq)}")

print(f"\nFinal deque: {list(dq)}")

# ✅ Expected output:
# After adding 0: [0]
# After adding 1: [0, 1]
# After adding 2: [0, 1, 2]
# After adding 3: [0, 1, 2, 3]
# After adding 4: [0, 1, 2, 3, 4]
# After adding 5: [1, 2, 3, 4, 5]
# After adding 6: [2, 3, 4, 5, 6]
# After adding 7: [3, 4, 5, 6, 7]
# After adding 8: [4, 5, 6, 7, 8]
# After adding 9: [5, 6, 7, 8, 9]
#
# Final deque: [5, 6, 7, 8, 9]

1.4 namedtuple — Lightweight, Readable Structs

from collections import namedtuple

Point = namedtuple("Point", "x y")
p = Point(3, 4)

print(f"Point: {p}")
print(f"X coordinate: {p.x}")
print(f"Y coordinate: {p.y}")

# ✅ Expected output:
# Point: Point(x=3, y=4)
# X coordinate: 3
# Y coordinate: 4

Use this when you need:

1.5 OrderedDict (for deterministic ordering)

Mostly replaced by Python 3.7+ dict, but still useful for:

1.6 ChainMap — Layered Configuration

from collections import ChainMap

user_config = {"theme": "dark", "language": "en"}
default_config = {"theme": "light", "language": "en", "timeout": 30}

settings = ChainMap(user_config, default_config)

print(f"Theme: {settings['theme']}")  # From user_config
print(f"Timeout: {settings['timeout']}")  # From default_config

# ✅ Expected output:
# Theme: dark
# Timeout: 30

🟢 Worked Example: all four collections tools on one problem

Individually these look like party tricks. Together they are how you take a pile of raw records apart. Below is one small web-server log analysed four ways — read the comments, then run it and match each printed line to the code that produced it.

# 🟢 WORKED EXAMPLE — four collections tools on one small server log
from collections import Counter, defaultdict, deque, namedtuple

# namedtuple builds a tiny immutable class in one line. The fields get names,
# so you write hit.status instead of hit[1] and can never mix the columns up.
Hit = namedtuple("Hit", "path status ms")     # ms = how long the request took

log = [
    Hit("/home", 200, 12),
    Hit("/login", 200, 40),
    Hit("/home", 200, 9),
    Hit("/admin", 403, 3),
    Hit("/login", 500, 220),
    Hit("/home", 200, 11),
    Hit("/admin", 403, 4),
]

print("First hit:", log[0])       # Hit(path='/home', status=200, ms=12) — prints its own field names
print("Its path :", log[0].path)  # /home

# --- Counter: how often does each value appear? -------------------------------
# Feed it any iterable and it tallies. Here the iterable is a generator
# expression pulling one field out of every record.
paths = Counter(hit.path for hit in log)
print("Hits per path:", paths)          # Counter({'/home': 3, '/login': 2, '/admin': 2})
print("Busiest page :", paths.most_common(1))   # [('/home', 3)] — already sorted for you

# --- defaultdict: group items without the "if key not in d" dance -------------
# defaultdict(list) means: asking for a key that does not exist CREATES an
# empty list there, so .append() always has something to append to.
times = defaultdict(list)
for hit in log:
    times[hit.path].append(hit.ms)
print("Grouped times:", dict(times))    # dict(...) just makes the print tidier

for path, ms_list in sorted(times.items()):        # sorted() = predictable order
    print(f"  {path}: {sum(ms_list) / len(ms_list):.1f} ms average")   # .1f = one decimal

# --- deque(maxlen=N): a rolling window that forgets old items for you ---------
recent = deque(maxlen=3)     # never holds more than 3; the oldest falls off the left
for hit in log:
    recent.append(hit.status)
print("Last 3 statuses:", list(recent))

# --- Counter again, this time on a filtered stream ---------------------------
errors = Counter(hit.path for hit in log if hit.status >= 400)   # 4xx and 5xx only
print("Failing paths:", errors)

# ✅ Expected output:
# First hit: Hit(path='/home', status=200, ms=12)
# Its path : /home
# Hits per path: Counter({'/home': 3, '/login': 2, '/admin': 2})
# Busiest page : [('/home', 3)]
# Grouped times: {'/home': [12, 9, 11], '/login': [40, 220], '/admin': [3, 4]}
#   /admin: 3.5 ms average
#   /home: 10.7 ms average
#   /login: 130.0 ms average
# Last 3 statuses: [500, 200, 403]
# Failing paths: Counter({'/admin': 2, '/login': 1})

Notice what you did not write: no manual counting dictionary, no "if the key is missing, create it first", no slicing a list to keep the last three. That is the whole point of the collections module.

🎯 Your Turn: count and group some orders

Same two tools, your data. Everything is written except the three pieces that actually use Counter and defaultdict. Fill in the blanks and run it — the expected output is at the bottom of the file so you can check yourself.

# 🎯 YOUR TURN — fill in the blanks marked with ___
from collections import Counter, defaultdict

# (country, order_total_in_pounds)
orders = [
    ("UK", 45), ("DE", 30), ("UK", 12),
    ("FR", 80), ("DE", 25), ("UK", 5),
]

# 1) How many orders came from each country?
#    Counter takes an iterable of the things you want tallied — here, just the country.
per_country = Counter(___ for country, total in orders)   # 👉 replace ___ with: country
print("Orders per country:", per_country)

# 2) How much money came from each country?
#    You want a dict whose missing keys start at 0 so += works on the first order.
revenue = defaultdict(___)          # 👉 replace ___ with the type that defaults to 0
for country, total in orders:
    revenue[country] += total
print("Revenue per country:", dict(revenue))

# 3) Which two countries placed the most orders?
print("Top 2 by order count:", per_country.most_common(___))   # 👉 how many do you want back?

# ✅ Expected output:
# Orders per country: Counter({'UK': 3, 'DE': 2, 'FR': 1})
# Revenue per country: {'UK': 62, 'DE': 55, 'FR': 80}
# Top 2 by order count: [('UK', 3), ('DE', 2)]

Watch out for blank 2: defaultdict wants the type itself, not a value — you pass int, not 0. Python then calls int() for you, which returns 0.

2. itertools — High-Performance Iterator Recipes

itertools is one of Python's strongest modules. It eliminates heavy loops and enables efficient pipelines.

2.1 Infinite Iterators

from itertools import count, cycle, repeat

# count: generating sequences
print("count example:")
for n in count(5):
    if n > 10:
        break
    print(n, end=" ")

print("\n\ncycle example (first 6):")
colors = cycle(['red', 'green', 'blue'])
for i, color in enumerate(colors):
    if i >= 6:
        break
    print(color, end=" ")

print("\n\nrepeat example:")
print(list(repeat("A", 3)))

# ✅ Expected output:
# count example:
# 5 6 7 8 9 10 
#
# cycle example (first 6):
# red green blue red green blue 
#
# repeat example:
# ['A', 'A', 'A']

2.2 Combinatorics (product, permutations, combinations)

from itertools import product, permutations, combinations

# Cartesian product
print("product([1,2], ['x','y']):")
for a, b in product([1,2], ["x","y"]):
    print(f"  ({a}, {b})")

# Permutations
print("\npermutations([1,2,3], 2):")
print(list(permutations([1,2,3], 2)))

# Combinations
print("\ncombinations([1,2,3], 2):")
print(list(combinations([1,2,3], 2)))

# ✅ Expected output:
# product([1,2], ['x','y']):
#   (1, x)
#   (1, y)
#   (2, x)
#   (2, y)
#
# permutations([1,2,3], 2):
# [(1, 2), (1, 3), (2, 1), (2, 3), (3, 1), (3, 2)]
#
# combinations([1,2,3], 2):
# [(1, 2), (1, 3), (2, 3)]

2.3 accumulate — Cumulative Operations

from itertools import accumulate
import operator

# Running sum
print("Running sum:", list(accumulate([1,2,3,4])))

# Running product
print("Running product:", list(accumulate([1,2,3,4], operator.mul)))

# Running max
print("Running max:", list(accumulate([3,1,4,1,5,9], max)))

# ✅ Expected output:
# Running sum: [1, 3, 6, 10]
# Running product: [1, 2, 6, 24]
# Running max: [3, 3, 4, 4, 5, 9]

2.4 groupby — Group Consecutive Items

from itertools import groupby

print("Grouping consecutive characters:")
for key, group in groupby("aabbccdd"):
    print(f"  {key}: {list(group)}")

# ✅ Expected output:
# Grouping consecutive characters:
#   a: ['a', 'a']
#   b: ['b', 'b']
#   c: ['c', 'c']
#   d: ['d', 'd']

2.5 islice — Slice Iterators Without Lists

from itertools import islice, count

# Take first 10 items from infinite iterator
print("First 10 from count(0):")
print(list(islice(count(0), 10)))

# ✅ Expected output:
# First 10 from count(0):
# [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]

2.6 chain — Combine Multiple Iterables

from itertools import chain

result = list(chain([1,2], [3,4], [5]))
print("chain([1,2], [3,4], [5]):", result)

# ✅ Expected output:
# chain([1,2], [3,4], [5]): [1, 2, 3, 4, 5]

2.7 tee — Duplicate Iterators

from itertools import tee

original = iter([1, 2, 3, 4, 5])
a, b = tee(original)

print("First copy:", list(a))
print("Second copy:", list(b))

# ✅ Expected output:
# First copy: [1, 2, 3, 4, 5]
# Second copy: [1, 2, 3, 4, 5]

3.1 lru_cache — Instant Caching

from functools import lru_cache
import time

@lru_cache(maxsize=100)
def fib(n):
    if n < 2:
        return n
    return fib(n-1) + fib(n-2)

start = time.time()
result = fib(35)
elapsed = time.time() - start

print(f"fib(35) = {result}")
print(f"Computed in {elapsed:.4f} seconds (with cache)")

3.2 partial — Pre-Fill Function Arguments

from functools import partial

def multiply(x, y):
    return x * y

times_10 = partial(multiply, 10)

print("times_10(5):", times_10(5))
print("times_10(7):", times_10(7))

# ✅ Expected output:
# times_10(5): 50
# times_10(7): 70

3.3 reduce — Functional Reduce Operations

from functools import reduce

result = reduce(lambda x, y: x + y, [1, 2, 3, 4, 5])
print("Sum with reduce:", result)

product = reduce(lambda x, y: x * y, [1, 2, 3, 4, 5])
print("Product with reduce:", product)

# ✅ Expected output:
# Sum with reduce: 15
# Product with reduce: 120

3.4 singledispatch — Generic Functions

from functools import singledispatch

@singledispatch
def process(obj):
    print("default:", obj)

@process.register(int)
def _(value):
    print("int:", value * 2)

@process.register(list)
def _(value):
    print("list:", len(value), "items")

process("hello")
process(42)
process([1, 2, 3])

# ✅ Expected output:
# default: hello
# int: 84
# list: 3 items

3.5 cached_property — Lazy Loaded Attributes

from functools import cached_property

class DataLoader:
    @cached_property
    def data(self):
        print("Loading data (expensive operation)...")
        return [1, 2, 3, 4, 5]

loader = DataLoader()
print("First access:", loader.data)
print("Second access:", loader.data)  # No "Loading" message

# ✅ Expected output:
# Loading data (expensive operation)...
# First access: [1, 2, 3, 4, 5]
# Second access: [1, 2, 3, 4, 5]

4. Combining All Three Modules Into Powerful Pipelines

Example: large-scale log analysis.

from collections import Counter
from itertools import islice
from functools import lru_cache

@lru_cache
def normalize(line):
    return line.lower().strip()

# Simulated log lines
log_lines = [
    "ERROR: Connection failed",
    "INFO: User logged in",
    "ERROR: Connection failed",
    "WARNING: Disk space low",
    "ERROR: Connection failed",
    "INFO: User logged out",
]

c = Counter()
for line in islice(log_lines, None):
    c[normalize(line)] += 1

print("Log analysis results:")
for msg, count in c.most_common(3):
    print(f"  {count}x: {msg}")

# ✅ Expected output:
# Log analysis results:
#   3x: error: connection failed
#   1x: info: user logged in
#   1x: warning: disk space low

Part 2: Advanced Patterns With Each Module

1.1 Subtracting / Combining Counters

from collections import Counter

c1 = Counter("banana")
c2 = Counter("bandana")

print("c1:", c1)
print("c2:", c2)
print("c1 + c2:", c1 + c2)
print("c1 - c2:", c1 - c2)

# ✅ Expected output:
# c1: Counter({'a': 3, 'n': 2, 'b': 1})
# c2: Counter({'a': 3, 'n': 2, 'b': 1, 'd': 1})
# c1 + c2: Counter({'a': 6, 'n': 4, 'b': 2, 'd': 1})
# c1 - c2: Counter()

2.1 Multi-Level defaultdict

from collections import defaultdict

nested = defaultdict(lambda: defaultdict(int))
nested["user1"]["views"] += 1
nested["user1"]["likes"] += 3
nested["user2"]["views"] += 5

print("User metrics:")
for user, metrics in nested.items():
    print(f"  {user}: {dict(metrics)}")

# ✅ Expected output:
# User metrics:
#   user1: {'views': 1, 'likes': 3}
#   user2: {'views': 5}

3.1 Efficient Sliding Window Statistics

from collections import deque

window = deque(maxlen=5)

def add_and_average(value):
    window.append(value)
    return sum(window) / len(window)

values = [10, 20, 30, 40, 50, 60, 70]
for v in values:
    avg = add_and_average(v)
    print(f"Added {v}, window: {list(window)}, avg: {avg:.1f}")

# ✅ Expected output:
# Added 10, window: [10], avg: 10.0
# Added 20, window: [10, 20], avg: 15.0
# Added 30, window: [10, 20, 30], avg: 20.0
# Added 40, window: [10, 20, 30, 40], avg: 25.0
# Added 50, window: [10, 20, 30, 40, 50], avg: 30.0
# Added 60, window: [20, 30, 40, 50, 60], avg: 40.0
# Added 70, window: [30, 40, 50, 60, 70], avg: 50.0

4. Advanced namedtuple Usage

from collections import namedtuple

# With default values
Point = namedtuple("Point", "x y", defaults=[0, 0])
p1 = Point()
p2 = Point(3, 4)

print(f"Default point: {p1}")
print(f"Custom point: {p2}")

# Replacing fields
p3 = p2._replace(x=10)
print(f"After replace: {p3}")

# From dictionary
data = {"x": 5, "y": 8}
p4 = Point(**data)
print(f"From dict: {p4}")

# ✅ Expected output:
# Default point: Point(x=0, y=0)
# Custom point: Point(x=3, y=4)
# After replace: Point(x=10, y=4)
# From dict: Point(x=5, y=8)

5.1 Batch Iteration (Chunking)

from itertools import islice

def chunks(iterable, size):
    iterator = iter(iterable)
    while chunk := list(islice(iterator, size)):
        yield chunk

data = range(1, 11)
print("Chunking 1-10 into groups of 3:")
for batch in chunks(data, 3):
    print(f"  Batch: {batch}")

# ✅ Expected output:
# Chunking 1-10 into groups of 3:
#   Batch: [1, 2, 3]
#   Batch: [4, 5, 6]
#   Batch: [7, 8, 9]
#   Batch: [10]

5.2 Combining Streams With zip_longest

from itertools import zip_longest

list1 = [1, 2, 3]
list2 = ['a', 'b', 'c', 'd', 'e']

merged = list(zip_longest(list1, list2, fillvalue=None))
print("Merged with zip_longest:")
for pair in merged:
    print(f"  {pair}")

# ✅ Expected output:
# Merged with zip_longest:
#   (1, 'a')
#   (2, 'b')
#   (3, 'c')
#   (None, 'd')
#   (None, 'e')

6.1 Combining partial + map

from functools import partial

def multiply(x, y):
    return x * y

times_10 = partial(multiply, 10)

result = list(map(times_10, range(1, 6)))
print("Multiplied by 10:", result)

# ✅ Expected output:
# Multiplied by 10: [10, 20, 30, 40, 50]

Part 3: High-Performance Data Pipelines

8. High-Performance Data Pipelines (Full Walkthrough)

Modern Python systems — data processing apps, ML pipelines, scrapers, log processors — rely on streaming, chunking, and lazy evaluation.

Complete Pipeline Example

from itertools import islice
from collections import Counter
from functools import lru_cache

# Stage 1: Streaming Reader (Generator)
def read_lines(lines):
    for line in lines:
        yield line

# Stage 2: Cleaning & Normalizing
def normalize(lines):
    for line in lines:
        line = line.strip().lower()
        if line:
            yield line

# Stage 3: Tokenizing
def tokenize(lines):
    for line in lines:
        for word in line.split():
            yield word

# Stage 4: Batching
def batches(iterable, size=3):
    it = iter(iterable)
    while chunk := list(islice(it, size)):
        yield chunk

# Stage 5: Aggregation
def count_words(batches):
    total = Counter()
    for batch in batches:
        total.update(batch)
    return total

# Sample data
sample_text = [
    "Python is amazing",
    "Python is powerful",
    "Data science with Python",
    "Machine learning is fun",
]

# Run pipeline
data = read_lines(sample_text)
clean = normalize(data)
words = tokenize(clean)
batch_stream = batches(words, 5)
counts = count_words(batch_stream)

print("Word frequencies:")
for word, count in counts.most_common(5):
    print(f"  {word}: {count}")

# ✅ Expected output:
# Word frequencies:
#   python: 3
#   is: 3
#   amazing: 1
#   powerful: 1
#   data: 1

🏁 Mini-Challenge: top three tags

No blanks this time — just a brief and an empty outline.

Each post below carries a comma-separated tag list, typed by humans, so the spacing and capitalisation are a mess. Produce the three most-used tags, tidied to lowercase. Aim for a lazy pipeline: chain.from_iterable to flatten, a generator expression to clean, and Counter to tally — no intermediate lists of your own.

# 🏁 MINI-CHALLENGE — write the pipeline yourself
from collections import Counter
from itertools import chain

posts = [
    "Python, Data ,ML",
    "python,web",
    " ML , python , data",
    "Python,data",
]

# 1. Flatten: turn the 4 strings into one stream of raw tags.
#    Hint: each post becomes a list with post.split(","), and
#    chain.from_iterable() glues those lists into a single stream.

# 2. Clean: strip the spaces off each tag and lowercase it,
#    so " ML " and "ML" and "ml" all count as the same tag.

# 3. Tally: hand the cleaned stream to Counter.

# 4. Print the top three, one per line, as "tag: count".

# your code here


# ✅ Expected output:
# python: 4
# data: 3
# ml: 2

🎓 Final Summary

In this comprehensive lesson, you learned expert-level usage of three critical Python modules:

📋 Quick Reference — Advanced Collections

ToolBest for
Counter(iterable)Count occurrences of elements
deque(maxlen=100)Fast append/pop from both ends
defaultdict(list)Dict with automatic default values
itertools.chain(*iterables)Combine multiple iterables
functools.lru_cacheCache function results by arguments

🎉 Great work! You've completed this lesson.

You can now use Counter, deque, defaultdict, itertools and functools to write more expressive and efficient Python.

Practice quiz

What is collections.Counter best suited for?

  • Sorting a list
  • Removing duplicates only
  • Counting occurrences of elements (frequencies)
  • Reversing a string

Answer: Counting occurrences of elements (frequencies). Counter counts occurrences — ideal for word frequencies, histograms, and event counting.

What does a defaultdict(list) give you when you access a missing key?

  • An automatically created empty list
  • A KeyError
  • None
  • An empty string

Answer: An automatically created empty list. defaultdict(list) creates a new empty list for any missing key, so you can append without checking.

What is a deque and what is its key advantage?

  • A sorted set with O(log n) lookup
  • A read-only tuple
  • A hash map with default values
  • A double-ended queue with O(1) append/pop on both ends

Answer: A double-ended queue with O(1) append/pop on both ends. deque is a double-ended queue offering O(1) appends and pops from both ends.

After dq = deque(maxlen=5) and appending 0..9 one at a time, what does list(dq) contain?

  • [0, 1, 2, 3, 4]
  • [5, 6, 7, 8, 9]
  • [0, 1, 2, ..., 9]
  • [9, 8, 7, 6, 5]

Answer: [5, 6, 7, 8, 9]. A maxlen deque discards items from the opposite end, keeping the last 5: [5,6,7,8,9].

What does namedtuple give you over a plain tuple?

  • Named fields with tuple speed and immutability
  • Mutability
  • Automatic sorting
  • Default dictionary behavior

Answer: Named fields with tuple speed and immutability. namedtuple adds readable field names while keeping tuple speed, immutability, and memory efficiency.

What does functools.lru_cache do to a function?

  • Runs it in parallel
  • Logs every call
  • Caches results keyed by arguments to avoid recomputation
  • Makes it asynchronous

Answer: Caches results keyed by arguments to avoid recomputation. lru_cache stores results by argument, returning the cached value on repeat calls.

What does functools.partial(multiply, 10) produce?

  • A copy of multiply
  • A new callable with the first argument pre-filled as 10
  • An error
  • The number 10

Answer: A new callable with the first argument pre-filled as 10. partial pre-fills arguments, so partial(multiply, 10)(5) calls multiply(10, 5).

What is list(itertools.accumulate([1, 2, 3, 4]))?

  • [1, 2, 3, 4]
  • [10, 6, 3, 1]
  • [24]
  • [1, 3, 6, 10]

Answer: [1, 3, 6, 10]. accumulate produces a running sum by default: 1, 1+2, 1+2+3, 1+2+3+4.

What does itertools.chain([1,2], [3,4], [5]) yield?

  • [[1,2],[3,4],[5]]
  • 1, 2, 3, 4, 5 as a single stream
  • [1, 3, 5]
  • Only the first iterable

Answer: 1, 2, 3, 4, 5 as a single stream. chain links multiple iterables into one continuous sequence: 1, 2, 3, 4, 5.

What does itertools.groupby group together?

  • All equal items anywhere in the iterable
  • Items by their hash value
  • Only consecutive equal items
  • Items into pairs

Answer: Only consecutive equal items. groupby groups only consecutive equal items, so unsorted input may produce repeated keys.

Continue this course