Text Analyzer

Analyze words, sentences, and keyword frequency using string methods, collections, and decorators.

Part of the free Python course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

Project Overview

A text analyzer takes a block of writing — an essay, a chapter, a support-ticket export, a folder of interview transcripts — and turns it into numbers you can act on: how long it is, how varied the vocabulary is, and which words carry the most weight. In this project you build one as a small command-line program driven by a menu: load a file, analyze it, run a find-and-replace across the whole text, and save the result back to disk.

When it works, you run python text_analyzer.py, press 1 and point it at a .txt file, then press 2. It prints a character count, a word count, how many of those words were unique, a sentence count, and the ten most frequent words with their tallies — followed by how long the whole pass took, because the analysis method is wrapped in a timing decorator you write yourself.

This is not a toy exercise. The same counting pass sits underneath tools people use every day: the word-count badge in a writing app, the keyword-density panel in an SEO checker, the "5 min read" estimate on a blog post, the vocabulary reports a teacher runs over student essays, and the first cleaning stage of nearly every natural-language pipeline. Writers, students, editors, and anyone triaging a pile of free-text feedback are the people who would actually reach for it.

Goal: Analyze text files and generate statistics.

Concepts: String methods, collections.Counter, decorators for timing.

Key Features:

Be honest with yourself about which of those the starter code already does. Counts, unique words, top words, and find-and-replace are finished and working. Reading level is not implemented at all, and the method called save_report currently writes the text back out rather than the statistics — the numbers are printed to the screen and then thrown away. Those two gaps are the first real work you own, and the Enhancement Ideas section shows where to start.

Core Concepts

Five ideas hold this project up. If any of them is fuzzy, the code will look like magic instead of like something you could have written, so read these before the code rather than after it.

1. Tokenizing: turning a string into a list of words

A computer does not see words, it sees characters, so before you can count anything you have to decide where one word stops and the next begins. The obvious answer, text.split(), splits on whitespace only, which means "language." and "language" become two different words and every comma stays glued to whatever it touches. This project uses the regular expression \b\w+\b instead: \w matches letters, digits and underscores, \b marks a boundary between word and non-word characters, so punctuation is left behind and only the word survives.

2. Counting: collections.Counter

You could count word frequencies with a plain dictionary and an if-statement for the first time you see each word. Counter is a dictionary subclass that does exactly that for you: hand it any iterable and it returns a mapping of item to how many times it appeared, with missing keys reporting zero instead of raising a KeyError. Its most_common(n) method hands back the top n as a list of (word, count) pairs already sorted from most to least frequent, which is the entire "top ten words" feature in one line.

3. Decorators: a function that wraps a function

In Python, functions are ordinary values you can pass around and return. A decorator takes advantage of that: it accepts a function, defines a new inner function that does something extra before and after calling the original, and returns that inner function in its place. Writing @timer above a method is just shorthand for reassigning the method to timer(method). That is why the timing code appears once and measures every method you tag with it, instead of being copy-pasted into each one.

4. File I/O and the encoding argument

Reading a file uses open() inside a with block, so the file is closed for you even if the code inside raises an error. The encoding="utf-8" argument matters more than beginners expect: without it, Python uses whatever your operating system defaults to, so the same program can read a file happily on one machine and crash on another. Text analysis in particular runs into curly quotes, accented names and emoji constantly, so state the encoding every time.

5. A class as a place to keep state

Every feature here needs the same thing: the text. Passing that string into five separate functions would mean threading it through every call and re-tokenizing it constantly. A class stores it once in self.text, and each method reads it from there — so loading a file changes what every later method sees, and a find-and-replace is a genuine edit rather than a value you have to remember to pass along.

Starter Code

This is a command-line app that uses input() for user interaction. The browser demo below shows a simplified version. For the full interactive experience, copy the code below and run it in your local Python environment (IDLE, VS Code, or terminal). Jump to the Browser Demo.

Here's a functional starting version of your app. Copy it to your local Python editor and run it.

Read it top to bottom before you run it and you will see three layers. The timer decorator sits alone at the top because it knows nothing about text — it could time any function, and keeping it separate is what makes it reusable. The TextAnalyzer class holds the text and every operation on it, so the data and the code that works on it live in one place. The block under if __name__ == "__main__": is the user interface and nothing else: it reads a menu choice and calls a method. That separation is deliberate — the class never calls input() or print() for menus, which is what will let you reuse it later behind a web form or a test suite.

import re
from collections import Counter
import time

# Decorator to measure execution time
def timer(func):
    def wrapper(*args, **kwargs):
        start = time.time()
        result = func(*args, **kwargs)
        print(f"\n⏱ Execution time: {time.time() - start:.3f}s")
        return result
    return wrapper

class TextAnalyzer:
    def __init__(self, text=""):
        self.text = text

    def load_file(self, filename):
        try:
            with open(filename, "r", encoding="utf-8") as f:
                self.text = f.read()
            print(f"📂 Loaded file: {filename}")
        except FileNotFoundError:
            print("❌ File not found.")

    @timer
    def analyze(self):
        words = re.findall(r"\b\w+\b", self.text.lower())
        sentences = re.split(r"[.!?]+", self.text)
        char_count = len(self.text)
        word_count = len(words)
        sentence_count = len([s for s in sentences if s.strip()])
        unique_words = len(set(words))
        freq = Counter(words).most_common(10)

        print(f"\n📊 Text Statistics:")
        print(f"Characters: {char_count}")
        print(f"Words: {word_count}")
        print(f"Unique Words: {unique_words}")
        print(f"Sentences: {sentence_count}")
        print("\n🔠 Most Common Words:")
        for w, c in freq:
            print(f"{w}: {c}")

    def find_replace(self, old, new):
        self.text = self.text.replace(old, new)
        print(f"✅ Replaced '{old}' with '{new}'")

    def save_report(self, filename="report.txt"):
        with open(filename, "w", encoding="utf-8") as f:
            f.write(self.text)
        print(f"💾 Report saved to {filename}")

# Demo Menu
if __name__ == "__main__":
    analyzer = TextAnalyzer()
    while True:
        print("\n--- TEXT ANALYZER ---")
        print("1. Load Text File")
        print("2. Analyze Text")
        print("3. Find & Replace")
        print("4. Save Report")
        print("5. Exit")
        choice = input("Choose option: ")

        if choice == "1":
            filename = input("Enter filename: ")
            analyzer.load_file(filename)
        elif choice == "2":
            analyzer.analyze()
        elif choice == "3":
            old = input("Find: ")
            new = input("Replace with: ")
            analyzer.find_replace(old, new)
        elif choice == "4":
            analyzer.save_report()
        elif choice == "5":
            print("👋 Exiting...")
            break
        else:
            print("❌ Invalid choice.")

The mistake to avoid with this file is running option 2 first. A fresh TextAnalyzer() starts with text="", so analyzing before loading anything prints a wall of zeros and an empty word list rather than an error — which looks like a broken analyzer when it is really an empty one. The same trap bites harder after a failed load: when load_file catches a missing filename it prints a message but leaves self.text untouched, so the next analysis silently reports on whatever was loaded before. Watch for the "Loaded file" line before you trust any numbers.

🚀 How to Run Locally:

Walking Through the Code

Here is the same program again, one piece at a time, with the reasoning that is invisible when you only see the finished file. Each piece below is lifted straight out of the starter code above — nothing new, just slowed down.

Step 1: Write the timing decorator

The decorator is defined first because Python has to know what timer means before it reaches the @timer line inside the class. It takes the function you are decorating, builds a wrapper around it that records a start time, calls the real function, prints how long that took, and — this is the line people forget — returns the original result so the caller still gets its value. The *args, **kwargs signature is what lets one decorator wrap methods that take no arguments and methods that take several.

def timer(func):
    def wrapper(*args, **kwargs):
        start = time.time()
        result = func(*args, **kwargs)
        print(f"\n⏱ Execution time: {time.time() - start:.3f}s")
        return result
    return wrapper

The classic mistake is writing func(*args, **kwargs) without assigning it to result and returning it. The decorated method then quietly returns None, and every caller that expected a number breaks somewhere far away from the decorator. A second, subtler cost: the wrapper replaces the original function, so the method's own name and docstring are lost — ask a decorated method for its __name__ and it answers "wrapper". Wrapping the inner function with functools.wraps(func) copies that identity across and is worth adding once you have this working.

Step 2: Hold the text in a class and load a file into it

The constructor defaults to an empty string so you can create an analyzer before you have any text, which is exactly what the menu needs. load_file replaces that string with the file's contents. The with block guarantees the file handle is closed, and the try/except FileNotFoundError turns a typo in a filename into a one-line message instead of a traceback that ends the program mid-menu.

class TextAnalyzer:
    def __init__(self, text=""):
        self.text = text

    def load_file(self, filename):
        try:
            with open(filename, "r", encoding="utf-8") as f:
                self.text = f.read()
            print(f"📂 Loaded file: {filename}")
        except FileNotFoundError:
            print("❌ File not found.")

Use f.read(), which returns one string, not f.readlines(), which returns a list. Beginners swap them constantly and the failure is confusing, because nothing goes wrong until the next method tries to call a string method on a list. Note too that catching only FileNotFoundError is a deliberate choice: a directory passed instead of a file, or a file saved in a different encoding, still raises — and you want to see those rather than silently analyze nothing.

Step 3: Do the counting in one pass

This is the heart of the project. re.findall returns every word as a list, and lowercasing first is what makes "Python" and "python" count as the same word. len(words) is the word count, len(set(words)) is the unique-word count because a set discards duplicates, and Counter(words).most_common(10) is the frequency table. Sentences come from splitting on runs of ., ! and ?; the comprehension that filters on s.strip() exists because splitting on a trailing period leaves an empty string at the end that would otherwise be counted as a sentence.

    @timer
    def analyze(self):
        words = re.findall(r"\b\w+\b", self.text.lower())
        sentences = re.split(r"[.!?]+", self.text)
        char_count = len(self.text)
        word_count = len(words)
        sentence_count = len([s for s in sentences if s.strip()])
        unique_words = len(set(words))
        freq = Counter(words).most_common(10)

Know the limits of this tokenizer before you trust a report built on it. \b\w+\b counts "don't" as the two words "don" and "t", and "e-mail" as "e" and "mail", because apostrophes and hyphens are not word characters. Sentence splitting is worse: "Mr. Smith left." is counted as two sentences, since the splitter cannot tell an abbreviation from a full stop. Neither is a bug you must fix today, but quoting an "average words per sentence" figure without knowing this is how people publish numbers they cannot defend.

Step 4: Find and replace across the whole text

Python strings are immutable, so str.replace cannot edit the text in place — it returns a brand-new string. Assigning that result back to self.text is what actually makes the change stick, and it makes every later analysis and save see the edited version.

    def find_replace(self, old, new):
        self.text = self.text.replace(old, new)
        print(f"✅ Replaced '{old}' with '{new}'")

Two things bite here. First, dropping the assignment — writing just self.text.replace(old, new) — computes the new string, throws it away, and still prints the cheerful success message, which is the most misleading kind of bug. Second, replace is case-sensitive and matches anywhere inside a word: replacing "python" in "Python and python" changes only the lowercase one, and replacing "cat" also rewrites the middle of "concatenate". If you need whole words or case-insensitive matching, that is a job for re.sub with a \b boundary, not for str.replace.

Step 5: Save the result, then drive it all from a menu

Saving mirrors loading: same with block, same explicit encoding, but "w" mode instead of "r". The menu is an infinite while True loop that prints options, reads a choice, dispatches to a method, and loops again — with break as the only way out. Comparing against the string "1" rather than the number 1 matters, because input() always hands you a string.

    def save_report(self, filename="report.txt"):
        with open(filename, "w", encoding="utf-8") as f:
            f.write(self.text)
        print(f"💾 Report saved to {filename}")

Mode "w" truncates the file the instant it is opened, so a second save silently destroys the first — use "a" if you meant to append. And notice the honest gap already mentioned above: this method writes self.text, not the statistics, so "report.txt" is really just a copy of your input. Making analyze build and return a string of results, rather than printing them and discarding them, is the single most useful change you can make to this program.

🎮 Browser Demo (Simplified Version)

This browser version demonstrates the core logic without user input. Click Run to see it work!

Two things change from the starter code, and both are worth noticing. The menu is gone, because a browser has no terminal for input() to read from — the text is passed straight into the constructor instead. And the one analysis method has been split into five small ones, each tagged with @timer, which is a much better demonstration of what a decorator buys you: you write the timing logic once and get a measurement for every method for free. This version also tokenizes in __init__ and stores the result in self.words, so the five methods share one pass over the text rather than each re-running the regular expression.

from collections import Counter
import re
import time

def timer(func):
    """Decorator to measure function execution time"""
    def wrapper(*args, **kwargs):
        start = time.time()
        result = func(*args, **kwargs)
        elapsed = time.time() - start
        print(f"\n⏱️  {func.__name__} took {elapsed:.4f}s")
        return result
    return wrapper

class TextAnalyzer:
    def __init__(self, text):
        self.text = text
        self.words = re.findall(r'\b\w+\b', text.lower())
    
    @timer
    def word_count(self):
        return len(self.words)
    
    @timer
    def char_count(self):
        return len(self.text)
    
    @timer
    def sentence_count(self):
        sentences = re.split(r'[.!?]+', self.text)
        return len([s for s in sentences if s.strip()])
    
    @timer
    def most_common_words(self, n=10):
        return Counter(self.words).most_common(n)
    
    @timer
    def avg_word_length(self):
        return sum(len(word) for word in self.words) / len(self.words)
    
    def generate_report(self):
        print("\n" + "="*50)
        print("📊 TEXT ANALYSIS REPORT")
        print("="*50)
        
        print(f"\n📝 Basic Statistics:")
        print(f"   Characters: {self.char_count()}")
        print(f"   Words: {self.word_count()}")
        print(f"   Sentences: {self.sentence_count()}")
        print(f"   Avg word length: {self.avg_word_length():.1f} chars")
        
        print(f"\n🔝 Top 10 Most Common Words:")
        for word, count in self.most_common_words(10):
            print(f"   {word}: {count}")

# Demo: Analyze sample text
print("🧠 Text Analyzer Demo\n")

sample_text = """
Python is an amazing programming language. Python is used for 
web development, data science, artificial intelligence, and more.
Learning Python opens many career opportunities in tech.
Python makes coding fun and accessible to everyone who wants to learn.
The Python community is supportive and there are countless resources 
available for beginners and advanced programmers alike.
"""

print("📄 Sample Text:")
print(sample_text[:100] + "...")

analyzer = TextAnalyzer(sample_text)
analyzer.generate_report()

Run it and read the order of the lines carefully — it is the most instructive thing on this page. The timing message for each method prints above the label it belongs to, not below it. That is not a bug: in print(f" Characters: {self.char_count()}") Python must evaluate self.char_count() before it can build the string, and that call is the wrapped one, so the decorator's own print happens first. Once you have seen it here you will recognise the same evaluation order everywhere else in Python.

What the Demo Prints

This is the actual output of the demo above, captured from a real run. Your timing figures will differ — they measure your machine on that particular second — but every count should match exactly, because the sample text is fixed.

🧠 Text Analyzer Demo

📄 Sample Text:

Python is an amazing programming language. Python is used for
web development, data science, artif...

==================================================
📊 TEXT ANALYSIS REPORT
==================================================

📝 Basic Statistics:

⏱️  char_count took 0.0000s
   Characters: 384

⏱️  word_count took 0.0000s
   Words: 55

⏱️  sentence_count took 0.0001s
   Sentences: 5

⏱️  avg_word_length took 0.0000s
   Avg word length: 5.8 chars

🔝 Top 10 Most Common Words:

⏱️  most_common_words took 0.0004s
   python: 5
   and: 4
   is: 3
   for: 2
   to: 2
   an: 1
   amazing: 1
   programming: 1
   language: 1
   used: 1

Read the numbers, not just the shape. 384 characters against 55 words is the giveaway that char_count measures the raw string — spaces, newlines and the blank line at the top of the triple-quoted sample all count. The sample text ends in a period, and the sentence count is 5 rather than 6, which is the empty-string filter in sentence_count doing its job. And the tail of the frequency table is a row of words that appeared once each: when counts tie, most_common keeps the order in which it first met them, so "an" before "amazing" is first-appearance order, not alphabetical.

Common Errors

These are the messages this project actually produces when it goes wrong, with what causes each. Learning to read the last line of a traceback first — the error type and its message — will save you more time than any other habit.

ZeroDivisionError: division by zero

Thrown by avg_word_length, which divides by len(self.words). If the text is empty — or contains only punctuation, so the regular expression matches nothing — that length is zero. Guard it by returning 0 when self.words is empty, before you divide. Any average you compute over user-supplied text needs this check.

TypeError: TextAnalyzer.__init__() missing 1 required positional argument: 'text'

You wrote TextAnalyzer() while using the browser demo's class, whose constructor requires the text up front. The starter code's version gives text a default of "", so it accepts an empty call. Mixing the two versions is the usual cause — check which class definition is actually in your file.

AttributeError: 'list' object has no attribute 'lower'

You used f.readlines() instead of f.read(). The first returns a list of lines, the second returns one string, and every method here expects a string. If you genuinely want the lines, join them back together with "".join(lines) before storing them in self.text.

TypeError: expected string or bytes-like object, got 'list'

The same mistake, one line earlier: re.findall was handed a list rather than a string. Whenever a traceback names a type you did not expect, print type(self.text) right before the failing line — that one line of debugging answers the question immediately.

UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 3

The file you loaded is not UTF-8 — it is very often a document exported from an older Windows editor in a legacy encoding, where an accented character is a single byte that UTF-8 cannot interpret. Try opening it with encoding="latin-1", or pass errors="replace" to accept the damage and keep going. Never respond by deleting the encoding argument; that only moves the crash to a different machine.

FileNotFoundError, or the printed "❌ File not found."

The starter code catches this one and prints its own message, which is friendlier but easy to skim past. Nine times in ten the file exists and the path is relative to the wrong directory: Python looks for it in the folder you ran the command from, not the folder the script lives in. Pass the full path while you are testing, and remember that self.text keeps its old value when the load fails.

No error at all — every count is 0

The worst kind of failure, because nothing complains. You analyzed before loading, or the load failed and you missed the message. It is also what you get if you decorate a method with @timer and forget to return the result inside the wrapper: the method still runs and still prints its timing, but hands back None. Silence is not success — check that the numbers are plausible for the text you fed in.

Enhancement Ideas 🚀

Do these in order; each one builds on the last, and all of them assume you have first closed the report gap described in Step 5.

1. File Upload Support

Allow users to upload text files and automatically analyze them. On the command line the useful version of this is batch mode: use the glob module to collect every .txt file in a folder, loop over them, and print one row of statistics per file. That turns a single-document tool into something you can point at a whole directory of drafts, and it forces you to handle the case where one file in the batch fails to decode without killing the rest of the run.

2. Sentiment Analysis

Use the textblob library to determine if text is positive, negative, or neutral. It gives you a polarity score on a scale from negative to positive, which you can bucket into three labels and print alongside your other statistics. Treat the number as a rough signal rather than a verdict: sentiment tools of this kind are trained on general text and miss sarcasm, negation and domain jargon routinely, so it belongs in a report next to the raw counts, not instead of them.

3. Reading Level Calculator

Implement the Flesch-Kincaid readability test to calculate the reading level. You already have two of the three inputs — word count and sentence count — so the only new work is counting syllables, which a simple heuristic (count vowel groups per word, subtract a silent trailing "e") approximates well enough for a first version. The Flesch Reading Ease score is 206.835 minus 1.015 times the average words per sentence minus 84.6 times the average syllables per word; higher scores mean easier text. This is the enhancement that makes the sentence-splitting weakness from Step 3 matter, because abbreviations inflate your sentence count and drag the score in the wrong direction.

4. Word Cloud Generation

Use matplotlib and wordcloud to create visual word clouds. Both are third-party packages you install with pip, so this is also your first taste of managing dependencies in a virtual environment. Do one piece of preparation first: strip out stop words — "the", "and", "is", "of" — or your cloud will be a picture of English grammar rather than of your document. That same stop-word list will immediately improve your top-ten table too, which is a good hint that it belongs in the analyzer rather than in the plotting code.

5. Compare Multiple Texts

Add functionality to compare two texts and find similarities/differences. The clean way to do it is with the tools you already have: build a set of words for each text, then set_a & set_b gives the shared vocabulary, set_a - set_b gives what is unique to the first, and the size of the intersection divided by the size of the union is a similarity score between 0 and 1. Counter objects support the same operators while keeping the counts, so counter_a & counter_b gives you shared words with the smaller of the two tallies — enough to build a plagiarism-style overlap check or an author-style comparison.

Next Steps

Work down this list in order. Each step is small enough to finish in a sitting, and each one leaves you with a program that still runs.

Finish even half of this and you will have written something you can explain line by line in an interview: a class that owns its data, a decorator you wrote rather than imported, a regular expression whose limits you can name, and error handling that exists because you hit the errors yourself.

Related lessons