Evaluating AI Models

Reviewed & published by Brayan K

By the end of this lesson you'll read a confusion matrix, pick the right metric for any task, and tell a great model from one that's only pretending.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

📊 Real-World Analogy: A School Report Card

Imagine grading a student with a single number — say "78% of answers correct." It hides a lot. Did they ace maths but fail science? Did the test even cover the right topics? A real report card breaks the grade into subjects, comments, and a comparison to the class, because one number can mislead.

Evaluating an AI model is exactly the same. One score — accuracy — can look brilliant while the model is useless. A proper evaluation is a report card: several metrics, each measuring a different subject, plus a check that you graded on questions the model never saw before. This lesson is how you write that report card.

1 The Confusion Matrix — Where Every Metric Begins

When a model answers yes/no questions, every prediction lands in one of four buckets. A true positive is a correct "yes"; a true negative a correct "no". A false positive is a false alarm (said yes, was no), and a false negative is a miss (said no, was yes). Lay those four counts in a grid and you have the confusion matrix.

You don't need any library to build one — just walk the two lists together and tally. Run this:

# A confusion matrix sorts every prediction into one of four buckets.
# We'll build it by hand from two lists — no libraries needed.

# 1 = "positive" (e.g. the patient IS sick), 0 = "negative" (healthy)
y_true = [1, 0, 1, 1, 0, 0, 1, 0, 0, 1]   # the real answers
y_pred = [1, 0, 0, 1, 0, 1, 1, 0, 0, 1]   # what the model guessed

# Four counts make up the confusion matrix:
tp = 0   # True Positive  — predicted 1, really 1  (correct hit)
tn = 0   # True Negative  — predicted 0, really 0  (correct pass)
fp = 0   # False Positive — predicted 1, really 0  (false alarm)
fn = 0   # False Negative — predicted 0, really 1  (a miss)

for actual, guess in zip(y_true, y_pred):
    if   actual == 1 and guess == 1: tp += 1
    elif actual == 0 and guess == 0: tn += 1
    elif actual == 0 and guess == 1: fp += 1
    elif actual == 1 and guess == 0: fn += 1

print("Confusion Matrix")
print(f"  True Positives  (TP): {tp}")   # 4
print(f"  True Negatives  (TN): {tn}")   # 4
print(f"  False Positives (FP): {fp}")   # 1  (one false alarm)
print(f"  False Negatives (FN): {fn}")   # 1  (one missed case)

# Every other classification metric is built from these four numbers.

2 Classification Metrics — Accuracy, Precision, Recall, F1

Those four counts give you four very different views of quality. Accuracy is the fraction of all predictions that were right. Precision asks "when the model says yes, how often is it correct?" Recall asks "of everything it should have caught, how much did it?" And F1 blends precision and recall into one number using the harmonic mean, so a model can't cheat by being great at one and terrible at the other.

accuracy = (TP + TN) / everything

precision = TP / (TP + FP)

recall = TP / (TP + FN)

F1 = 2 · precision · recall / (precision + recall)

# From the four confusion-matrix counts you can derive every metric.
tp, tn, fp, fn = 4, 4, 1, 1

# Accuracy: of ALL predictions, how many were correct?
accuracy = (tp + tn) / (tp + tn + fp + fn)

# Precision: of everything FLAGGED positive, how many really were?
#   "When the model says yes, how often is it right?"
precision = tp / (tp + fp)

# Recall (sensitivity): of all REAL positives, how many did we catch?
#   "Of everything we should have caught, how much did we?"
recall = tp / (tp + fn)

# F1 score: the harmonic mean — one number balancing precision and recall.
f1 = 2 * precision * recall / (precision + recall)

print(f"Accuracy:  {accuracy:.2f}")    # 0.80  (8 of 10 correct)
print(f"Precision: {precision:.2f}")   # 0.80  (4 of 5 flagged were right)
print(f"Recall:    {recall:.2f}")      # 0.80  (caught 4 of 5 real cases)
print(f"F1 score:  {f1:.2f}")          # 0.80

# ROC-AUC (0.5 = random guessing, 1.0 = perfect) needs probabilities and
# a sweep of thresholds, so it's normally read from a library — see below.

🛠️ The Real Tool: scikit-learn

You hand-rolled the metrics above to understand them — but in real projects you call a library. Here is the exact same data scored by scikit-learn, including the ROC-AUC you can't easily do by hand. The # Expected output comment shows what you'd see if you ran it locally.

# In real projects you don't hand-roll these — scikit-learn does it for you.
# This is what the worked example above looks like with the library.
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score, f1_score, roc_auc_score
)

y_true = [1, 0, 1, 1, 0, 0, 1, 0, 0, 1]
y_pred = [1, 0, 0, 1, 0, 1, 1, 0, 0, 1]
# Predicted PROBABILITY of class 1 (ROC-AUC needs scores, not labels):
y_prob = [0.9, 0.1, 0.4, 0.8, 0.2, 0.6, 0.7, 0.3, 0.1, 0.95]

print("accuracy :", accuracy_score(y_true, y_pred))
print("precision:", precision_score(y_true, y_pred))
print("recall   :", recall_score(y_true, y_pred))
print("f1       :", f1_score(y_true, y_pred))
print("roc_auc  :", roc_auc_score(y_true, y_prob))

# Expected output:
# accuracy : 0.8
# precision: 0.8
# recall   : 0.8
# f1       : 0.8000000000000002
# roc_auc  : 0.9166666666666667

Notice the metrics match your hand calculations exactly. The library just saves you the arithmetic — and adds ROC-AUC, which needs the probability column y_prob.

# 🎯 YOUR TURN — compute precision and recall from confusion counts
# A spam filter was tested on 100 emails. Here are its four counts:

tp = 35   # spam correctly caught
fp = 5    # good emails wrongly flagged as spam (false alarms)
fn = 10   # spam that slipped through (misses)
tn = 50   # good emails correctly let through

# 1) Precision = TP / (TP + FP)   "when it flags spam, how often is it right?"
precision = ___          # 👉 replace ___ using tp and fp

# 2) Recall = TP / (TP + FN)      "of all the spam, how much did it catch?"
recall = ___             # 👉 replace ___ using tp and fn

print(f"Precision: {precision:.2f}")
print(f"Recall:    {recall:.2f}")

# ✅ Expected output:
# Precision: 0.88
# Recall:    0.78

3 Regression Metrics — MAE, MSE, RMSE, R²

Classification predicts a category; regression predicts a number — a price, a temperature, an age. So instead of counting hits and misses, you measure how far off the predictions are.

MAE (mean absolute error) is the average size of the error in the data's own units. MSE squares the errors first, so big misses count extra; RMSE takes the square root of MSE to get back into the original units. R² reports the fraction of the data's variation the model explains — 1.0 is perfect, 0 is no better than always guessing the average.

# Regression predicts a NUMBER (a price, a temperature), not a class.
# So we measure how far off we are, on average.

y_true = [3.0, -0.5, 2.0, 7.0, 4.5]   # the real values
y_pred = [2.5,  0.0, 2.0, 8.0, 4.0]   # the model's predictions
n = len(y_true)

# Mean Absolute Error: average size of the error (same units as the data).
mae = sum(abs(t - p) for t, p in zip(y_true, y_pred)) / n

# Mean Squared Error: average of the SQUARED errors (punishes big misses).
mse = sum((t - p) ** 2 for t, p in zip(y_true, y_pred)) / n

# Root Mean Squared Error: square root of MSE → back into the data's units.
rmse = mse ** 0.5

# R-squared: fraction of variance explained (1.0 = perfect, 0 = no better
# than always guessing the mean).
mean_true = sum(y_true) / n
ss_res = sum((t - p) ** 2 for t, p in zip(y_true, y_pred))   # error left over
ss_tot = sum((t - mean_true) ** 2 for t in y_true)           # total spread
r2 = 1 - ss_res / ss_tot

print(f"MAE : {mae:.3f}")     # 0.400
print(f"MSE : {mse:.3f}")     # 0.260
print(f"RMSE: {rmse:.3f}")    # 0.510
print(f"R^2 : {r2:.3f}")      # 0.954

# Rule of thumb: report MAE for an easy-to-explain "typical error",
# and RMSE when large mistakes should count extra.
# 🎯 YOUR TURN — compute RMSE from two lists
# A model predicted house prices (in £1000s). Compare to the real prices.

y_true = [200, 350, 500, 275]   # real prices
y_pred = [210, 330, 540, 260]   # predicted prices
n = len(y_true)

# 1) Mean Squared Error: average of the squared differences.
mse = ___ / n            # 👉 sum((t - p) ** 2 for t, p in zip(y_true, y_pred))

# 2) RMSE: the square root of MSE (use ** 0.5).
rmse = ___               # 👉 take the square root of mse

print(f"MSE : {mse:.1f}")
print(f"RMSE: {rmse:.2f}")

# ✅ Expected output:
# MSE : 525.0
# RMSE: 22.91

4 Cross-Validation — Don't Trust a Single Split

Score a model on one random test split and you might just get lucky — or unlucky. k-fold cross-validation fixes that: it chops the data into k equal parts (folds), trains on k−1 of them and tests on the one held out, then rotates so every row gets tested exactly once. You report the average of the k scores.

The spread of those scores matters too. A small standard deviation means the result is stable; a big one means your model's quality swings with the split — a warning sign.

# Cross-validation gives a more trustworthy score than a single split.
# k-fold splits the data into k parts: train on k-1, test on the held-out 1,
# and rotate so every row is tested exactly once. Here we fake the scores
# you'd get back from each fold to show how the averaging works.

fold_scores = [0.82, 0.79, 0.85, 0.81, 0.78]   # accuracy from each of 5 folds
k = len(fold_scores)

mean_score = sum(fold_scores) / k

# Standard deviation tells you how STABLE the model is across folds.
variance = sum((s - mean_score) ** 2 for s in fold_scores) / k
std_dev = variance ** 0.5

print(f"{k}-fold scores: {fold_scores}")
print(f"Mean accuracy : {mean_score:.3f}")        # 0.810
print(f"Std deviation : {std_dev:.3f}")           # 0.024
print(f"Report as     : {mean_score:.2f} +/- {std_dev:.2f}")  # 0.81 +/- 0.02

# A low std deviation means the score is reliable, not a lucky split.

⚖️ Overfitting vs Underfitting (Bias-Variance)

The whole reason you cross-validate is to catch two opposite failures. Underfitting (high bias) is a model too simple to learn the pattern — it scores poorly on both the training data and new data. Overfitting (high variance) is a model that memorised the training data, including its noise — it scores brilliantly on training but badly on anything new.

🎯 Mini-Challenge: Score a Classifier

Time to fly solo. You're given a confusion matrix in the comments. Compute all four classification metrics from scratch — only the outline is provided, no formulas filled in.

# 🎯 MINI-CHALLENGE: Score a binary classifier
# A model was tested on 200 cases. Its confusion matrix:
#   TP = 60,  FP = 20,  FN = 30,  TN = 90
#
# 1. Store the four counts in variables tp, fp, fn, tn
# 2. Compute accuracy  = (tp + tn) / (tp + tn + fp + fn)
# 3. Compute precision = tp / (tp + fp)
# 4. Compute recall    = tp / (tp + fn)
# 5. Compute f1        = 2 * precision * recall / (precision + recall)
# 6. Print each value rounded to 2 decimals
#
# ✅ Expected output:
# Accuracy:  0.75
# Precision: 0.75
# Recall:    0.67
# F1:        0.71

# your code here

5 Common Mistakes (And How to Fix Them)

These four traps catch almost every beginner. Spot them before they fool you.

❌ Trusting accuracy on imbalanced data

If 99% of cases are negative, a model that always says "negative" hits 99% accuracy and catches nothing:

# 990 healthy, 10 sick; model predicts "healthy" for everyone
accuracy = 990 / 1000   # 0.99  ← looks amazing
recall   = 0 / 10       # 0.00  ← caught zero sick patients!

✅ Fix: report precision, recall, F1 and ROC-AUC alongside accuracy.

❌ Evaluating on the training set

Scoring on the same data the model learned from rewards memorisation, not understanding:

model.fit(X_train, y_train)
score = model.score(X_train, y_train)   # ❌ inflated — it has seen this

✅ Fix: always score on a held-out test set the model has never seen.

score = model.score(X_test, y_test)     # ✅ honest

❌ Reporting one lucky split, with no cross-validation

A single train/test split can hand you a flattering (or unfair) number by chance.

✅ Fix: use k-fold cross-validation and report the mean ± standard deviation, not one number.

❌ Optimising the wrong metric

Maximising precision on a cancer screen means you'll miss sick patients (low recall) to avoid false alarms — the opposite of what matters here.

✅ Fix: choose the metric from the cost of errors. Disease detection → recall. Spam filter → precision. Unsure → F1.

📋 Quick Reference

MetricTaskFormula / MeaningUse When
AccuracyClassification(TP+TN)/allClasses are balanced
PrecisionClassificationTP/(TP+FP)False positives costly
RecallClassificationTP/(TP+FN)False negatives costly
F1ClassificationHarmonic mean of P & RBoth matter equally
ROC-AUCBinary classif.0.5 random, 1.0 perfectThreshold-independent
MAERegressionavg |y−ŷ|Easy-to-explain error
RMSERegression√(avg squared error)Big misses count extra
R²Regressionvariance explained1.0 perfect, 0 = mean

❓ Frequently Asked Questions

Lesson complete — you can now write a model's report card!

You can build a confusion matrix by hand, derive accuracy, precision, recall and F1, measure regression error with MAE/RMSE/R², cross-validate for a trustworthy score, and read the bias-variance gap to spot overfitting. Most importantly, you know that the metric you choose shapes the model you build.

Practice quiz

What is a false positive in a confusion matrix?

  • Predicted negative when it was really positive
  • Predicted positive and it was positive
  • Predicted positive when it was really negative
  • Predicted negative and it was negative

Answer: Predicted positive when it was really negative. A false positive is a false alarm: the model said positive but the true label was negative.

How is precision calculated?

  • TP / (TP + FP)
  • TP / (TP + FN)
  • (TP + TN) / total
  • TN / (TN + FP)

Answer: TP / (TP + FP). Precision = TP / (TP + FP): of everything flagged positive, how many really were positive.

How is recall (sensitivity) calculated?

  • TP / (TP + FP)
  • TN / (TN + FN)
  • (TP + TN) / total
  • TP / (TP + FN)

Answer: TP / (TP + FN). Recall = TP / (TP + FN): of all real positives, how many the model caught.

What is the F1 score?

  • The arithmetic mean of precision and recall
  • The harmonic mean of precision and recall
  • The same as accuracy
  • The square root of precision times accuracy

Answer: The harmonic mean of precision and recall. F1 = 2 * P * R / (P + R), the harmonic mean, so a model cannot win by being great at only one.

Why is accuracy misleading on imbalanced data?

  • A model that always predicts the majority class can score high while catching no positives
  • It is always too low
  • It cannot be computed
  • It only works for regression

Answer: A model that always predicts the majority class can score high while catching no positives. With 99% negatives, always predicting negative gives 99% accuracy but zero recall on the rare class.

What does an ROC-AUC of 0.5 indicate?

  • A perfect classifier
  • Severe overfitting
  • Random guessing
  • Perfect recall

Answer: Random guessing. ROC-AUC ranges from 0.5 (random) to 1.0 (perfect ranking of positives above negatives).

Which regression metric is in the same units as the data and penalises large errors more than MAE?

  • R-squared
  • RMSE
  • Accuracy
  • Precision

Answer: RMSE. RMSE is the square root of MSE, so it is in the data's units and squares (penalises) big misses.

What does an R-squared of 1.0 mean?

  • The model explains none of the variance
  • The model is overfit
  • The errors are infinite
  • The model explains all of the variance (perfect fit)

Answer: The model explains all of the variance (perfect fit). R-squared is the fraction of variance explained; 1.0 is perfect and 0 is no better than predicting the mean.

What does k-fold cross-validation do?

  • Trains on the test set
  • Splits data into k folds, rotating which fold is held out so every row is tested once
  • Removes outliers
  • Increases the dataset size

Answer: Splits data into k folds, rotating which fold is held out so every row is tested once. k-fold rotates the held-out fold so every row is tested exactly once, giving a more reliable averaged score.

A model scores 99% on training but 70% on the test set. What is happening?

  • Underfitting
  • Perfect generalisation
  • Overfitting (high variance)
  • Data leakage from the test set

Answer: Overfitting (high variance). A large train-test gap is the classic sign of overfitting: the model memorised the training data.

Continue this course

Frequently asked questions

What is a confusion matrix?

A 2x2 table that sorts predictions into four buckets: true positives, true negatives, false positives, and false negatives. Every classification metric — accuracy, precision, recall, F1 — is calculated from these four counts.

What is the difference between precision and recall?

Precision asks 'when the model flags something positive, how often is it right?' (TP / (TP + FP)). Recall asks 'of all the real positives, how many did we catch?' (TP / (TP + FN)). Optimise precision when false alarms are costly, recall when missing a real case is costly.

Why is accuracy misleading on imbalanced data?

If 99% of cases are negative, a model that always predicts 'negative' scores 99% accuracy while catching zero positives. Accuracy hides this failure; precision, recall, F1, and ROC-AUC expose it.

When should I use RMSE versus MAE?

Both measure regression error in the data's own units. MAE is the average absolute error and is easy to explain. RMSE squares the errors first, so it penalises large mistakes more heavily — use it when big misses are especially bad.

What does k-fold cross-validation do?

It splits your data into k parts, trains on k-1 and tests on the held-out part, then rotates so every row is tested once. Averaging the k scores gives a more reliable estimate than a single train/test split, and the spread tells you how stable the model is.

What is the difference between overfitting and underfitting?

Underfitting (high bias) means the model is too simple and does poorly on both training and test data. Overfitting (high variance) means it memorised the training data and does well there but poorly on new data. The gap between training and test scores is the giveaway.