Model Monitoring in Production
Reviewed & published by Brayan K
Shipping a model is the start, not the finish. By the end of this lesson you'll detect data drift, track prediction quality over time, and wire up alerts that tell you to retrain — before users feel the damage.
Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
What You'll Learn in This Lesson
- Tell data drift apart from concept drift, with real examples
- Detect drift by comparing the mean and std of two batches
- Track prediction quality with a rolling-accuracy metric
- Trigger an alert when accuracy dips below a threshold
- Recognise label lag and why input drift warns you sooner
- Know when drift should trigger a model retrain
1 Why Models Degrade
A model that was 95% accurate on launch day can silently slide to 70% within weeks. Nothing in your code changed — the world changed, and your model didn't. There are two failures you must watch for.
Data drift — the inputs change shape. You trained on summer shoppers; now it's winter and the incoming feature values look different. The model still works correctly, it just hasn't seen this kind of input before.
Concept drift — the relationship between inputs and the right answer changes. A "good salary" meant 50k in 2010 and 80k today, so the same input should now produce a different label. The model is now answering the wrong question.
The trap: most teams only watch accuracy. But accuracy needs labels (the true answer), and labels often arrive weeks later — a problem called label lag. Input drift, by contrast, is visible the instant a request arrives. That's why you monitor both: drift warns you early, accuracy confirms the damage.
2 Detecting Drift by Comparing Two Batches
The simplest, label-free drift check is this: keep a reference batch (a sample of your training data) and, for each new current batch of live traffic, compare their summary statistics. If the mean (average) or standard deviation (how spread out the values are) move beyond a tolerance you set, the inputs have drifted.
Read the worked example below line by line — every function is plain Python, no libraries. Then run it and confirm the output matches the # Expected output comment at the bottom.
# ============================================
# DATA DRIFT vs CONCEPT DRIFT (plain Python)
# ============================================
# Data drift = the INPUTS change shape over time.
# (You trained on summer shoppers; now it is winter.)
# Concept drift = the INPUT -> OUTPUT relationship changes.
# ("good salary" meant 50k in 2010, 80k today.)
#
# This example detects DATA drift the simplest way there is:
# compare the mean and standard deviation of two batches.
def mean(values):
return sum(values) / len(values)
def std(values):
m = mean(values)
variance = sum((v - m) ** 2 for v in values) / len(values)
return variance ** 0.5
def check_drift(reference, current, mean_tol=2.0, std_tol=2.0):
"""Flag drift when mean or std move more than the tolerance."""
ref_mean, ref_std = mean(reference), std(reference)
cur_mean, cur_std = mean(current), std(current)
mean_shift = abs(cur_mean - ref_mean)
std_shift = abs(cur_std - ref_std)
drifted = mean_shift > mean_tol or std_shift > std_tol
return ref_mean, cur_mean, mean_shift, std_shift, drifted
# Reference batch = the data the model was trained on
reference = [50, 51, 49, 52, 48, 50, 51, 49, 50, 50]
# Batch A: looks just like training data -> no drift
batch_a = [50, 49, 51, 50, 52, 48, 50, 51, 49, 50]
# Batch B: shifted upward -> drift!
batch_b = [60, 62, 58, 64, 61, 59, 63, 60, 62, 61]
for label, batch in [("Batch A", batch_a), ("Batch B", batch_b)]:
ref_m, cur_m, m_shift, s_shift, drifted = check_drift(reference, batch)
flag = "DRIFT DETECTED" if drifted else "OK"
print(f"{label}: ref_mean={ref_m:.1f} cur_mean={cur_m:.1f} "
f"mean_shift={m_shift:.1f} std_shift={s_shift:.1f} -> {flag}")
# Expected output:
# Batch A: ref_mean=50.0 cur_mean=50.0 mean_shift=0.0 std_shift=0.1 -> OK
# Batch B: ref_mean=50.0 cur_mean=61.0 mean_shift=11.0 std_shift=0.4 -> DRIFT DETECTED🎯 Your Turn: Flag the Drift
# 🎯 YOUR TURN — detect drift by comparing two batches
# Fill in each ___ , then run it.
def mean(values):
return sum(values) / len(values)
reference = [100, 102, 98, 101, 99, 100, 101, 99]
current = [130, 133, 128, 131, 129, 132, 130, 131]
ref_mean = mean(reference)
cur_mean = mean(___) # 👉 pass the CURRENT batch here
mean_shift = abs(cur_mean - ref_mean) # 👉 how far the average moved
THRESHOLD = 10
drifted = mean_shift ___ THRESHOLD # 👉 use a comparison: > or <
print(f"ref_mean={ref_mean:.1f} cur_mean={cur_mean:.1f} shift={mean_shift:.1f}")
print("DRIFT DETECTED" if drifted else "OK")
# ✅ Expected output:
# ref_mean=100.0 cur_mean=130.5 shift=30.5
# DRIFT DETECTED3 Tracking Prediction Quality Over Time
Once labels do arrive, you can measure how well predictions are landing. But a single day's accuracy is noisy — one weird batch can make it jump or dip for reasons that don't matter. The fix is a rolling window: average the last few days so a genuine downward trend stands out from the daily wobble.
You then set an alert threshold — a line in the sand. When the rolling value crosses below it, you raise an alert (log it, post to Slack, or page on-call). The next example builds exactly that: a rolling accuracy that trips an alarm below 0.85.
# ============================================
# ROLLING ACCURACY + THRESHOLD ALERT (plain Python)
# ============================================
# Accuracy rarely falls off a cliff. It sags slowly as the world
# drifts away from your training data. A ROLLING window smooths out
# the day-to-day noise so a real downward trend is visible.
def rolling_accuracy(daily_accuracy, window=3):
"""Average of the last 'window' days, computed day by day."""
rolled = []
for i in range(len(daily_accuracy)):
start = max(0, i - window + 1)
chunk = daily_accuracy[start:i + 1]
rolled.append(sum(chunk) / len(chunk))
return rolled
# One accuracy reading per day for two weeks (0.0 - 1.0)
daily = [0.92, 0.91, 0.90, 0.91, 0.89, 0.88,
0.86, 0.84, 0.83, 0.81, 0.79, 0.78, 0.76, 0.74]
ALERT_THRESHOLD = 0.85 # page the on-call team below this
rolled = rolling_accuracy(daily, window=3)
print("Day raw rolling status")
alerted = False
for day, (raw, roll) in enumerate(zip(daily, rolled), start=1):
if roll < ALERT_THRESHOLD:
status = "ALERT - retrain!"
alerted = True
else:
status = "ok"
print(f"{day:>3} {raw:.2f} {roll:.2f} {status}")
print()
print("Alert raised" if alerted else "All clear")
# Expected output:
# Day raw rolling status
# 1 0.92 0.92 ok
# 2 0.91 0.92 ok
# 3 0.90 0.91 ok
# 4 0.91 0.91 ok
# 5 0.89 0.90 ok
# 6 0.88 0.89 ok
# 7 0.86 0.88 ok
# 8 0.84 0.86 ok
# 9 0.83 0.84 ALERT - retrain!
# 10 0.81 0.83 ALERT - retrain!
# 11 0.79 0.81 ALERT - retrain!
# 12 0.78 0.79 ALERT - retrain!
# 13 0.76 0.78 ALERT - retrain!
# 14 0.74 0.76 ALERT - retrain!
#
# Alert raised🎯 Your Turn: Raise the Alert
# 🎯 YOUR TURN — raise an alert when rolling accuracy dips
# Fill in each ___ , then run it.
daily = [0.90, 0.89, 0.88, 0.84, 0.82, 0.80]
window = 2
ALERT_THRESHOLD = 0.85
for i in range(len(daily)):
start = max(0, i - window + 1)
chunk = daily[start:i + 1]
rolling = sum(chunk) / len(___) # 👉 divide by how many days are in 'chunk'
if rolling ___ ALERT_THRESHOLD: # 👉 alert when BELOW the threshold
status = "ALERT"
else:
status = "ok"
print(f"Day {i+1}: rolling={rolling:.2f} -> {status}")
# ✅ Expected output:
# Day 1: rolling=0.90 -> ok
# Day 2: rolling=0.90 -> ok
# Day 3: rolling=0.89 -> ok
# Day 4: rolling=0.86 -> ok
# Day 5: rolling=0.83 -> ALERT
# Day 6: rolling=0.81 -> ALERT4 Logging, Alerting, and the Tools Teams Actually Use
The plain-Python examples show you the idea. In real systems you don't hand-roll the maths — you log every prediction and its inputs, push the numbers to a metric store, draw them on a dashboard, and let an alerting rule page you. A production stack usually layers up like this:
┌──────────────────────────────────────┐
│ Alerting (PagerDuty / Slack) │ <- pages a human on critical drift
├──────────────────────────────────────┤
│ Dashboard (Grafana / DataDog) │ <- humans watch trends here
├──────────────────────────────────────┤
│ Metric store (Prometheus/BigQuery) │ <- PSI, accuracy, latency over time
├──────────────────────────────────────┤
│ Collectors (input/output loggers) │ <- log every request + prediction
├──────────────────────────────────────┤
│ Model inference service │ <- your model answering requests
└──────────────────────────────────────┘For the drift maths itself, purpose-built libraries do the heavy lifting:
- Evidently / WhyLabs / NannyML — drift reports and statistical tests
- MLflow — track model versions and metrics across runs
- Great Expectations — validate incoming data (nulls, ranges, types)
- Fairlearn / AIF360 — fairness metrics across groups
Here's the same drift check from Section 2, but expressed with Evidently. It's read-only — the tool isn't installed in the editor — so study it as the production version of what you already built by hand.
# ============================================
# THE SAME CHECK WITH EVIDENTLY (a monitoring tool)
# ============================================
# In production you don't hand-roll drift maths. Tools like Evidently,
# WhyLabs, or NannyML run the statistical tests, render a report, and
# wire into alerting for you. Here is the Evidently equivalent.
#
# (Read-only: 'pip install evidently' is not available in this editor.)
import pandas as pd
from evidently.report import Report
from evidently.metric_preset import DataDriftPreset
# reference_df = the training data; current_df = today's live traffic
reference_df = pd.DataFrame({"income": [50, 51, 49, 52, 48, 50, 51, 49]})
current_df = pd.DataFrame({"income": [60, 62, 58, 64, 61, 59, 63, 60]})
report = Report(metrics=[DataDriftPreset()])
report.run(reference_data=reference_df, current_data=current_df)
# Pull out the machine-readable result so you can alert on it
result = report.as_dict()
drift = result["metrics"][0]["result"]
print("Columns drifted:", drift["number_of_drifted_columns"])
print("Dataset drift?:", drift["dataset_drift"])
# Expected output:
# Columns drifted: 1
# Dataset drift?: True🔁 Retraining Triggers — Closing the Loop
Monitoring is only useful if it leads to action. The action is usually a retrain — fit a fresh model on recent data so it re-learns the world as it is now. Don't retrain on a single bad reading; tie it to a tiered, persistent signal:
| Signal | Tier | Action |
|---|---|---|
| Mean/std shift within tolerance | Info | Log, keep watching |
| Moderate drift, a few days | Warning | Investigate inputs |
| Rolling accuracy below target, sustained | Critical | Retrain on fresh data |
| Pipeline returns nulls / broken feature | Critical | Fix data, then retrain |
The persistence rule ("sustained across a window") is what stops you retraining on a one-off spike — and it's the same idea as the rolling window you coded above.
5 Common Mistakes (And How to Fix Them)
Monitoring fails in predictable ways. Here are the four that bite teams most often:
❌ No monitoring at all
The model ships and nobody watches it. Accuracy quietly rots and the first signal is an angry customer.
✅ Fix: log every prediction with its inputs from day one, even if the only "dashboard" is a daily print of mean/std and rolling accuracy.
You watch accuracy only. But accuracy needs labels, and by the time it drops the inputs have been drifting for weeks.
✅ Fix: monitor input statistics too — they're available immediately and warn you before accuracy moves.
You wait for ground-truth labels that take months to arrive (did the loan default? did the user churn?), so your alerts are always late.
✅ Fix: lean on label-free signals (input drift, prediction-distribution shift) as your early warning; treat accuracy as confirmation, not detection.
Every tiny wobble fires a page. The team mutes the channel — and then misses the alert that actually mattered.
✅ Fix: tier alerts (info / warning / critical), only page for critical, and require the signal to persist across a rolling window before firing.
🎯 Mini-Challenge: Build a Tiny Watcher
Time to fade the scaffolding. You've detected drift and tracked rolling accuracy separately — now combine them into one daily watcher. The starter below gives you only a comment outline and the data. Write the loop yourself, then check it against the expected output in the comments.
# 🎯 MINI-CHALLENGE: a tiny monitoring loop
# Combine both signals you learned into one watcher.
#
# 1. You are given a list of (input_mean, accuracy) tuples, one per day.
# 2. For each day, flag DATA DRIFT if input_mean is more than 5 away
# from the reference mean of 50.
# 3. Also flag LOW ACCURACY if accuracy is below 0.85.
# 4. Print the day number and which alerts (if any) fired.
#
# ✅ Expected (for the data below):
# Day 1: ok
# Day 2: ok
# Day 3: DATA DRIFT
# Day 4: DATA DRIFT, LOW ACCURACY
reference_mean = 50
days = [
(50, 0.91), # input_mean, accuracy
(53, 0.88),
(58, 0.86),
(61, 0.80),
]
# your code here📋 Quick Reference — Model Monitoring
| What to watch | How | Tool | Alert when |
|---|---|---|---|
| Data drift | Mean / std vs reference, PSI | Evidently, WhyLabs | shift > tolerance |
| Concept drift | Rolling accuracy over time | MLflow, Prometheus | below baseline target |
| Latency | p50 / p95 / p99 timings | Grafana, DataDog | p99 > 200ms |
| Feature health | Null %, range checks | Great Expectations | > 5% nulls |
| Retrain trigger | Sustained critical signal | CI/CD pipeline | persists across window |
❓ Frequently Asked Questions
Lesson complete — your models now have vitals and alarms!
You can tell data drift from concept drift, detect drift by comparing the mean and std of two batches, smooth prediction quality with a rolling-accuracy metric, trip an alert below a threshold, and decide when a sustained signal should trigger a retrain. That's the full monitoring loop, from raw signal to action.
Practice quiz
What is data drift?
- The model's code changes over time
- The output labels are deleted
- The distribution of the input features shifts away from the training data
- The server latency increases
Answer: The distribution of the input features shifts away from the training data. Data drift means the inputs change shape compared to training, even though the model itself is unchanged.
What is concept drift?
- The relationship between inputs and the correct output changes
- The inputs change shape
- The model file becomes corrupted
- The dataset gets larger
Answer: The relationship between inputs and the correct output changes. Concept drift is when the same input should now map to a different output, e.g. what counts as a 'good salary' changes.
Why can you detect data drift without any labels?
- Because labels are never needed
- Because drift only affects outputs
- Because the model reports it automatically
- Because you compare input statistics (mean, std) against the training reference
Answer: Because you compare input statistics (mean, std) against the training reference. Input drift is measured purely from the incoming feature distribution, so no ground-truth labels are required.
What is label lag?
- The time to load a model
- The delay between making a prediction and learning whether it was correct
- The gap between two model versions
- The latency of an API call
Answer: The delay between making a prediction and learning whether it was correct. Label lag is the delay before ground truth arrives, which makes accuracy-based alerts late.
Why use a rolling window for accuracy instead of the raw daily value?
- It smooths out day-to-day noise so a genuine downward trend stands out
- It uses less memory
- It makes accuracy always higher
- It removes the need for labels
Answer: It smooths out day-to-day noise so a genuine downward trend stands out. A rolling average reduces noise so you alert on real trends rather than random single-day wobble.
In the monitoring stack, what fires when accuracy crosses below the alert threshold?
- The model retrains itself instantly
- The dataset is deleted
- An alert is raised (log, Slack, or page on-call)
- The server restarts
Answer: An alert is raised (log, Slack, or page on-call). Crossing the threshold raises an alert so a human or pipeline can act before users feel the damage.
Why monitor input drift in addition to accuracy?
- Input drift is harder to compute
- Input drift is visible immediately, while accuracy needs labels that may arrive late
- Accuracy is never useful
- Input drift requires no data
Answer: Input drift is visible immediately, while accuracy needs labels that may arrive late. Input drift is an early smoke alarm; falling accuracy is the fire that confirms the damage later.
When should sustained drift trigger an automatic retrain?
- After a single bad reading
- Never — retraining is manual only
- Whenever any input changes at all
- When a tiered signal persists across a window (e.g. rolling accuracy below target)
Answer: When a tiered signal persists across a window (e.g. rolling accuracy below target). Retraining should be tied to a persistent, tiered signal so a one-off spike does not trigger it.
What is alert fatigue and how do you avoid it?
- Too few alerts; add more pages
- Too many noisy alerts cause people to ignore them; tier alerts and require persistence
- Alerts that never fire; lower all thresholds
- A hardware failure in the alerting system
Answer: Too many noisy alerts cause people to ignore them; tier alerts and require persistence. Tiering alerts (info/warning/critical) and requiring the signal to persist prevents teams from muting alerts.
Which tools are commonly used specifically for drift reports and statistical tests?
- NumPy and pandas
- Grafana and PagerDuty
- Evidently, WhyLabs, and NannyML
- Git and Docker
Answer: Evidently, WhyLabs, and NannyML. Evidently, WhyLabs, and NannyML are purpose-built for drift detection and statistical monitoring.
Continue this course
- Previous: Serving ML Models: TorchServe, FastAPI, TensorFlow Serving
- Next: MLOps Fundamentals: Pipelines, CI/CD, Versioning — Automate ML pipelines with MLflow, DVC, and CI/CD for model releases
- Quick reference: AI & Machine Learning cheat sheet
- From the blog: Model Deployment: From Jupyter to Production
Frequently asked questions
What is the difference between data drift and concept drift?
Data drift means the inputs change shape — the distribution of features your model sees moves away from the training data (e.g. a new customer demographic). Concept drift means the relationship between inputs and the correct output changes, so the same input should now map to a different prediction (e.g. what counts as a 'good salary' rises over time). Data drift you can spot from inputs alone; concept drift usually shows up as falling accuracy once labels arrive.
How do I detect drift without any labels?
Compare the statistics of incoming inputs against the training data — mean, standard deviation, min/max, or a binned distribution. If those summary numbers move beyond a tolerance, the inputs have drifted. This is exactly what the plain-Python example does, and it needs no ground-truth labels, which is why it is your first line of defence in production.
Why use a rolling window for accuracy instead of the raw number?
A single day's accuracy is noisy — one unusual batch can make it spike or dip for reasons that don't matter. A rolling average over the last few days smooths that noise so a genuine downward trend stands out. You alert on the rolling value, not the raw value, to avoid firing on random wobble.
What is label lag and why does it matter for monitoring?
Label lag is the delay between making a prediction and learning whether it was correct. A loan model may not know if a borrower defaults for months, so accuracy-based alerts arrive late. That is why you also monitor input drift, which is available immediately — it warns you before the (delayed) accuracy metric confirms a problem.
When should drift trigger an automatic retrain?
Tie retraining to a tiered threshold rather than a single number. Minor drift logs and is watched; moderate drift opens an investigation; major sustained drift (or rolling accuracy below your service-level target) triggers retraining on fresh data. Always require the signal to persist across a window so you don't retrain on a one-off spike.
How do I avoid alert fatigue?
Tier your alerts — info, warning, critical — and only page a human for critical, sustained problems. Use rolling windows and 'must persist for N periods' rules so transient blips stay silent. If every wobble pages the team, people start ignoring alerts, and the one that matters gets missed too.