MLOps Fundamentals
Reviewed & published by Brayan K
Automate the entire ML lifecycle — from experiment tracking to automated retraining pipelines and model versioning.
Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
What You'll Learn in This Lesson
- • Track experiments reproducibly with MLflow-style logging
- • Build automated ML pipelines with validation gates
- • Version data with DVC alongside Git for code
- • Manage model lifecycle: staging → production → archived
- • Run A/B tests to safely promote new models
1️⃣ The MLOps Lifecycle
MLOps bridges the gap between "it works on my laptop" and "it serves 10M users":
Data → Feature Store → Training → Evaluation → Registry → Deploy → Monitor
↑ |
└──────────────── Automated Retraining Trigger ←──────────────────────┘| Component | Tool | Purpose |
|---|---|---|
| Experiment Tracking | MLflow, W&B | Log params, metrics, artifacts |
| Data Versioning | DVC, LakeFS | Track dataset versions |
| Pipeline Orchestration | Airflow, Kubeflow | Automate train→deploy |
| Feature Store | Feast, Tecton | Reusable feature pipelines |
| Model Registry | MLflow Registry | Version + lifecycle management |
import numpy as np
import json
from datetime import datetime
# ============================================
# EXPERIMENT TRACKING WITH MLFLOW-STYLE LOGGING
# ============================================
np.random.seed(42)
print("=== ML Experiment Tracking ===")
print()
print("Without experiment tracking, you lose track of which")
print("hyperparameters produced which results. MLflow solves this.")
print()
class ExperimentTracker:
"""Simplified MLflow-style experiment tracker."""
def __init__(self, experiment_name):
self.name = experiment_name
self.runs = []
def start_run(self, run_name):
self.current_run = {
"name": run_name,
"params": {},
"metrics": {},
"timestamp": datetime.now().isoformat()
}
return self
def log_param(self, key, value):
self.current_run["params"][key] = value
def log_metric(self, key, value):
self.current_run["metrics"][key] = round(value, 4)
def end_run(self):
self.runs.append(self.current_run)
print(f" ✅ Run '{self.current_run['name']}' logged")
def get_best_run(self, metric, higher_is_better=True):
return sorted(self.runs,
key=lambda r: r["metrics"].get(metric, 0),
reverse=higher_is_better)[0]
# Create experiment
tracker = ExperimentTracker("loan_default_prediction")
# Run multiple experiments with different hyperparameters
configs = [
{"name": "baseline_lr", "model": "LogisticRegression", "C": 1.0, "max_iter": 100},
{"name": "tuned_lr", "model": "LogisticRegression", "C": 0.1, "max_iter": 500},
{"name": "random_forest", "model": "RandomForest", "n_estimators": 100, "max_depth": 10},
{"name": "xgboost_v1", "model": "XGBoost", "learning_rate": 0.1, "n_estimators": 200},
{"name": "xgboost_v2", "model": "XGBoost", "learning_rate": 0.05, "n_estimators": 500},
]
# Simulated results
results = [
{"accuracy": 0.82, "f1": 0.78, "auc": 0.85},
{"accuracy": 0.84, "f1": 0.81, "auc": 0.87},
{"accuracy": 0.88, "f1": 0.85, "auc": 0.91},
{"accuracy": 0.90, "f1": 0.87, "auc": 0.93},
{"accuracy": 0.91, "f1": 0.89, "auc": 0.94},
]
print("Logging experiment runs:")
for config, result in zip(configs, results):
tracker.start_run(config["name"])
for k, v in config.items():
if k != "name":
tracker.log_param(k, v)
for k, v in result.items():
tracker.log_metric(k, v)
tracker.end_run()
print()
print("=== Experiment Results ===")
print(f"{'Run':<20s} {'Model':<20s} {'AUC':>6s} {'F1':>6s}")
print("-" * 55)
for run in tracker.runs:
print(f"{run['name']:<20s} {run['params']['model']:<20s} {run['metrics']['auc']:>6.3f} {run['metrics']['f1']:>6.3f}")
best = tracker.get_best_run("auc")
print(f"\n🏆 Best run: {best['name']} (AUC={best['metrics']['auc']})")
print(f" Params: {json.dumps(best['params'], indent=2)}")import numpy as np
# ============================================
# ML PIPELINE AUTOMATION (CI/CD for ML)
# ============================================
np.random.seed(42)
print("=== Automated ML Pipeline ===")
print()
print("An ML pipeline automates: data → train → evaluate → deploy")
print("Think of it like a factory assembly line for models.")
print()
class MLPipeline:
"""Simplified ML pipeline with stages."""
def __init__(self, name):
self.name = name
self.stages = []
self.artifacts = {}
def add_stage(self, name, func):
self.stages.append({"name": name, "func": func})
def run(self):
print(f"🚀 Pipeline '{self.name}' started")
print()
for i, stage in enumerate(self.stages):
print(f" Stage {i+1}/{len(self.stages)}: {stage['name']}")
try:
result = stage["func"](self.artifacts)
self.artifacts.update(result or {})
print(f" ✅ Passed")
except Exception as e:
print(f" ❌ Failed: {e}")
print(f"\n🛑 Pipeline halted at stage '{stage['name']}'")
return False
print(f"\n✅ Pipeline '{self.name}' completed successfully!")
return True
# Define pipeline stages
def data_validation(artifacts):
data = np.random.randn(1000, 5)
null_pct = 0.002
print(f" Rows: {len(data)}, Features: {data.shape[1]}, Null%: {null_pct:.1%}")
if null_pct > 0.05:
raise ValueError(f"Too many nulls: {null_pct:.1%}")
return {"data": data, "rows": len(data)}
def feature_engineering(artifacts):
n_features = artifacts["data"].shape[1]
new_features = n_features + 3 # polynomial + interaction
print(f" Original: {n_features} → Engineered: {new_features} features")
return {"n_features": new_features}
def model_training(artifacts):
accuracy = 0.91 + np.random.uniform(-0.02, 0.02)
print(f" Model: XGBoost, Accuracy: {accuracy:.3f}")
return {"accuracy": accuracy, "model_version": "v2.4.1"}
def model_evaluation(artifacts):
acc = artifacts["accuracy"]
threshold = 0.85
print(f" Accuracy: {acc:.3f} vs threshold: {threshold}")
if acc < threshold:
raise ValueError(f"Accuracy {acc:.3f} below threshold {threshold}")
return {"passed_eval": True}
def model_registry(artifacts):
version = artifacts["model_version"]
print(f" Registered model {version} as 'production'")
return {"deployed": True}
# Build and run pipeline
pipeline = MLPipeline("loan_default_v2")
pipeline.add_stage("Data Validation", data_validation)
pipeline.add_stage("Feature Engineering", feature_engineering)
pipeline.add_stage("Model Training", model_training)
pipeline.add_stage("Model Evaluation", model_evaluation)
pipeline.add_stage("Model Registry", model_registry)
pipeline.run()
print()
print("=== DVC (Data Version Control) ===")
print()
print("DVC tracks data files like Git tracks code:")
print(" dvc init")
print(" dvc add data/train.csv # Hash + .gitignore large file")
print(" dvc push # Upload to S3/GCS")
print(" dvc pull # Download on another machine")
print(" git checkout v1.0 && dvc checkout # Reproduce old experiment")import numpy as np
import json
# ============================================
# MODEL VERSIONING & A/B TESTING
# ============================================
np.random.seed(42)
print("=== Model Registry & Versioning ===")
print()
class ModelRegistry:
"""Track model versions and their lifecycle stages."""
def __init__(self):
self.models = {}
def register(self, name, version, metrics, stage="staging"):
key = f"{name}/{version}"
self.models[key] = {
"name": name, "version": version,
"metrics": metrics, "stage": stage
}
print(f" Registered {key} → {stage}")
def promote(self, name, version, new_stage):
key = f"{name}/{version}"
old = self.models[key]["stage"]
self.models[key]["stage"] = new_stage
print(f" {key}: {old} → {new_stage}")
def get_production(self, name):
for key, model in self.models.items():
if model["name"] == name and model["stage"] == "production":
return model
return None
registry = ModelRegistry()
registry.register("fraud_detector", "v1.0", {"f1": 0.85, "auc": 0.90})
registry.register("fraud_detector", "v2.0", {"f1": 0.89, "auc": 0.93})
registry.register("fraud_detector", "v2.1", {"f1": 0.91, "auc": 0.95})
print()
registry.promote("fraud_detector", "v2.1", "production")
registry.promote("fraud_detector", "v2.0", "archived")
print()
print("=== A/B Testing Models ===")
print()
print("Route traffic between model versions to compare real performance.")
print()
n_requests = 10000
split = 0.9 # 90% to current champion
champion_correct = np.random.binomial(1, 0.89, int(n_requests * split))
challenger_correct = np.random.binomial(1, 0.92, int(n_requests * (1 - split)))
champion_acc = champion_correct.mean()
challenger_acc = challenger_correct.mean()
print(f" Champion (v2.0): {int(n_requests * split):,} requests → accuracy {champion_acc:.3f}")
print(f" Challenger (v2.1): {int(n_requests * (1-split)):,} requests → accuracy {challenger_acc:.3f}")
print()
# Statistical significance check
from_diff = challenger_acc - champion_acc
se = np.sqrt(champion_acc * (1-champion_acc) / len(champion_correct) +
challenger_acc * (1-challenger_acc) / len(challenger_correct))
z_score = from_diff / se
print(f" Improvement: +{from_diff:.3f}")
print(f" Z-score: {z_score:.2f}")
print(f" Significant: {'✅ Yes (z > 1.96)' if abs(z_score) > 1.96 else '❌ Not yet'}")
if abs(z_score) > 1.96 and challenger_acc > champion_acc:
print(f"\n 🏆 Promote challenger v2.1 to production!")
else:
print(f"\n ⏳ Continue A/B test, need more data")🚦 Worked example: promotion is a gate, not a button
The lifecycle above says models are versioned and promoted. Here is what that means in code you can read: a registry, a production pointer, and a gate. Watch v2 — it is genuinely better than v1 and is still rejected, because a 0.005 gain is indistinguishable from noise, and shipping noise is how a system slowly gets worse while every dashboard says it improved.
# WORKED EXAMPLE — a model registry with a promotion gate, in plain Python.
# "MLOps" is largely bookkeeping like this, done reliably.
MIN_GAIN = 0.01 # a candidate must beat production by at least this much
registry = {} # version -> record
production = None # which version is live
def register(version, f1):
registry[version] = {"f1": f1, "stage": "staging"}
def promote(version):
global production
candidate = registry[version]["f1"]
if production is not None:
incumbent = registry[production]["f1"]
if candidate < incumbent + MIN_GAIN:
print(f"{version}: rejected (f1 {candidate:.3f} does not beat {production} at {incumbent:.3f} by {MIN_GAIN})")
return
registry[production]["stage"] = "archived"
registry[version]["stage"] = "production"
production = version
print(f"{version}: promoted (f1 {candidate:.3f})")
register("v1", 0.850); promote("v1")
register("v2", 0.855); promote("v2") # better, but not by enough
register("v3", 0.900); promote("v3")
print()
print(f"{'version':<8}{'f1':>6} stage")
for v, rec in registry.items():
print(f"{v:<8}{rec['f1']:>6.3f} {rec['stage']}")
print()
print("The gate is the point: without it v2 would have shipped for a 0.005 gain,")
print("and every future comparison would be against noise.")
# ✅ Expected output:
# v1: promoted (f1 0.850)
# v2: rejected (f1 0.855 does not beat v1 at 0.850 by 0.01)
# v3: promoted (f1 0.900)
#
# version f1 stage
# v1 0.850 archived
# v2 0.855 staging
# v3 0.900 production
#
# The gate is the point: without it v2 would have shipped for a 0.005 gain,
# and every future comparison would be against noise.🎯 Your turn: put the module together
This is the monitoring loop turned into ten lines. The data is chosen so that one week drifts without losing accuracy and another loses accuracy without drifting; if your trigger fires exactly twice, both rules are wired correctly.
# 🎯 YOUR TURN — the automated retraining trigger from the lifecycle diagram.
# Each week you get the mean of the incoming data and the live accuracy.
# Retrain when the inputs have drifted OR accuracy has dropped too far.
# Fill in the three ___ blanks.
reference_mean = 50.0 # what the training data looked like
DRIFT_LIMIT = 5.0 # how far the input mean may move before we retrain
ACCURACY_FLOOR = 0.90
weeks = [(50.4, 0.94), (51.2, 0.93), (56.1, 0.92), (50.8, 0.88), (49.9, 0.95)]
version = 1
for week, (input_mean, accuracy) in enumerate(weeks, start=1):
# 1) drift is how far this week's input mean sits from the reference
shift = abs(input_mean - ___)
# 👉 replace ___ with reference_mean
# 2) retrain if the drift is too large, or accuracy fell below the floor
if shift > ___ or accuracy < ACCURACY_FLOOR:
# 👉 replace ___ with DRIFT_LIMIT
# 3) bump the model version to record the retrain
version ___ 1
# 👉 replace ___ with +=
print(f"week {week}: RETRAIN -> v{version} (shift {shift:.1f}, accuracy {accuracy:.2f})")
else:
print(f"week {week}: ok (shift {shift:.1f}, accuracy {accuracy:.2f})")
print(f"final version: v{version}")
# ✅ Expected output:
# week 1: ok (shift 0.4, accuracy 0.94)
# week 2: ok (shift 1.2, accuracy 0.93)
# week 3: RETRAIN -> v2 (shift 6.1, accuracy 0.92)
# week 4: RETRAIN -> v3 (shift 0.8, accuracy 0.88)
# week 5: ok (shift 0.1, accuracy 0.95)
# final version: v3⚠️ Common Mistakes
📋 Quick Reference — MLOps
| Concept | Command / Tool |
|---|---|
| Track experiment | mlflow.log_param(), mlflow.log_metric() |
| Version data | dvc add data.csv → dvc push |
| Register model | mlflow.register_model() |
| Promote model | transition_model_version_stage("Production") |
| A/B test | Route 90/10 traffic, check z-score > 1.96 |
🎉 Lesson Complete!
You've mastered MLOps fundamentals! Next, learn how to build recommendation systems that power Netflix, Spotify, and Amazon.
Practice quiz
What problem does MLOps primarily solve?
- Making models more accurate than humans
- Replacing data scientists
- Bridging the gap between 'works on my laptop' and reliable production at scale
- Eliminating the need for data
Answer: Bridging the gap between 'works on my laptop' and reliable production at scale. MLOps automates and operationalises the ML lifecycle so models are reproducible, tested, and reliably deployed.
What is the purpose of experiment tracking (e.g. MLflow)?
- To log parameters, metrics, and artifacts so you know which config produced which result
- To speed up model inference
- To compress the model
- To version the source code only
Answer: To log parameters, metrics, and artifacts so you know which config produced which result. Experiment tracking records hyperparameters and metrics across runs so results are reproducible and comparable.
What does DVC (Data Version Control) do?
- Trains models automatically
- Serves models over HTTP
- Monitors live predictions
- Tracks data file versions alongside Git tracking code
Answer: Tracks data file versions alongside Git tracking code. DVC versions large data files (hashing and storing them) so datasets can be reproduced like code in Git.
What does CI/CD bring to an ML pipeline?
- It removes the need for testing
- It automates the data to train to evaluate to deploy flow with validation gates
- It only applies to web apps, not ML
- It increases model size
Answer: It automates the data to train to evaluate to deploy flow with validation gates. CI/CD for ML automates the pipeline stages and halts on failures, making deployment repeatable and safe.
Why add a validation gate (e.g. accuracy threshold) in a pipeline?
- To halt deployment when the model fails to meet a quality bar
- To make training faster
- To version the data
- To reduce GPU cost
Answer: To halt deployment when the model fails to meet a quality bar. A validation gate stops a model that scores below threshold from being promoted to production.
What is a model registry used for?
- Storing raw training data
- Routing user traffic
- Tracking model versions and their lifecycle stages (staging, production, archived)
- Logging server errors
Answer: Tracking model versions and their lifecycle stages (staging, production, archived). A model registry versions models and manages their lifecycle from staging through production to archived.
What is the typical lifecycle progression for a registered model?
- production then staging then archived
- staging then production then archived
- archived then production then staging
- production then archived then staging
Answer: staging then production then archived. Models usually move from staging (testing) to production (live) and eventually to archived when retired.
What is A/B testing of models?
- Training two models on the same data
- Compressing a model into two parts
- Versioning two datasets
- Routing live traffic between versions to compare real-world performance
Answer: Routing live traffic between versions to compare real-world performance. A/B testing splits live traffic between a champion and challenger model to measure which performs better.
In the A/B test example, when is the improvement considered statistically significant?
- When accuracy is above 90%
- When the absolute z-score exceeds 1.96
- When there are more than 1000 requests
- Whenever the challenger wins once
Answer: When the absolute z-score exceeds 1.96. A z-score above 1.96 corresponds to roughly 95% confidence that the difference is not due to chance.
Why is training only in notebooks and deploying manually a problem?
- Notebooks are too slow
- Notebooks cannot run Python
- It is not reproducible — if it is not in a pipeline, it cannot be reliably reproduced
- It uses too much memory
Answer: It is not reproducible — if it is not in a pipeline, it cannot be reliably reproduced. Manual notebook workflows are not reproducible; pipelines make the process repeatable and auditable.
Continue this course
- Previous: Monitoring Models in Production (Drift, Outliers, Bias)
- Next: Building Recommender Systems (Content, Collaborative, Hybrid) — Build content-based, collaborative filtering, and hybrid recommendation engines
- Quick reference: AI & Machine Learning cheat sheet
- From the blog: Model Deployment: From Jupyter to Production