Linear Regression
Reviewed & published by Brayan K
Your first real machine-learning algorithm. By the end you'll fit a line to data by hand, measure its error with a cost function, watch gradient descent learn the line for you, split data to test honestly, and run it all in one line with scikit-learn.
Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
What You'll Learn in This Lesson
- What a line y = wx + b represents as a model
- How the MSE cost function scores a fit
- What gradient descent does, step by step
- How to code a tiny training loop in plain Python
- Why you split data into training and test sets
- How to do it for real with sklearn's LinearRegression
🌍 Real-World Analogy
Imagine a seasoned estate agent who has watched hundreds of houses sell. Over time they develop a gut feeling: "bigger houses sell for more." Linear regression turns that gut feeling into an exact formula — something like price = 150 × size + 50000.
The agent's instinct is fuzzy; the formula is precise and repeatable. That is the whole job of this lesson: take a scatter of dots and find the single straight line that best summarises the trend, so you can predict the next dot you haven't seen yet.
1 The Model Is Just a Line: y = wx + b
A linear-regression model is nothing more than two numbers. You feed it an input x (a feature, like study hours) and it returns a prediction y using the straight-line formula:
- • x — the input feature you measure
- • w — the weight (slope): how much y changes per unit of x
- • b — the bias (intercept): the value of y when x is 0
- • y — the prediction the model produces
"Training" the model just means picking the w and b that draw the line closest to all your data points. Everything else in this lesson is about how we pick them. The example below starts with a hand-guessed line and a function that applies it.
2 Scoring the Fit: the MSE Cost Function
How do you know one line is better than another? You need a single number that says "this is how wrong you are." That number is the cost function, and the standard one for regression is Mean Squared Error (MSE):
MSE = average of (actual − predicted)²
For each point you take the gap between the real value and the model's guess (the error), square it, then average across all points. Squaring forces every error to be positive (so a +5 and a −5 don't cancel) and makes big misses hurt much more than small ones. Lower MSE means a better fit. The worked example below prints each error and then the MSE.
# === Fitting a line y = w*x + b, then measuring how good it is ===
# Plain Python only — no numpy needed. Data is just two lists.
# Dataset: study hours -> exam score
hours = [1, 2, 3, 4, 5] # the feature (x)
scores = [52, 58, 67, 71, 79] # the target (y)
# A "model" is just two numbers: a slope w and an intercept b.
# Read it as: "score = w * hours + b".
w = 6.5 # each extra hour adds ~6.5 points (a guess for now)
b = 46.0 # a student who studied 0 hours scores ~46 (a guess for now)
# predict(x) applies the line to one input.
def predict(x):
return w * x + b # the straight-line formula
# Show each prediction next to the real value.
print("hours actual predicted error")
for x, y in zip(hours, scores):
yhat = predict(x) # the model's guess
error = y - yhat # how far off we were (can be + or -)
print(f"{x:>5} {y:>6} {yhat:>9.1f} {error:>+6.1f}")
# MSE = Mean Squared Error = average of (actual - predicted) squared.
# Squaring makes every error positive and punishes big misses harder.
total = 0.0
for x, y in zip(hours, scores):
total += (y - predict(x)) ** 2 # add up each squared error
mse = total / len(hours) # divide by how many points
print()
print(f"MSE = {mse:.2f}") # lower MSE = the line fits better
# Expected output: MSE = 1.30 (this line happens to fit quite well)🎯 Your Turn: build the MSE calculator
# 🎯 YOUR TURN — finish the MSE (Mean Squared Error) calculator
# Fill in each blank marked with ___ . Plain Python, no numpy.
actual = [10, 20, 30, 40] # the real values
predicted = [12, 18, 33, 39] # a model's guesses
total = 0.0
for a, p in zip(actual, predicted):
# 👉 1) the error is (actual - predicted); square it and add to total
total += (a - p) ___ 2 # 👉 replace ___ with the power operator (**)
# 👉 2) MSE is the average, so divide the total by the number of points
mse = total / ___ # 👉 replace ___ with len(actual)
print(f"MSE = {mse:.2f}")
# ✅ Expected output: MSE = 4.503 Gradient Descent: How the Model Learns
Guessing w and b by hand doesn't scale. Gradient descent is the algorithm that finds them automatically. Picture the MSE as a valley: every (w, b) is a spot on the hillside, and the lowest point is the best line. Gradient descent is like walking downhill blindfolded — feel which way the ground slopes, take a small step that way, and repeat.
- Gradient — the direction the cost increases fastest. You step the opposite way (downhill).
- Learning rate — how big each step is. Too big and you overshoot the valley; too small and training crawls.
- Epoch — one full pass over the data that nudges w and b once.
# === Gradient descent: let the computer FIND w and b for us ===
# This is the core loop behind how almost every ML model learns.
# Still plain Python — only the 'lists + arithmetic' you already know.
# Same study-hours data
hours = [1, 2, 3, 4, 5]
scores = [52, 58, 67, 71, 79]
n = len(hours)
# Start from a deliberately bad guess and let the loop fix it.
w = 0.0
b = 0.0
learning_rate = 0.01 # step size: how big a nudge we take each round
# One "epoch" = one full pass over the data that nudges w and b downhill.
for epoch in range(2000):
# 1) Predict for every point and collect the errors.
err_sum = 0.0 # will become the gradient for b
err_x_sum = 0.0 # will become the gradient for w
for x, y in zip(hours, scores):
yhat = w * x + b # current prediction
error = yhat - y # signed error (predicted - actual)
err_sum += error # accumulate for b's slope
err_x_sum += error * x # accumulate for w's slope
# 2) Gradients tell us which way is "downhill" on the MSE surface.
grad_w = (2 / n) * err_x_sum
grad_b = (2 / n) * err_sum
# 3) Step a little in the opposite (downhill) direction.
w -= learning_rate * grad_w
b -= learning_rate * grad_b
# After 2000 small steps, w and b have settled near the best line.
print(f"Learned model: score = {w:.2f} * hours + {b:.2f}")
print(f"Prediction for 6 hours: {w*6 + b:.1f}")
# Expected output (approx):
# Learned model: score = 6.70 * hours + 45.30
# Prediction for 6 hours: 85.5🎯 Your Turn: complete the update step
# 🎯 YOUR TURN — complete ONE gradient-descent step
# The loop body is written for you except for the two update lines.
xs = [1, 2, 3, 4]
ys = [3, 5, 7, 9] # the true line is y = 2x + 1
n = len(xs)
w = 0.0
b = 0.0
learning_rate = 0.05
for epoch in range(1000):
err_sum = err_x_sum = 0.0
for x, y in zip(xs, ys):
error = (w * x + b) - y # predicted - actual
err_sum += error
err_x_sum += error * x
grad_w = (2 / n) * err_x_sum
grad_b = (2 / n) * err_sum
# 👉 Move w and b DOWNHILL: subtract learning_rate * gradient
w = w - learning_rate * ___ # 👉 replace ___ with grad_w
b = b - learning_rate * ___ # 👉 replace ___ with grad_b
print(f"w = {w:.2f}, b = {b:.2f}")
# ✅ Expected output (approx): w = 2.00, b = 1.004 Training vs Test: the Honesty Check
A model that scores its own homework will always claim an A. To get an honest measure of how it performs on data it has never seen, you split your data into two parts before training:
- Training set (usually ~80%) — the model learns w and b from these.
- Test set (usually ~20%) — held back, used only to score the finished model.
If the model does great on the training set but poorly on the test set, it has overfit — it memorised the specific points instead of learning the underlying trend. The test score is the number you actually trust. In real projects, train_test_split from scikit-learn does this split for you in one line.
5 The Real Way: scikit-learn
You now understand what happens under the hood — so in real projects you let a library do it. scikit-learn (imported as sklearn) gives you LinearRegression: create it, call .fit(X, y) to train, and .predict(...) to use it. It runs the maths far faster and more reliably than hand-rolled code.
Read the example below as a worked walkthrough — it needs scikit-learn installed locally to run, and the comments show the expected output. Notice the w and b it finds match what your gradient-descent loop learned above.
# === The real-world way: scikit-learn does the maths for you ===
# In practice you do NOT hand-write gradient descent. You call a library.
# (This needs scikit-learn installed: pip install scikit-learn)
from sklearn.linear_model import LinearRegression
# sklearn expects X as a list of rows; each row is one sample's features.
X = [[1], [2], [3], [4], [5]] # study hours (one feature per row)
y = [52, 58, 67, 71, 79] # exam scores
model = LinearRegression() # create the (untrained) model
model.fit(X, y) # .fit() finds the best w and b for you
print(f"slope w = {model.coef_[0]:.2f}")
print(f"intercept b = {model.intercept_:.2f}")
# .predict() takes the same row shape and returns predictions.
predictions = model.predict([[6], [7]])
print(f"6 hours -> {predictions[0]:.1f}")
print(f"7 hours -> {predictions[1]:.1f}")
# Expected output:
# slope w = 6.70
# intercept b = 45.30
# 6 hours -> 85.5
# 7 hours -> 92.2Common Errors (And How to Fix Them)
❌ Not splitting the data
You report 99% accuracy, ship the model, and it fails on real users.
✅ Fix: hold back a test set the model never trains on, and judge it only by the test score. Use train_test_split(X, y, test_size=0.2).
❌ Features on wildly different scales
One feature is in the thousands (square feet) and another is single digits (bedrooms). Gradient descent zig-zags or diverges, and weights become hard to compare.
✅ Fix: scale features to a similar range first (e.g. StandardScaler) before training. A smaller learning rate also helps stop the cost blowing up.
❌ Assuming the relationship is linear
Your data curves, but you force a straight line through it. The MSE stays high no matter how long you train.
✅ Fix: plot the data first. If it bends, add polynomial features, transform a column (e.g. a log), or pick a model that can capture curves.
❌ Extrapolating far beyond the training range
You trained on houses 600–2000 sq ft, then predict a 10,000 sq ft mansion. The line keeps going forever, but reality doesn't.
✅ Fix: only trust predictions inside (or near) the range you trained on. Flag inputs far outside it as unreliable.
📋 Quick Reference
| Term | Formula / Code | Meaning |
|---|---|---|
| Model | y = w*x + b | A line: weight w, bias b |
| MSE (cost) | mean((y - ŷ)²) | How wrong the line is (lower is better) |
| Gradient step | w -= lr * grad_w | Nudge parameters downhill |
| Learning rate | lr = 0.01 | Step size each epoch |
| Train / test split | train_test_split(...) | Learn on one slice, score on another |
| sklearn fit | model.fit(X, y) | Finds best w and b for you |
| sklearn predict | model.predict(X) | Apply the trained line |
Mini-Challenge: pick the better line
No blanks this time — just a brief and a comment outline. Write the two helper functions, score both candidate lines with MSE, and print which one fits better. Check your answer against the expected result in the comments.
# 🎯 MINI-CHALLENGE: compare two candidate lines and pick the better one
#
# Data (temperature -> ice creams sold):
# temps = [15, 18, 20, 25, 30]
# sales = [22, 33, 40, 62, 81]
#
# 1. Write predict(x, w, b) that returns w * x + b
# 2. Write mse(w, b) that loops over the data and returns the Mean Squared Error
# 3. Score line A: w = 3.0, b = -25
# 4. Score line B: w = 3.0, b = -20
# 5. Print both MSEs and print which line fits better (lower MSE wins)
#
# ✅ Expected: line B has the lower MSE, so it fits better.
temps = [15, 18, 20, 25, 30]
sales = [22, 33, 40, 62, 81]
# your code here🎉 Lesson Complete!
You've built your first ML model from the ground up: you fit a line y = wx + b, scored it with the MSE cost function, wrote a gradient-descent loop that learns w and b on its own, understood why training and test sets matter, and saw how scikit-learn does it all in three lines.
Practice quiz
In the model y = w*x + b, what does w represent?
- The intercept (value of y when x is 0)
- The prediction error
- The weight or slope: how much y changes per unit of x
- The learning rate
Answer: The weight or slope: how much y changes per unit of x. w is the weight (slope) — it controls how much the prediction y rises for each unit increase in x.
What does the bias term b represent in y = w*x + b?
- The value of y when x is 0
- The slope of the line
- The mean squared error
- The number of epochs
Answer: The value of y when x is 0. b is the intercept — the predicted value of y when the input x equals 0.
What does Mean Squared Error (MSE) measure?
- The sum of all predictions
- The slope of the line
- The number of data points
- The average of (actual - predicted) squared
Answer: The average of (actual - predicted) squared. MSE averages the squared errors; lower MSE means the line fits the data better.
Why are the errors squared in MSE?
- To make training faster
- To make every error positive and penalise large misses more
- To turn the output into a probability
- To reduce the number of features
Answer: To make every error positive and penalise large misses more. Squaring stops +5 and -5 from cancelling and makes big errors hurt far more than small ones.
What is gradient descent?
- A loop that nudges parameters downhill on the cost surface to reduce error
- A way to split data into train and test
- A method to square the errors
- A library function that plots data
Answer: A loop that nudges parameters downhill on the cost surface to reduce error. Gradient descent repeatedly steps the parameters in the direction that lowers the cost (downhill).
What does the learning rate control?
- How many features the model uses
- The number of test rows
- The size of each step taken during gradient descent
- The intercept value
Answer: The size of each step taken during gradient descent. The learning rate is the step size: too large overshoots the minimum, too small makes training crawl.
To move 'downhill' in gradient descent, you update a parameter by...
- Adding the learning rate times the gradient
- Subtracting the learning rate times the gradient
- Multiplying by the gradient
- Dividing by the gradient
Answer: Subtracting the learning rate times the gradient. You subtract learning_rate * gradient so the parameter moves opposite to the increasing-cost direction.
What is one epoch in training?
- One data point
- One prediction
- One test/train split
- One full pass over the training data that nudges the parameters
Answer: One full pass over the training data that nudges the parameters. An epoch is a complete pass over the dataset used to update the parameters once (in batch gradient descent).
What does it mean if a model fits the training set well but the test set poorly?
- It has underfit
- It has overfit — memorised noise instead of the real pattern
- It needs a smaller learning rate
- The MSE is negative
Answer: It has overfit — memorised noise instead of the real pattern. Overfitting means the model memorised the training points and fails to generalise to unseen data.
When is linear regression the wrong choice?
- When predicting a continuous number
- When you have a train/test split
- When the true relationship is not roughly a straight line
- When you use scikit-learn
Answer: When the true relationship is not roughly a straight line. If the data curves, a straight line fits poorly; you need polynomial features, a transform, or another model.
Continue this course
- Previous: Data Preprocessing
- Next: Classification Basics — Categorise data into classes using logistic regression and k-NN
- Quick reference: AI & Machine Learning cheat sheet › Common Algorithms
- From the blog: Getting Started with Machine Learning in Python
Frequently asked questions
What does linear regression actually predict?
A continuous number — a value on a sliding scale rather than a category. Think house prices, exam scores, temperatures, or sales figures. If the answer you want is 'how much' or 'how many', linear regression is a sensible first tool. If the answer is a label like 'spam vs not spam' or 'cat vs dog', that is classification instead, which you'll meet in the next lesson.
What is the cost function and why square the errors?
The cost function is a single number that measures how wrong the model is across all the data, and the standard choice for regression is Mean Squared Error (MSE): the average of (actual - predicted) squared. Squaring does two jobs at once — it makes every error positive so they can't cancel out, and it punishes large mistakes far more than small ones (an error of 10 contributes 100, an error of 1 contributes only 1). Training means finding the w and b that make this number as small as possible.
What is gradient descent in one sentence?
It is a loop that repeatedly nudges the model's parameters a tiny step 'downhill' on the cost surface — compute which direction lowers the error (the gradient), take a small step that way (controlled by the learning rate), and repeat until the cost stops shrinking. It is the same core idea that trains neural networks, just with many more parameters.
Why do I split data into training and test sets?
Because a model that has memorised the exact answers it was trained on can look perfect and still fail on anything new. You fit (train) on one slice of the data and check accuracy on a separate test slice the model never saw. If it does well on training data but badly on the test set, it has overfit — it memorised noise instead of learning the real pattern. The test score is your honest estimate of real-world performance.
When is linear regression the wrong choice?
When the real relationship is not roughly a straight line. If sales rise then plateau, or price grows exponentially with size, a straight line will fit poorly no matter how you train it. Plot your data first — if it curves, you need polynomial features, a transform (like taking a log), or a different model entirely. Forcing a line onto a curve gives confident-looking but wrong predictions.