Linear Regression

Reviewed & published by Brayan K

Your first real machine-learning algorithm. By the end you'll fit a line to data by hand, measure its error with a cost function, watch gradient descent learn the line for you, split data to test honestly, and run it all in one line with scikit-learn.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

🌍 Real-World Analogy

Imagine a seasoned estate agent who has watched hundreds of houses sell. Over time they develop a gut feeling: "bigger houses sell for more." Linear regression turns that gut feeling into an exact formula — something like price = 150 × size + 50000.

The agent's instinct is fuzzy; the formula is precise and repeatable. That is the whole job of this lesson: take a scatter of dots and find the single straight line that best summarises the trend, so you can predict the next dot you haven't seen yet.

1 The Model Is Just a Line: y = wx + b

A linear-regression model is nothing more than two numbers. You feed it an input x (a feature, like study hours) and it returns a prediction y using the straight-line formula:

"Training" the model just means picking the w and b that draw the line closest to all your data points. Everything else in this lesson is about how we pick them. The example below starts with a hand-guessed line and a function that applies it.

2 Scoring the Fit: the MSE Cost Function

How do you know one line is better than another? You need a single number that says "this is how wrong you are." That number is the cost function, and the standard one for regression is Mean Squared Error (MSE):

MSE = average of (actual − predicted)²

For each point you take the gap between the real value and the model's guess (the error), square it, then average across all points. Squaring forces every error to be positive (so a +5 and a −5 don't cancel) and makes big misses hurt much more than small ones. Lower MSE means a better fit. The worked example below prints each error and then the MSE.

# === Fitting a line y = w*x + b, then measuring how good it is ===
# Plain Python only — no numpy needed. Data is just two lists.

# Dataset: study hours -> exam score
hours  = [1, 2, 3, 4, 5]          # the feature (x)
scores = [52, 58, 67, 71, 79]     # the target  (y)

# A "model" is just two numbers: a slope w and an intercept b.
# Read it as: "score = w * hours + b".
w = 6.5   # each extra hour adds ~6.5 points (a guess for now)
b = 46.0  # a student who studied 0 hours scores ~46 (a guess for now)

# predict(x) applies the line to one input.
def predict(x):
    return w * x + b          # the straight-line formula

# Show each prediction next to the real value.
print("hours  actual  predicted  error")
for x, y in zip(hours, scores):
    yhat = predict(x)         # the model's guess
    error = y - yhat          # how far off we were (can be + or -)
    print(f"{x:>5}  {y:>6}  {yhat:>9.1f}  {error:>+6.1f}")

# MSE = Mean Squared Error = average of (actual - predicted) squared.
# Squaring makes every error positive and punishes big misses harder.
total = 0.0
for x, y in zip(hours, scores):
    total += (y - predict(x)) ** 2   # add up each squared error
mse = total / len(hours)             # divide by how many points

print()
print(f"MSE = {mse:.2f}")             # lower MSE = the line fits better
# Expected output: MSE = 1.30  (this line happens to fit quite well)

🎯 Your Turn: build the MSE calculator

# 🎯 YOUR TURN — finish the MSE (Mean Squared Error) calculator
# Fill in each blank marked with ___ . Plain Python, no numpy.

actual    = [10, 20, 30, 40]   # the real values
predicted = [12, 18, 33, 39]   # a model's guesses

total = 0.0
for a, p in zip(actual, predicted):
    # 👉 1) the error is (actual - predicted); square it and add to total
    total += (a - p) ___ 2        # 👉 replace ___ with the power operator (**)

# 👉 2) MSE is the average, so divide the total by the number of points
mse = total / ___                 # 👉 replace ___ with len(actual)

print(f"MSE = {mse:.2f}")
# ✅ Expected output: MSE = 4.50

3 Gradient Descent: How the Model Learns

Guessing w and b by hand doesn't scale. Gradient descent is the algorithm that finds them automatically. Picture the MSE as a valley: every (w, b) is a spot on the hillside, and the lowest point is the best line. Gradient descent is like walking downhill blindfolded — feel which way the ground slopes, take a small step that way, and repeat.

# === Gradient descent: let the computer FIND w and b for us ===
# This is the core loop behind how almost every ML model learns.
# Still plain Python — only the 'lists + arithmetic' you already know.

# Same study-hours data
hours  = [1, 2, 3, 4, 5]
scores = [52, 58, 67, 71, 79]
n = len(hours)

# Start from a deliberately bad guess and let the loop fix it.
w = 0.0
b = 0.0
learning_rate = 0.01   # step size: how big a nudge we take each round

# One "epoch" = one full pass over the data that nudges w and b downhill.
for epoch in range(2000):
    # 1) Predict for every point and collect the errors.
    err_sum   = 0.0   # will become the gradient for b
    err_x_sum = 0.0   # will become the gradient for w
    for x, y in zip(hours, scores):
        yhat = w * x + b      # current prediction
        error = yhat - y      # signed error (predicted - actual)
        err_sum   += error    # accumulate for b's slope
        err_x_sum += error * x  # accumulate for w's slope

    # 2) Gradients tell us which way is "downhill" on the MSE surface.
    grad_w = (2 / n) * err_x_sum
    grad_b = (2 / n) * err_sum

    # 3) Step a little in the opposite (downhill) direction.
    w -= learning_rate * grad_w
    b -= learning_rate * grad_b

# After 2000 small steps, w and b have settled near the best line.
print(f"Learned model: score = {w:.2f} * hours + {b:.2f}")
print(f"Prediction for 6 hours: {w*6 + b:.1f}")
# Expected output (approx):
# Learned model: score = 6.70 * hours + 45.30
# Prediction for 6 hours: 85.5

🎯 Your Turn: complete the update step

# 🎯 YOUR TURN — complete ONE gradient-descent step
# The loop body is written for you except for the two update lines.

xs = [1, 2, 3, 4]
ys = [3, 5, 7, 9]          # the true line is y = 2x + 1
n  = len(xs)

w = 0.0
b = 0.0
learning_rate = 0.05

for epoch in range(1000):
    err_sum = err_x_sum = 0.0
    for x, y in zip(xs, ys):
        error = (w * x + b) - y     # predicted - actual
        err_sum   += error
        err_x_sum += error * x
    grad_w = (2 / n) * err_x_sum
    grad_b = (2 / n) * err_sum

    # 👉 Move w and b DOWNHILL: subtract learning_rate * gradient
    w = w - learning_rate * ___     # 👉 replace ___ with grad_w
    b = b - learning_rate * ___     # 👉 replace ___ with grad_b

print(f"w = {w:.2f}, b = {b:.2f}")
# ✅ Expected output (approx): w = 2.00, b = 1.00

4 Training vs Test: the Honesty Check

A model that scores its own homework will always claim an A. To get an honest measure of how it performs on data it has never seen, you split your data into two parts before training:

If the model does great on the training set but poorly on the test set, it has overfit — it memorised the specific points instead of learning the underlying trend. The test score is the number you actually trust. In real projects, train_test_split from scikit-learn does this split for you in one line.

5 The Real Way: scikit-learn

You now understand what happens under the hood — so in real projects you let a library do it. scikit-learn (imported as sklearn) gives you LinearRegression: create it, call .fit(X, y) to train, and .predict(...) to use it. It runs the maths far faster and more reliably than hand-rolled code.

Read the example below as a worked walkthrough — it needs scikit-learn installed locally to run, and the comments show the expected output. Notice the w and b it finds match what your gradient-descent loop learned above.

# === The real-world way: scikit-learn does the maths for you ===
# In practice you do NOT hand-write gradient descent. You call a library.
# (This needs scikit-learn installed: pip install scikit-learn)
from sklearn.linear_model import LinearRegression

# sklearn expects X as a list of rows; each row is one sample's features.
X = [[1], [2], [3], [4], [5]]   # study hours (one feature per row)
y = [52, 58, 67, 71, 79]        # exam scores

model = LinearRegression()      # create the (untrained) model
model.fit(X, y)                 # .fit() finds the best w and b for you

print(f"slope w     = {model.coef_[0]:.2f}")
print(f"intercept b = {model.intercept_:.2f}")

# .predict() takes the same row shape and returns predictions.
predictions = model.predict([[6], [7]])
print(f"6 hours -> {predictions[0]:.1f}")
print(f"7 hours -> {predictions[1]:.1f}")

# Expected output:
# slope w     = 6.70
# intercept b = 45.30
# 6 hours -> 85.5
# 7 hours -> 92.2

Common Errors (And How to Fix Them)

❌ Not splitting the data

You report 99% accuracy, ship the model, and it fails on real users.

✅ Fix: hold back a test set the model never trains on, and judge it only by the test score. Use train_test_split(X, y, test_size=0.2).

❌ Features on wildly different scales

One feature is in the thousands (square feet) and another is single digits (bedrooms). Gradient descent zig-zags or diverges, and weights become hard to compare.

✅ Fix: scale features to a similar range first (e.g. StandardScaler) before training. A smaller learning rate also helps stop the cost blowing up.

❌ Assuming the relationship is linear

Your data curves, but you force a straight line through it. The MSE stays high no matter how long you train.

✅ Fix: plot the data first. If it bends, add polynomial features, transform a column (e.g. a log), or pick a model that can capture curves.

❌ Extrapolating far beyond the training range

You trained on houses 600–2000 sq ft, then predict a 10,000 sq ft mansion. The line keeps going forever, but reality doesn't.

✅ Fix: only trust predictions inside (or near) the range you trained on. Flag inputs far outside it as unreliable.

📋 Quick Reference

TermFormula / CodeMeaning
Modely = w*x + bA line: weight w, bias b
MSE (cost)mean((y - ŷ)²)How wrong the line is (lower is better)
Gradient stepw -= lr * grad_wNudge parameters downhill
Learning ratelr = 0.01Step size each epoch
Train / test splittrain_test_split(...)Learn on one slice, score on another
sklearn fitmodel.fit(X, y)Finds best w and b for you
sklearn predictmodel.predict(X)Apply the trained line

Mini-Challenge: pick the better line

No blanks this time — just a brief and a comment outline. Write the two helper functions, score both candidate lines with MSE, and print which one fits better. Check your answer against the expected result in the comments.

# 🎯 MINI-CHALLENGE: compare two candidate lines and pick the better one
#
# Data (temperature -> ice creams sold):
#   temps = [15, 18, 20, 25, 30]
#   sales = [22, 33, 40, 62, 81]
#
# 1. Write predict(x, w, b) that returns w * x + b
# 2. Write mse(w, b) that loops over the data and returns the Mean Squared Error
# 3. Score line A: w = 3.0, b = -25
# 4. Score line B: w = 3.0, b = -20
# 5. Print both MSEs and print which line fits better (lower MSE wins)
#
# ✅ Expected: line B has the lower MSE, so it fits better.

temps = [15, 18, 20, 25, 30]
sales = [22, 33, 40, 62, 81]

# your code here

🎉 Lesson Complete!

You've built your first ML model from the ground up: you fit a line y = wx + b, scored it with the MSE cost function, wrote a gradient-descent loop that learns w and b on its own, understood why training and test sets matter, and saw how scikit-learn does it all in three lines.

Practice quiz

In the model y = w*x + b, what does w represent?

  • The intercept (value of y when x is 0)
  • The prediction error
  • The weight or slope: how much y changes per unit of x
  • The learning rate

Answer: The weight or slope: how much y changes per unit of x. w is the weight (slope) — it controls how much the prediction y rises for each unit increase in x.

What does the bias term b represent in y = w*x + b?

  • The value of y when x is 0
  • The slope of the line
  • The mean squared error
  • The number of epochs

Answer: The value of y when x is 0. b is the intercept — the predicted value of y when the input x equals 0.

What does Mean Squared Error (MSE) measure?

  • The sum of all predictions
  • The slope of the line
  • The number of data points
  • The average of (actual - predicted) squared

Answer: The average of (actual - predicted) squared. MSE averages the squared errors; lower MSE means the line fits the data better.

Why are the errors squared in MSE?

  • To make training faster
  • To make every error positive and penalise large misses more
  • To turn the output into a probability
  • To reduce the number of features

Answer: To make every error positive and penalise large misses more. Squaring stops +5 and -5 from cancelling and makes big errors hurt far more than small ones.

What is gradient descent?

  • A loop that nudges parameters downhill on the cost surface to reduce error
  • A way to split data into train and test
  • A method to square the errors
  • A library function that plots data

Answer: A loop that nudges parameters downhill on the cost surface to reduce error. Gradient descent repeatedly steps the parameters in the direction that lowers the cost (downhill).

What does the learning rate control?

  • How many features the model uses
  • The number of test rows
  • The size of each step taken during gradient descent
  • The intercept value

Answer: The size of each step taken during gradient descent. The learning rate is the step size: too large overshoots the minimum, too small makes training crawl.

To move 'downhill' in gradient descent, you update a parameter by...

  • Adding the learning rate times the gradient
  • Subtracting the learning rate times the gradient
  • Multiplying by the gradient
  • Dividing by the gradient

Answer: Subtracting the learning rate times the gradient. You subtract learning_rate * gradient so the parameter moves opposite to the increasing-cost direction.

What is one epoch in training?

  • One data point
  • One prediction
  • One test/train split
  • One full pass over the training data that nudges the parameters

Answer: One full pass over the training data that nudges the parameters. An epoch is a complete pass over the dataset used to update the parameters once (in batch gradient descent).

What does it mean if a model fits the training set well but the test set poorly?

  • It has underfit
  • It has overfit — memorised noise instead of the real pattern
  • It needs a smaller learning rate
  • The MSE is negative

Answer: It has overfit — memorised noise instead of the real pattern. Overfitting means the model memorised the training points and fails to generalise to unseen data.

When is linear regression the wrong choice?

  • When predicting a continuous number
  • When you have a train/test split
  • When the true relationship is not roughly a straight line
  • When you use scikit-learn

Answer: When the true relationship is not roughly a straight line. If the data curves, a straight line fits poorly; you need polynomial features, a transform, or another model.

Continue this course

Frequently asked questions

What does linear regression actually predict?

A continuous number — a value on a sliding scale rather than a category. Think house prices, exam scores, temperatures, or sales figures. If the answer you want is 'how much' or 'how many', linear regression is a sensible first tool. If the answer is a label like 'spam vs not spam' or 'cat vs dog', that is classification instead, which you'll meet in the next lesson.

What is the cost function and why square the errors?

The cost function is a single number that measures how wrong the model is across all the data, and the standard choice for regression is Mean Squared Error (MSE): the average of (actual - predicted) squared. Squaring does two jobs at once — it makes every error positive so they can't cancel out, and it punishes large mistakes far more than small ones (an error of 10 contributes 100, an error of 1 contributes only 1). Training means finding the w and b that make this number as small as possible.

What is gradient descent in one sentence?

It is a loop that repeatedly nudges the model's parameters a tiny step 'downhill' on the cost surface — compute which direction lowers the error (the gradient), take a small step that way (controlled by the learning rate), and repeat until the cost stops shrinking. It is the same core idea that trains neural networks, just with many more parameters.

Why do I split data into training and test sets?

Because a model that has memorised the exact answers it was trained on can look perfect and still fail on anything new. You fit (train) on one slice of the data and check accuracy on a separate test slice the model never saw. If it does well on training data but badly on the test set, it has overfit — it memorised noise instead of learning the real pattern. The test score is your honest estimate of real-world performance.

When is linear regression the wrong choice?

When the real relationship is not roughly a straight line. If sales rise then plateau, or price grows exponentially with size, a straight line will fit poorly no matter how you train it. Plot your data first — if it curves, you need polynomial features, a transform (like taking a log), or a different model entirely. Forcing a line onto a curve gives confident-looking but wrong predictions.