Time Series Forecasting

Reviewed & published by Brayan K

By the end of this lesson you'll break a series into trend, seasonality, and noise, smooth it, build a forecast, split it correctly for time, and measure how good your prediction is.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

🌦️ Real-World Analogy: Predicting the Weather

Guessing tomorrow's weather is like driving while glancing at the rear-view mirror. You read recent patterns — the last few days were warm, summer is the warm season, today was a fluke cold snap — and project them forward.

Time-series forecasting formalises that intuition. It splits the past into three parts and projects each one ahead:

The crucial difference from ordinary machine learning: observations are not independent. Yesterday's value heavily shapes today's, so order is everything. Lose the order and you lose the signal.

1 The Three Building Blocks: Trend, Seasonality, Noise

Any series can be thought of as those three pieces added together: value = trend + seasonality + noise. Pulling them apart is called decomposition, and it tells you what you can actually predict (trend and seasonality) versus what you can't (noise).

If the seasonal swing grows as the level grows (e.g. a busy shop's December spike gets bigger every year), the pieces multiply instead of adding — that's a multiplicative model. Otherwise it's additive.

2 Smoothing: Moving Averages and Exponential Smoothing

Raw data is jagged. Smoothing averages nearby points so the underlying shape shows through. A moving average (MA) replaces each point with the mean of itself and the previous few — a window of, say, 3. A bigger window is smoother but lags further behind turns.

Exponential smoothing is the smarter cousin: instead of treating every value in the window equally, it weights recent values more, with the weight decaying as you go back. One number, alpha (between 0 and 1), controls it — high alpha reacts fast to recent change, low alpha stays calm and slow. The formula is just s = alpha * actual + (1 - alpha) * previous.

The worked example below uses plain Python so you can see exactly how a moving average is computed — then it builds the simplest possible forecast.

# Worked example: moving average + a naive (lag-1) forecast
# Plain Python only — no libraries needed.

sales = [100, 120, 130, 115, 140, 160, 150, 170]

# 1) Moving average smooths out noise so you can see the trend.
#    Average each value with the (window-1) values before it.
def moving_average(data, window):
    result = []
    for i in range(len(data)):
        start = max(0, i - window + 1)   # clamp so early points don't go negative
        chunk = data[start:i + 1]        # the window ending at position i
        result.append(sum(chunk) / len(chunk))
    return result

ma3 = moving_average(sales, 3)
print("Raw :", sales)
print("MA-3:", [round(x, 1) for x in ma3])
# Expected output:
# Raw : [100, 120, 130, 115, 140, 160, 150, 170]
# MA-3: [100.0, 110.0, 116.7, 121.7, 128.3, 138.3, 150.0, 160.0]

# 2) Naive (lag-1) forecast: "tomorrow = today".
#    The forecast for each day is simply the previous actual value.
naive = [None] + sales[:-1]   # shift the series forward by one step
print("Naive forecast:", naive)
# Expected output:
# Naive forecast: [None, 100, 120, 130, 115, 140, 160, 150]

# 3) Measure the error of that forecast with MAE (mean absolute error).
errors = [abs(sales[i] - naive[i]) for i in range(1, len(sales))]
mae = sum(errors) / len(errors)
print("Per-step errors:", errors)
print("MAE:", round(mae, 2))
# Expected output:
# Per-step errors: [20, 10, 15, 25, 20, 10, 20]
# MAE: 17.14

🎯 Your Turn 1: Compute a Moving Average

Fill in the two blanks so the function returns a 3-day moving average. One blank is the divisor that turns a sum into an average; the other is the window size.

# 🎯 YOUR TURN — compute a 3-day moving average over a list
# Fill in each ___ . Run it and check against the expected output.

temps = [18, 20, 22, 19, 24, 26, 23]

def moving_average(data, window):
    result = []
    for i in range(len(data)):
        start = max(0, i - window + 1)
        chunk = data[start:i + 1]           # the slice ending at i
        result.append(sum(chunk) / ___)     # 👉 divide by how many items are in 'chunk'
    return result

window = ___                                 # 👉 use a window of 3
ma = moving_average(temps, window)
print([round(x, 1) for x in ma])

# ✅ Expected output:
# [18.0, 19.0, 20.0, 20.3, 21.7, 23.0, 24.3]

3 Lag Features: Turning a Series into an ML Table

Ordinary ML models expect a table of rows and feature columns — they don't understand "time". The trick is to build lag features: for each row, the inputs are the previous values (t-1, t-2, t-3…) and the target is the current value. Once your series looks like a normal table, you can throw any model at it — linear regression, random forest, XGBoost.

In pandas, .shift(1) pushes the series forward one step (so each row sees the past), and .rolling(3).mean() adds a rolling average feature. The early rows are NaN until enough history exists.

import pandas as pd

# Lag features: turn a time series into a supervised ML table.
# Each row uses PAST values (t-1, t-2, t-3) to predict the value at t.
s = pd.Series([100, 120, 130, 115, 140, 160, 150, 170], name="sales")

df = pd.DataFrame({"y": s})
df["lag1"] = s.shift(1)          # value 1 step ago
df["lag2"] = s.shift(2)          # value 2 steps ago
df["roll3"] = s.shift(1).rolling(3).mean()  # avg of the 3 prior values
print(df)

# Expected output (first rows have NaN until enough history exists):
#        y   lag1   lag2   roll3
# 0    100    NaN    NaN     NaN
# 1    120  100.0    NaN     NaN
# 2    130  120.0  100.0     NaN
# 3    115  130.0  120.0  116.666667
# 4    140  115.0  130.0  121.666667
# ...
# IMPORTANT: every feature uses .shift(1) first, so row t never sees y[t].
# That is what prevents leakage from the future.

4 Stationarity: Why Classical Models Want a Flat Series

A series is stationary when its statistical behaviour doesn't drift: the mean, the variance, and the way values relate to their neighbours all stay roughly constant over time. A series with a clear upward trend is not stationary — its mean keeps climbing.

Classical models like ARIMA assume stationarity, so you make a series stationary by differencing it: replace each value with the change from the previous value (y[t] - y[t-1]). That removes a linear trend. You can test for stationarity with the Augmented Dickey-Fuller (ADF) test — a p-value below 0.05 means "stationary enough".

from statsmodels.tsa.stattools import adfuller

series = [10, 12, 15, 18, 22, 27, 33, 40]   # clear upward trend
print(round(adfuller(series)[1], 3))   # ADF p-value
# Expected output:
# 0.998   <- high p-value => NOT stationary (it has a trend)

# Difference it to remove the trend:
diff = [series[i] - series[i-1] for i in range(1, len(series))]
print(diff)
# Expected output:
# [2, 3, 3, 4, 5, 6, 7]   <- the trend is now flatter / more stationary

5 Train / Test Split for Time — Never Shuffle!

In normal ML you shuffle your data before splitting. For time series you must not. The whole point is to predict the future from the past, so your test set has to be strictly later in time than your training set. Train on the earliest 80%, evaluate on the most recent 20%.

Shuffling — or using scikit-learn's default train_test_split(..., shuffle=True) — lets the model peek at future values during training. That's look-ahead leakage, and it produces beautiful accuracy that collapses in production.

import pandas as pd

# Train/test split for TIME — split by position, never shuffle.
s = pd.Series(range(1, 13))          # 12 months of data

split = int(len(s) * 0.8)            # keep the last 20% as the test set
train = s.iloc[:split]               # earliest 80%  -> the past
test  = s.iloc[split:]               # latest  20%  -> the future you predict

print("train:", list(train))
print("test :", list(test))
# Expected output:
# train: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]
# test : [11, 12]
# Note: NO shuffle=True, NO random split. The test set is strictly LATER
# in time than the training set, exactly like real forecasting.

🎯 Your Turn 2: Naive Forecast + MAE

Build a 1-step naive forecast (tomorrow = today) by shifting the series, then measure its Mean Absolute Error. Fill in the two blanks and check the expected output.

# 🎯 YOUR TURN — make a 1-step naive forecast and measure its error
# The naive forecast says "tomorrow equals today": forecast[i] = actual[i-1].

actual = [50, 54, 53, 58, 60, 57]

# 1) Build the naive forecast by shifting the series one step forward.
forecast = [None] + actual[___]    # 👉 take every value EXCEPT the last one

# 2) Mean Absolute Error: average of |actual - forecast|, skipping index 0.
diffs = [abs(actual[i] - forecast[i]) for i in range(1, len(actual))]
mae = sum(diffs) / ___              # 👉 divide by how many diffs there are: len(diffs)

print("Forecast:", forecast)
print("MAE:", round(mae, 2))

# ✅ Expected output:
# Forecast: [None, 50, 54, 53, 58, 60]
# MAE: 3.0

6 Classical Models: ARIMA & Exponential Smoothing

ARIMA(p, d, q) combines three ideas: AR (predict from p past values), I (difference the series d times to make it stationary), and MA (correct using q past forecast errors). It's a strong, interpretable baseline for trended data.

from statsmodels.tsa.arima.model import ARIMA

# Classical model: ARIMA(p, d, q) on a stationary-ish series.
#   p = how many lagged values (AR), d = how many times to difference
#   to remove trend, q = how many lagged errors (MA).
series = [112, 118, 132, 129, 121, 135, 148, 148, 136, 119, 104, 118]

model = ARIMA(series, order=(1, 1, 1))   # 1 AR term, 1 difference, 1 MA term
fit = model.fit()
forecast = fit.forecast(steps=3)         # predict the next 3 steps
print(forecast.round(1))
# Expected output (values approximate):
# [124.8 122.3 123.0]
# d=1 differences the series once to remove the trend (stationarity).
# Check stationarity first with the ADF test:
#   from statsmodels.tsa.stattools import adfuller
#   adfuller(series)[1]   # p-value < 0.05 => already stationary

Exponential smoothing (in its full Holt-Winters form) tracks a smoothed level, a trend, and a seasonal component, weighting recent data most. It's often the easiest way to get a solid seasonal forecast.

from statsmodels.tsa.holtwinters import ExponentialSmoothing

# Exponential smoothing weights recent points more than old ones.
# Holt-Winters adds a trend and a seasonal component on top.
monthly = [266, 146, 183, 119, 180, 169, 232, 225, 193, 123, 337, 186,
           194, 150, 210, 273, 191, 287, 226, 304, 290, 422, 265, 342]

model = ExponentialSmoothing(
    monthly,
    trend="add",            # add a linear trend
    seasonal="add",         # add a seasonal swing
    seasonal_periods=12,    # 12 months per cycle
)
fit = model.fit()
print(fit.forecast(3).round(0))
# Expected output (values approximate):
# [320. 240. 300.]
# 'add' = additive; use 'mul' when the seasonal swing grows with the level.

🤖 ML and Deep Learning: When to Reach for an LSTM

Two broad routes go beyond classical models:

7 Evaluating a Forecast: MAE, RMSE, MAPE

A forecast is only as good as the number you score it with. Three metrics cover most cases, all computed on the test set (the future the model didn't see):

actual   = [100, 110, 120, 130]
forecast = [ 98, 115, 118, 125]

errors = [actual[i] - forecast[i] for i in range(len(actual))]

mae  = sum(abs(e) for e in errors) / len(errors)
rmse = (sum(e * e for e in errors) / len(errors)) ** 0.5
mape = sum(abs(errors[i]) / actual[i] for i in range(len(actual))) / len(actual) * 100

print("MAE :", round(mae, 2))
print("RMSE:", round(rmse, 2))
print("MAPE:", round(mape, 2), "%")
# Expected output:
# MAE : 3.5
# RMSE: 3.81
# MAPE: 3.01 %

8 Common Errors (And How to Fix Them)

These four mistakes ruin more forecasts than any modelling choice:

❌ Shuffling time data before the split

A random split scatters future rows into the training set:

# ❌ Wrong — order destroyed, future leaks into training
from sklearn.model_selection import train_test_split
X_tr, X_te, y_tr, y_te = train_test_split(X, y)   # shuffle=True by default!

✅ Fix: split by position, keeping time order:

split = int(len(X) * 0.8)
X_tr, X_te = X[:split], X[split:]   # train = past, test = future

❌ Data leakage from the future

Building a feature that includes the value you're predicting (or scaling using the whole dataset):

df["roll3"] = df["y"].rolling(3).mean()   # ❌ includes today's y!

✅ Fix: shift first so every feature is strictly past:

df["roll3"] = df["y"].shift(1).rolling(3).mean()   # ✅ only past values

A plain trend model misses the weekly/yearly cycle and forecasts a smooth line through every peak and dip.

✅ Fix: model the season — use a seasonal-naive baseline, Holt-Winters with seasonal_periods, or add a month/day_of_week feature.

❌ Fitting ARIMA to a non-stationary series

A strong trend with d=0 gives wild, drifting forecasts.

✅ Fix: difference the series (raise d) until the ADF p-value drops below 0.05, then fit.

📋 Quick Reference

MethodBest ForLimitation
Naive / seasonal-naiveThe baseline you must beatIgnores trend & structure
Moving averageVisualising trend, smoothingLags behind, no forecast
Exponential smoothingShort-term, seasonal forecastsNo external features
ARIMATrended, near-stationary dataNeeds differencing / tuning
Lag features + MLAdding calendar / external dataFeature-engineering effort
LSTM / GRULong, complex sequencesNeeds lots of data
MetricMeaningUse when
MAEAvg absolute error (data units)You want a plain, robust number
RMSELike MAE but penalises big missesLarge errors are costly
MAPEAvg error as a %Comparing across series (actuals not near 0)

🎯 Mini-Challenge: Seasonal-Naive Forecast + RMSE

Support is faded now — only a comment outline is given. Write a seasonal-naive forecast (this period equals the value one full season ago) and score it with RMSE. The expected answer is in the comments so you can self-check.

# 🎯 MINI-CHALLENGE: seasonal-naive forecast + RMSE
#
# A "seasonal-naive" forecast for weekly data says: this week ≈ the same
# week one season ago. With a season length of 3, forecast[i] = actual[i-3].
#
# Steps:
#   1. Given:  actual = [10, 12, 14, 11, 13, 16, 12, 15, 18]   (season = 3)
#   2. Build forecast: for each i >= 3, forecast[i] = actual[i - 3]
#      (the first 3 entries have no forecast — skip them).
#   3. Compute RMSE over the forecastable points:
#        - square each (actual[i] - forecast[i])
#        - take the mean of those squares
#        - take the square root  (use value ** 0.5)
#   4. print the RMSE rounded to 2 decimals.
#
# ✅ Expected output (season = 3):
# RMSE: 1.58

# your code here

Lesson complete — you can forecast the future!

You can decompose a series into trend, seasonality, and noise; smooth it with moving averages and exponential smoothing; build lag features; split correctly by time without leakage; choose between ARIMA, exponential smoothing, and an LSTM; and score the result with MAE, RMSE, and MAPE.

Practice quiz

What three components is a time series commonly decomposed into?

  • Mean, median, and mode
  • Input, hidden, and output
  • Trend, seasonality, and noise
  • Train, validation, and test

Answer: Trend, seasonality, and noise. Decomposition splits a series into trend (slow drift), seasonality (repeating cycle), and noise (the unpredictable residual).

Why must you NOT shuffle time-series data before splitting train/test?

  • Shuffling lets the model see future values during training, causing look-ahead leakage
  • Shuffling is too slow on large datasets
  • Shuffling changes the data types
  • Shuffling is fine for time series

Answer: Shuffling lets the model see future values during training, causing look-ahead leakage. Order carries the signal; shuffling leaks future information into training, giving fake-high accuracy that collapses in production. Split by time instead.

What does it mean for a series to be stationary?

  • It never changes value
  • It has exactly one season per year
  • It contains no noise
  • Its mean, variance, and autocorrelation stay roughly constant over time

Answer: Its mean, variance, and autocorrelation stay roughly constant over time. A stationary series has constant statistical behaviour over time — no trend or changing seasonality. ARIMA assumes stationarity.

How do you typically make a trended series stationary for ARIMA?

  • Multiply every value by a constant
  • Difference it: replace each value with the change from the previous value
  • Shuffle the values randomly
  • Take the logarithm of the index

Answer: Difference it: replace each value with the change from the previous value. Differencing (y[t] - y[t-1]) removes a linear trend; the 'I' (integrated) term in ARIMA(p, d, q) controls how many times you difference.

What do the p, d, and q in ARIMA(p, d, q) stand for?

  • AR lagged values, number of differences, and MA lagged errors
  • Precision, depth, and quality
  • Periods, days, and quarters
  • Probability, distance, and quantile

Answer: AR lagged values, number of differences, and MA lagged errors. p = autoregressive lagged values, d = times to difference for stationarity, q = moving-average lagged forecast errors.

How does exponential smoothing differ from a simple moving average?

  • It weights every value in the window equally
  • It ignores recent values
  • It weights recent values more heavily, controlled by alpha
  • It only works on stationary series

Answer: It weights recent values more heavily, controlled by alpha. A moving average weights all window values equally and lags; exponential smoothing weights recent values more (via alpha), reacting faster.

What is the naive (lag-1) forecast?

  • The average of all past values
  • Tomorrow's prediction equals today's actual value
  • A forecast from a neural network
  • The median of the test set

Answer: Tomorrow's prediction equals today's actual value. The naive forecast simply says 'tomorrow = today' (forecast[t] = actual[t-1]). It is the baseline any real model must beat.

What is the key difference between RMSE and MAE?

  • RMSE is always negative
  • MAE can only be used on percentages
  • They are identical
  • RMSE squares errors so it punishes large misses more than MAE

Answer: RMSE squares errors so it punishes large misses more than MAE. MAE is the mean absolute error; RMSE squares the errors before averaging and rooting, so big misses are penalised more heavily.

When is MAPE (Mean Absolute Percentage Error) problematic?

  • When the data has a trend
  • When actual values are near zero, because the percentage blows up
  • When you compare across different series
  • When the series is stationary

Answer: When actual values are near zero, because the percentage blows up. MAPE expresses error as a percentage of the actual value, so it becomes unstable or huge when actuals are close to zero.

When does reaching for an LSTM pay off over ARIMA or exponential smoothing?

  • Always — LSTMs beat classical models on every dataset
  • Only on stationary series with no seasonality
  • Only with long, complex sequences and thousands of observations
  • When you have fewer than 50 data points

Answer: Only with long, complex sequences and thousands of observations. On small business datasets a naive baseline plus ARIMA or smoothing usually wins; LSTMs only pay off with long, complex sequences and lots of data.

Continue this course

Frequently asked questions

What is time-series forecasting?

It is predicting future values of a quantity that is recorded in time order — sales, temperature, traffic — by learning patterns (trend, seasonality, and how each value depends on recent values) from past observations.

Why can't I shuffle time-series data before splitting train/test?

Because order carries the signal. Shuffling lets the model 'see the future' during training, which leaks information and gives fake-high accuracy. Always split by time: train on the earliest part, test on the most recent part.

What is stationarity and why does it matter?

A stationary series has a constant mean, variance, and autocorrelation over time — no trend or changing seasonality. Classical models like ARIMA assume stationarity, so you often difference the series (subtract the previous value) to remove trend before modelling.

Should I use a moving average or exponential smoothing?

A moving average weights every value in its window equally and lags behind turns. Exponential smoothing weights recent values more (controlled by alpha), so it reacts faster. Use a moving average to visualise trend; use exponential smoothing when you need a short forecast.

Do I need an LSTM, or is ARIMA enough?

Start simple. A naive or seasonal-naive baseline plus ARIMA or exponential smoothing beats a neural network on most small business datasets. LSTMs only pay off with long, complex sequences and thousands of observations.

Which error metric should I report — MAE, RMSE, or MAPE?

MAE is the average absolute error in the original units. RMSE is similar but punishes large misses more. MAPE expresses error as a percentage, which is comparable across series but blows up when actual values are near zero. Report MAE/RMSE for one series, MAPE to compare across series.