Time Series Forecasting
Reviewed & published by Brayan K
By the end of this lesson you'll break a series into trend, seasonality, and noise, smooth it, build a forecast, split it correctly for time, and measure how good your prediction is.
Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
What You'll Learn in This Lesson
- Break a series into trend, seasonality, and noise
- Smooth data with moving averages and exponential smoothing
- Build lag features so any ML model can forecast
- Split train/test by time — and why you must never shuffle
- Pick a model: naive, ARIMA, exponential smoothing, or LSTM
- Score forecasts with MAE, RMSE, and MAPE
🌦️ Real-World Analogy: Predicting the Weather
Guessing tomorrow's weather is like driving while glancing at the rear-view mirror. You read recent patterns — the last few days were warm, summer is the warm season, today was a fluke cold snap — and project them forward.
Time-series forecasting formalises that intuition. It splits the past into three parts and projects each one ahead:
- • Trend — the slow drift up or down (the climate is warming).
- • Seasonality — patterns that repeat on a fixed cycle (summers are hot every year).
- • Noise — the random wobble you can't predict (an unexpected rainy afternoon).
The crucial difference from ordinary machine learning: observations are not independent. Yesterday's value heavily shapes today's, so order is everything. Lose the order and you lose the signal.
1 The Three Building Blocks: Trend, Seasonality, Noise
Any series can be thought of as those three pieces added together: value = trend + seasonality + noise. Pulling them apart is called decomposition, and it tells you what you can actually predict (trend and seasonality) versus what you can't (noise).
- Trend — is the level generally rising or falling over the whole span?
- Seasonality — is there a repeating shape every 7 days, 12 months, etc.?
- Noise (also called the residual) — whatever's left once you remove trend and seasonality. A good model leaves only noise behind.
If the seasonal swing grows as the level grows (e.g. a busy shop's December spike gets bigger every year), the pieces multiply instead of adding — that's a multiplicative model. Otherwise it's additive.
2 Smoothing: Moving Averages and Exponential Smoothing
Raw data is jagged. Smoothing averages nearby points so the underlying shape shows through. A moving average (MA) replaces each point with the mean of itself and the previous few — a window of, say, 3. A bigger window is smoother but lags further behind turns.
Exponential smoothing is the smarter cousin: instead of treating every value in the window equally, it weights recent values more, with the weight decaying as you go back. One number, alpha (between 0 and 1), controls it — high alpha reacts fast to recent change, low alpha stays calm and slow. The formula is just s = alpha * actual + (1 - alpha) * previous.
The worked example below uses plain Python so you can see exactly how a moving average is computed — then it builds the simplest possible forecast.
# Worked example: moving average + a naive (lag-1) forecast
# Plain Python only — no libraries needed.
sales = [100, 120, 130, 115, 140, 160, 150, 170]
# 1) Moving average smooths out noise so you can see the trend.
# Average each value with the (window-1) values before it.
def moving_average(data, window):
result = []
for i in range(len(data)):
start = max(0, i - window + 1) # clamp so early points don't go negative
chunk = data[start:i + 1] # the window ending at position i
result.append(sum(chunk) / len(chunk))
return result
ma3 = moving_average(sales, 3)
print("Raw :", sales)
print("MA-3:", [round(x, 1) for x in ma3])
# Expected output:
# Raw : [100, 120, 130, 115, 140, 160, 150, 170]
# MA-3: [100.0, 110.0, 116.7, 121.7, 128.3, 138.3, 150.0, 160.0]
# 2) Naive (lag-1) forecast: "tomorrow = today".
# The forecast for each day is simply the previous actual value.
naive = [None] + sales[:-1] # shift the series forward by one step
print("Naive forecast:", naive)
# Expected output:
# Naive forecast: [None, 100, 120, 130, 115, 140, 160, 150]
# 3) Measure the error of that forecast with MAE (mean absolute error).
errors = [abs(sales[i] - naive[i]) for i in range(1, len(sales))]
mae = sum(errors) / len(errors)
print("Per-step errors:", errors)
print("MAE:", round(mae, 2))
# Expected output:
# Per-step errors: [20, 10, 15, 25, 20, 10, 20]
# MAE: 17.14🎯 Your Turn 1: Compute a Moving Average
Fill in the two blanks so the function returns a 3-day moving average. One blank is the divisor that turns a sum into an average; the other is the window size.
# 🎯 YOUR TURN — compute a 3-day moving average over a list
# Fill in each ___ . Run it and check against the expected output.
temps = [18, 20, 22, 19, 24, 26, 23]
def moving_average(data, window):
result = []
for i in range(len(data)):
start = max(0, i - window + 1)
chunk = data[start:i + 1] # the slice ending at i
result.append(sum(chunk) / ___) # 👉 divide by how many items are in 'chunk'
return result
window = ___ # 👉 use a window of 3
ma = moving_average(temps, window)
print([round(x, 1) for x in ma])
# ✅ Expected output:
# [18.0, 19.0, 20.0, 20.3, 21.7, 23.0, 24.3]3 Lag Features: Turning a Series into an ML Table
Ordinary ML models expect a table of rows and feature columns — they don't understand "time". The trick is to build lag features: for each row, the inputs are the previous values (t-1, t-2, t-3…) and the target is the current value. Once your series looks like a normal table, you can throw any model at it — linear regression, random forest, XGBoost.
In pandas, .shift(1) pushes the series forward one step (so each row sees the past), and .rolling(3).mean() adds a rolling average feature. The early rows are NaN until enough history exists.
import pandas as pd
# Lag features: turn a time series into a supervised ML table.
# Each row uses PAST values (t-1, t-2, t-3) to predict the value at t.
s = pd.Series([100, 120, 130, 115, 140, 160, 150, 170], name="sales")
df = pd.DataFrame({"y": s})
df["lag1"] = s.shift(1) # value 1 step ago
df["lag2"] = s.shift(2) # value 2 steps ago
df["roll3"] = s.shift(1).rolling(3).mean() # avg of the 3 prior values
print(df)
# Expected output (first rows have NaN until enough history exists):
# y lag1 lag2 roll3
# 0 100 NaN NaN NaN
# 1 120 100.0 NaN NaN
# 2 130 120.0 100.0 NaN
# 3 115 130.0 120.0 116.666667
# 4 140 115.0 130.0 121.666667
# ...
# IMPORTANT: every feature uses .shift(1) first, so row t never sees y[t].
# That is what prevents leakage from the future.4 Stationarity: Why Classical Models Want a Flat Series
A series is stationary when its statistical behaviour doesn't drift: the mean, the variance, and the way values relate to their neighbours all stay roughly constant over time. A series with a clear upward trend is not stationary — its mean keeps climbing.
Classical models like ARIMA assume stationarity, so you make a series stationary by differencing it: replace each value with the change from the previous value (y[t] - y[t-1]). That removes a linear trend. You can test for stationarity with the Augmented Dickey-Fuller (ADF) test — a p-value below 0.05 means "stationary enough".
from statsmodels.tsa.stattools import adfuller
series = [10, 12, 15, 18, 22, 27, 33, 40] # clear upward trend
print(round(adfuller(series)[1], 3)) # ADF p-value
# Expected output:
# 0.998 <- high p-value => NOT stationary (it has a trend)
# Difference it to remove the trend:
diff = [series[i] - series[i-1] for i in range(1, len(series))]
print(diff)
# Expected output:
# [2, 3, 3, 4, 5, 6, 7] <- the trend is now flatter / more stationary5 Train / Test Split for Time — Never Shuffle!
In normal ML you shuffle your data before splitting. For time series you must not. The whole point is to predict the future from the past, so your test set has to be strictly later in time than your training set. Train on the earliest 80%, evaluate on the most recent 20%.
Shuffling — or using scikit-learn's default train_test_split(..., shuffle=True) — lets the model peek at future values during training. That's look-ahead leakage, and it produces beautiful accuracy that collapses in production.
import pandas as pd
# Train/test split for TIME — split by position, never shuffle.
s = pd.Series(range(1, 13)) # 12 months of data
split = int(len(s) * 0.8) # keep the last 20% as the test set
train = s.iloc[:split] # earliest 80% -> the past
test = s.iloc[split:] # latest 20% -> the future you predict
print("train:", list(train))
print("test :", list(test))
# Expected output:
# train: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]
# test : [11, 12]
# Note: NO shuffle=True, NO random split. The test set is strictly LATER
# in time than the training set, exactly like real forecasting.🎯 Your Turn 2: Naive Forecast + MAE
Build a 1-step naive forecast (tomorrow = today) by shifting the series, then measure its Mean Absolute Error. Fill in the two blanks and check the expected output.
# 🎯 YOUR TURN — make a 1-step naive forecast and measure its error
# The naive forecast says "tomorrow equals today": forecast[i] = actual[i-1].
actual = [50, 54, 53, 58, 60, 57]
# 1) Build the naive forecast by shifting the series one step forward.
forecast = [None] + actual[___] # 👉 take every value EXCEPT the last one
# 2) Mean Absolute Error: average of |actual - forecast|, skipping index 0.
diffs = [abs(actual[i] - forecast[i]) for i in range(1, len(actual))]
mae = sum(diffs) / ___ # 👉 divide by how many diffs there are: len(diffs)
print("Forecast:", forecast)
print("MAE:", round(mae, 2))
# ✅ Expected output:
# Forecast: [None, 50, 54, 53, 58, 60]
# MAE: 3.06 Classical Models: ARIMA & Exponential Smoothing
ARIMA(p, d, q) combines three ideas: AR (predict from p past values), I (difference the series d times to make it stationary), and MA (correct using q past forecast errors). It's a strong, interpretable baseline for trended data.
from statsmodels.tsa.arima.model import ARIMA
# Classical model: ARIMA(p, d, q) on a stationary-ish series.
# p = how many lagged values (AR), d = how many times to difference
# to remove trend, q = how many lagged errors (MA).
series = [112, 118, 132, 129, 121, 135, 148, 148, 136, 119, 104, 118]
model = ARIMA(series, order=(1, 1, 1)) # 1 AR term, 1 difference, 1 MA term
fit = model.fit()
forecast = fit.forecast(steps=3) # predict the next 3 steps
print(forecast.round(1))
# Expected output (values approximate):
# [124.8 122.3 123.0]
# d=1 differences the series once to remove the trend (stationarity).
# Check stationarity first with the ADF test:
# from statsmodels.tsa.stattools import adfuller
# adfuller(series)[1] # p-value < 0.05 => already stationaryExponential smoothing (in its full Holt-Winters form) tracks a smoothed level, a trend, and a seasonal component, weighting recent data most. It's often the easiest way to get a solid seasonal forecast.
from statsmodels.tsa.holtwinters import ExponentialSmoothing
# Exponential smoothing weights recent points more than old ones.
# Holt-Winters adds a trend and a seasonal component on top.
monthly = [266, 146, 183, 119, 180, 169, 232, 225, 193, 123, 337, 186,
194, 150, 210, 273, 191, 287, 226, 304, 290, 422, 265, 342]
model = ExponentialSmoothing(
monthly,
trend="add", # add a linear trend
seasonal="add", # add a seasonal swing
seasonal_periods=12, # 12 months per cycle
)
fit = model.fit()
print(fit.forecast(3).round(0))
# Expected output (values approximate):
# [320. 240. 300.]
# 'add' = additive; use 'mul' when the seasonal swing grows with the level.🤖 ML and Deep Learning: When to Reach for an LSTM
Two broad routes go beyond classical models:
- Lag features + a tree model (XGBoost, LightGBM, random forest). Build the lag table from Section 3, add calendar features (day_of_week, month, holiday flags), and train any tabular model. This is the workhorse for most real forecasting.
- LSTM / GRU recurrent neural networks. An LSTM (Long Short-Term Memory) reads the sequence step by step and keeps a memory of what it has seen, so it can learn long, complex patterns directly — no manual lag engineering. The cost: it needs lots of data (thousands of points), careful scaling, and far more tuning.
7 Evaluating a Forecast: MAE, RMSE, MAPE
A forecast is only as good as the number you score it with. Three metrics cover most cases, all computed on the test set (the future the model didn't see):
- MAE (Mean Absolute Error) — average of |actual - forecast|. Same units as the data, easy to explain.
- RMSE (Root Mean Squared Error) — square the errors, average, square-root. Punishes big misses harder than MAE.
- MAPE (Mean Absolute Percentage Error) — the average error as a percentage of the actual value. Comparable across different series, but unstable when actuals are near zero.
actual = [100, 110, 120, 130]
forecast = [ 98, 115, 118, 125]
errors = [actual[i] - forecast[i] for i in range(len(actual))]
mae = sum(abs(e) for e in errors) / len(errors)
rmse = (sum(e * e for e in errors) / len(errors)) ** 0.5
mape = sum(abs(errors[i]) / actual[i] for i in range(len(actual))) / len(actual) * 100
print("MAE :", round(mae, 2))
print("RMSE:", round(rmse, 2))
print("MAPE:", round(mape, 2), "%")
# Expected output:
# MAE : 3.5
# RMSE: 3.81
# MAPE: 3.01 %8 Common Errors (And How to Fix Them)
These four mistakes ruin more forecasts than any modelling choice:
❌ Shuffling time data before the split
A random split scatters future rows into the training set:
# ❌ Wrong — order destroyed, future leaks into training
from sklearn.model_selection import train_test_split
X_tr, X_te, y_tr, y_te = train_test_split(X, y) # shuffle=True by default!✅ Fix: split by position, keeping time order:
split = int(len(X) * 0.8)
X_tr, X_te = X[:split], X[split:] # train = past, test = future❌ Data leakage from the future
Building a feature that includes the value you're predicting (or scaling using the whole dataset):
df["roll3"] = df["y"].rolling(3).mean() # ❌ includes today's y!✅ Fix: shift first so every feature is strictly past:
df["roll3"] = df["y"].shift(1).rolling(3).mean() # ✅ only past valuesA plain trend model misses the weekly/yearly cycle and forecasts a smooth line through every peak and dip.
✅ Fix: model the season — use a seasonal-naive baseline, Holt-Winters with seasonal_periods, or add a month/day_of_week feature.
❌ Fitting ARIMA to a non-stationary series
A strong trend with d=0 gives wild, drifting forecasts.
✅ Fix: difference the series (raise d) until the ADF p-value drops below 0.05, then fit.
📋 Quick Reference
| Method | Best For | Limitation |
|---|---|---|
| Naive / seasonal-naive | The baseline you must beat | Ignores trend & structure |
| Moving average | Visualising trend, smoothing | Lags behind, no forecast |
| Exponential smoothing | Short-term, seasonal forecasts | No external features |
| ARIMA | Trended, near-stationary data | Needs differencing / tuning |
| Lag features + ML | Adding calendar / external data | Feature-engineering effort |
| LSTM / GRU | Long, complex sequences | Needs lots of data |
| Metric | Meaning | Use when |
|---|---|---|
| MAE | Avg absolute error (data units) | You want a plain, robust number |
| RMSE | Like MAE but penalises big misses | Large errors are costly |
| MAPE | Avg error as a % | Comparing across series (actuals not near 0) |
🎯 Mini-Challenge: Seasonal-Naive Forecast + RMSE
Support is faded now — only a comment outline is given. Write a seasonal-naive forecast (this period equals the value one full season ago) and score it with RMSE. The expected answer is in the comments so you can self-check.
# 🎯 MINI-CHALLENGE: seasonal-naive forecast + RMSE
#
# A "seasonal-naive" forecast for weekly data says: this week ≈ the same
# week one season ago. With a season length of 3, forecast[i] = actual[i-3].
#
# Steps:
# 1. Given: actual = [10, 12, 14, 11, 13, 16, 12, 15, 18] (season = 3)
# 2. Build forecast: for each i >= 3, forecast[i] = actual[i - 3]
# (the first 3 entries have no forecast — skip them).
# 3. Compute RMSE over the forecastable points:
# - square each (actual[i] - forecast[i])
# - take the mean of those squares
# - take the square root (use value ** 0.5)
# 4. print the RMSE rounded to 2 decimals.
#
# ✅ Expected output (season = 3):
# RMSE: 1.58
# your code hereLesson complete — you can forecast the future!
You can decompose a series into trend, seasonality, and noise; smooth it with moving averages and exponential smoothing; build lag features; split correctly by time without leakage; choose between ARIMA, exponential smoothing, and an LSTM; and score the result with MAE, RMSE, and MAPE.
Practice quiz
What three components is a time series commonly decomposed into?
- Mean, median, and mode
- Input, hidden, and output
- Trend, seasonality, and noise
- Train, validation, and test
Answer: Trend, seasonality, and noise. Decomposition splits a series into trend (slow drift), seasonality (repeating cycle), and noise (the unpredictable residual).
Why must you NOT shuffle time-series data before splitting train/test?
- Shuffling lets the model see future values during training, causing look-ahead leakage
- Shuffling is too slow on large datasets
- Shuffling changes the data types
- Shuffling is fine for time series
Answer: Shuffling lets the model see future values during training, causing look-ahead leakage. Order carries the signal; shuffling leaks future information into training, giving fake-high accuracy that collapses in production. Split by time instead.
What does it mean for a series to be stationary?
- It never changes value
- It has exactly one season per year
- It contains no noise
- Its mean, variance, and autocorrelation stay roughly constant over time
Answer: Its mean, variance, and autocorrelation stay roughly constant over time. A stationary series has constant statistical behaviour over time — no trend or changing seasonality. ARIMA assumes stationarity.
How do you typically make a trended series stationary for ARIMA?
- Multiply every value by a constant
- Difference it: replace each value with the change from the previous value
- Shuffle the values randomly
- Take the logarithm of the index
Answer: Difference it: replace each value with the change from the previous value. Differencing (y[t] - y[t-1]) removes a linear trend; the 'I' (integrated) term in ARIMA(p, d, q) controls how many times you difference.
What do the p, d, and q in ARIMA(p, d, q) stand for?
- AR lagged values, number of differences, and MA lagged errors
- Precision, depth, and quality
- Periods, days, and quarters
- Probability, distance, and quantile
Answer: AR lagged values, number of differences, and MA lagged errors. p = autoregressive lagged values, d = times to difference for stationarity, q = moving-average lagged forecast errors.
How does exponential smoothing differ from a simple moving average?
- It weights every value in the window equally
- It ignores recent values
- It weights recent values more heavily, controlled by alpha
- It only works on stationary series
Answer: It weights recent values more heavily, controlled by alpha. A moving average weights all window values equally and lags; exponential smoothing weights recent values more (via alpha), reacting faster.
What is the naive (lag-1) forecast?
- The average of all past values
- Tomorrow's prediction equals today's actual value
- A forecast from a neural network
- The median of the test set
Answer: Tomorrow's prediction equals today's actual value. The naive forecast simply says 'tomorrow = today' (forecast[t] = actual[t-1]). It is the baseline any real model must beat.
What is the key difference between RMSE and MAE?
- RMSE is always negative
- MAE can only be used on percentages
- They are identical
- RMSE squares errors so it punishes large misses more than MAE
Answer: RMSE squares errors so it punishes large misses more than MAE. MAE is the mean absolute error; RMSE squares the errors before averaging and rooting, so big misses are penalised more heavily.
When is MAPE (Mean Absolute Percentage Error) problematic?
- When the data has a trend
- When actual values are near zero, because the percentage blows up
- When you compare across different series
- When the series is stationary
Answer: When actual values are near zero, because the percentage blows up. MAPE expresses error as a percentage of the actual value, so it becomes unstable or huge when actuals are close to zero.
When does reaching for an LSTM pay off over ARIMA or exponential smoothing?
- Always — LSTMs beat classical models on every dataset
- Only on stationary series with no seasonality
- Only with long, complex sequences and thousands of observations
- When you have fewer than 50 data points
Answer: Only with long, complex sequences and thousands of observations. On small business datasets a naive baseline plus ARIMA or smoothing usually wins; LSTMs only pay off with long, complex sequences and lots of data.
Continue this course
- Previous: Model Selection & Tuning
- Next: Gradient Boosting (XGBoost & LightGBM) — Sequential weak learners, XGBoost/LightGBM and key hyperparameters
- Quick reference: AI & Machine Learning cheat sheet
Frequently asked questions
What is time-series forecasting?
It is predicting future values of a quantity that is recorded in time order — sales, temperature, traffic — by learning patterns (trend, seasonality, and how each value depends on recent values) from past observations.
Why can't I shuffle time-series data before splitting train/test?
Because order carries the signal. Shuffling lets the model 'see the future' during training, which leaks information and gives fake-high accuracy. Always split by time: train on the earliest part, test on the most recent part.
What is stationarity and why does it matter?
A stationary series has a constant mean, variance, and autocorrelation over time — no trend or changing seasonality. Classical models like ARIMA assume stationarity, so you often difference the series (subtract the previous value) to remove trend before modelling.
Should I use a moving average or exponential smoothing?
A moving average weights every value in its window equally and lags behind turns. Exponential smoothing weights recent values more (controlled by alpha), so it reacts faster. Use a moving average to visualise trend; use exponential smoothing when you need a short forecast.
Do I need an LSTM, or is ARIMA enough?
Start simple. A naive or seasonal-naive baseline plus ARIMA or exponential smoothing beats a neural network on most small business datasets. LSTMs only pay off with long, complex sequences and thousands of observations.
Which error metric should I report — MAE, RMSE, or MAPE?
MAE is the average absolute error in the original units. RMSE is similar but punishes large misses more. MAPE expresses error as a percentage, which is comparable across series but blows up when actual values are near zero. Report MAE/RMSE for one series, MAPE to compare across series.