Dimensionality Reduction
Reviewed & published by Brayan K
Squeeze data with hundreds of features down to a handful you can plot and model — using mean-centring, PCA, and t-SNE/UMAP — without throwing away the signal.
Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
What You'll Learn in This Lesson
- Why the curse of dimensionality breaks distance-based models
- How PCA finds principal components — the directions of most variance
- How to read explained variance to choose how many components to keep
- When to reach for t-SNE or UMAP to visualise clusters in 2D
- The difference between feature selection and feature extraction
- Why you must scale data before PCA, and how to avoid data leakage
🌍 Real-World Analogy: A Shadow on the Wall
Hold up your hand and shine a torch at it. The shadow on the wall is your 3D hand flattened to 2D. You lose depth, but if you turn your hand to a good angle, the shadow still tells you it is a hand — five fingers, a thumb, the shape is all there.
Dimensionality reduction is choosing the best angle to cast the shadow. Your data might live in 100 dimensions, but a well-chosen 2D shadow can keep the parts that matter — the clusters, the trends — and drop the rest. PCA picks the angle that keeps the shadow as spread out as possible (maximum variance), because a spread-out shadow preserves the most information. A summary of a long book keeps the plot and loses the filler; a good reduction keeps the structure and loses the noise.
1 The Curse of Dimensionality
A dimension is just a feature — one column in your dataset. Add more columns and you add more dimensions. That sounds harmless, but space gets enormous fast. To cover a 1D line with points spaced 0.1 apart you need 10 points; a 2D square needs 100; a 10-dimensional cube needs 10 billion. Your dataset of a few thousand rows is a tiny sprinkle in that vastness.
The painful consequence: in high dimensions, every point ends up roughly the same distance from every other point. Anything that relies on "near" versus "far" — k-nearest-neighbours, k-means, clustering, many distance kernels — starts to fall apart because "near" stops meaning anything. Models also need exponentially more data to fill the space, so they overfit.
2 PCA — Variance and Principal Components
PCA (Principal Component Analysis) is the workhorse of reduction. The idea in one line: find the direction in which your data is most spread out, call it the first principal component, then the next-most-spread direction at right angles to it, and so on.
Why chase spread? Because variance is information. A direction where every point has the same value tells you nothing; a direction where points spread far apart separates them. The first principal component is the single axis that captures the most variance — the best 1D "shadow" of your data. The second captures the most of what is left, and is perpendicular to the first.
The recipe is exactly the worked example below: mean-centre the data (so variance is measured from the centre), then project each point onto the chosen axis with a dot product. Project onto the top one or two components and you have reduced the data while keeping the most spread.
📐 The two operations that make PCA work:
- Mean-centre: subtract each column's mean so the cloud sits at the origin.
- Project: take the dot product of each point with a unit axis to collapse it to one number.
Worked example. This is PCA's engine in plain Python — no libraries. Read every comment, then run it. It mean-centres a small 2D dataset and projects the points onto a single axis.
# Dimensionality reduction by hand: mean-centre, then PROJECT onto one axis.
# No libraries needed — this is the whole idea of PCA in miniature.
# A tiny 2D dataset: each row is a point (x, y).
data = [
[2.0, 1.0],
[3.0, 2.5],
[4.0, 3.0],
[5.0, 4.5],
[6.0, 5.0],
]
n = len(data) # 5 points
# Step 1 — find the mean (centre) of each column.
mean_x = sum(row[0] for row in data) / n
mean_y = sum(row[1] for row in data) / n
print("Mean point:", round(mean_x, 2), round(mean_y, 2)) # 4.0 3.2
# Step 2 — MEAN-CENTRE: subtract the mean so the cloud sits around (0, 0).
# PCA always centres first — variance is measured from the centre.
centered = [[row[0] - mean_x, row[1] - mean_y] for row in data]
print("Centred points:")
for p in centered:
print(" ", round(p[0], 2), round(p[1], 2))
# -2.0 -2.2 / -1.0 -0.7 / 0.0 -0.2 / 1.0 1.3 / 2.0 1.8
# Step 3 — PROJECT each 2D point onto ONE axis (a unit direction vector).
# Projection = dot product with the axis. This turns 2D into 1D.
# The data rises along a diagonal, so we pick a diagonal-ish axis.
import math
axis = [0.8, 0.6] # length = sqrt(0.8^2 + 0.6^2) = 1.0 (unit)
print("Axis length:", round(math.hypot(axis[0], axis[1]), 2)) # 1.0
print("1D coordinates (one number per point):")
projected = []
for p in centered:
coord = p[0] * axis[0] + p[1] * axis[1] # dot product → a single number
projected.append(coord)
print(" ", round(coord, 2))
# -2.92 / -1.22 / -0.12 / 1.58 / 2.68
# Step 4 — how much spread (variance) survived the squeeze?
mean_proj = sum(projected) / n
variance = sum((c - mean_proj) ** 2 for c in projected) / n
print("Variance kept on this axis:", round(variance, 2)) # 4.063 Explained Variance — How Many Components?
Doing the eigen-maths by hand is fine for two dimensions, but in practice you use scikit-learn. The number you care about most is the explained variance ratio: the fraction of total variance each principal component keeps. Add them up and you know how much information survives at each cut-off.
A common rule is to keep enough components to retain 90–95% of the cumulative variance. Read it straight off explained_variance_ratio_. Run the worked example below: four correlated-plus-noise features collapse to two components that hold ~95% of the information.
# The same idea with scikit-learn's PCA — and the number that matters most:
# EXPLAINED VARIANCE (how much information each component keeps).
import numpy as np
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
rng = np.random.default_rng(0)
# 4 features, but height/weight/shoe are correlated → real info is ~1-2 dims.
n = 200
height = rng.normal(170, 10, n)
weight = height * 0.6 + rng.normal(0, 4, n) # tied to height
shoe = height * 0.07 + rng.normal(0, 0.4, n) # tied to height
noise = rng.normal(0, 1, n) # pure noise feature
X = np.column_stack([height, weight, shoe, noise])
# ALWAYS scale before PCA — otherwise the big-numbered feature dominates.
X_scaled = StandardScaler().fit_transform(X)
pca = PCA() # keep all components so we can read the spectrum
pca.fit(X_scaled)
ratios = pca.explained_variance_ratio_
print("Explained variance per component:")
for i, r in enumerate(ratios, start=1):
print(f" PC{i}: {r*100:5.1f}% (cumulative {ratios[:i].sum()*100:5.1f}%)")
# Pick the smallest number of components that keep >= 90% of the variance.
cum = np.cumsum(ratios)
k = int(np.argmax(cum >= 0.90)) + 1
print(f"Components needed for 90% variance: {k}")
X_2d = PCA(n_components=2).fit_transform(X_scaled)
print("Reduced shape:", X_2d.shape) # (200, 2)
# Expected output (approx):
# PC1: ~70% (cumulative ~70%)
# PC2: ~25% (cumulative ~95%)
# PC3: ~ 4% (cumulative ~99%)
# PC4: ~ 1% (cumulative 100.0%)
# Components needed for 90% variance: 2
# Reduced shape: (200, 2)4 t-SNE and UMAP — Reduction for Your Eyes
PCA is linear: it can only cast straight shadows. When clusters curl around each other, a straight shadow smears them together. t-SNE and UMAP are non-linear methods built for one job — making a 2D or 3D picture in which similar points sit close together so clusters become visible.
t-SNE focuses on local structure: it tries to keep each point's nearest neighbours nearby. The perplexity setting controls roughly how many neighbours each point balances (typical 5–50). UMAP does a similar job, usually runs faster, and tends to preserve more of the global layout.
# t-SNE turns high-dimensional data into a 2D PICTURE so you can SEE clusters.
# It is for VISUALISATION ONLY — never feed t-SNE output into a model.
import numpy as np
from sklearn.datasets import load_digits
from sklearn.manifold import TSNE
digits = load_digits() # 1797 images of digits, each is 64 pixels (64-D)
X, y = digits.data, digits.target
print("Original:", X.shape) # (1797, 64)
# perplexity ~ "how many neighbours each point balances" (typical 5-50).
tsne = TSNE(n_components=2, perplexity=30, init="pca", random_state=0)
X_2d = tsne.fit_transform(X)
print("Embedded:", X_2d.shape) # (1797, 2)
# Same images of the same digit land near each other → tight clusters.
for digit in [0, 1, 2]:
pts = X_2d[y == digit]
cx, cy = pts.mean(axis=0)
print(f" digit {digit}: cluster centre ({cx:6.1f}, {cy:6.1f})")
# Expected output (exact numbers vary by version — clusters are what matter):
# Original: (1797, 64)
# Embedded: (1797, 2)
# digit 0: cluster centre ( .. , .. )
# digit 1: cluster centre ( .. , .. )
# digit 2: cluster centre ( .. , .. )
#
# UMAP is the popular alternative — same job, usually faster, and it keeps
# more GLOBAL layout. Swap in: from umap import UMAP; UMAP().fit_transform(X)5 Feature Selection vs Feature Extraction
There are two different ways to end up with fewer features, and people mix them up constantly:
Keep a subset of the original columns and drop the rest. The survivors are still your real, named features, so the result stays interpretable. Examples: drop low-variance columns, drop one of two highly correlated columns, or use model importances.
Build new features by combining the originals. PCA, t-SNE, UMAP, and autoencoders all do this. You usually pack more information into fewer features, but the new features (like "PC1") are harder to read.
Rule of thumb: if you need to explain the result to a human or a regulator, prefer selection. If you only need a compact, accurate input for a model or a plot, extraction usually wins.
🧭 When (and When Not) to Reduce Dimensions
Reach for reduction when:
- You want to visualise high-dimensional data in 2D or 3D (t-SNE / UMAP).
- Training is slow or models overfit because there are too many features.
- Features are highly correlated and you want to decorrelate them (PCA).
- You want to strip noise so the signal is cleaner.
- You only have a handful of features already — reduction may just lose information.
- Interpretability is essential — PCA components are hard to explain.
- A tree-based model (random forest, gradient boosting) is doing fine; these handle many features and irrelevant columns well, so PCA often hurts more than it helps.
Now you try. Fill in the blanks marked ___. Mean-centring is step 1 of every PCA.
# 🎯 YOUR TURN — mean-centre a tiny dataset, then check it worked.
# Mean-centring is step 1 of PCA: shift the cloud so its centre sits at (0, 0).
data = [
[10.0, 2.0],
[12.0, 4.0],
[14.0, 6.0],
[16.0, 8.0],
]
n = len(data)
# 1) Work out the mean of each column.
mean_x = ___ # 👉 sum of the x values (row[0]) divided by n
mean_y = ___ # 👉 sum of the y values (row[1]) divided by n
print("Mean:", mean_x, mean_y)
# 2) Subtract the mean from every point to centre the data.
centered = [[row[0] - mean_x, row[1] - mean_y] for row in data]
# 3) The mean of centred data must be (0, 0) — that is the whole point.
cx = sum(p[0] for p in centered) / n
cy = sum(p[1] for p in centered) / n
print("Centre after centring:", round(cx, 4), round(cy, 4))
# ✅ Expected output:
# Mean: 13.0 5.0
# Centre after centring: 0.0 0.0Your turn again. Project the centred points onto one axis with a dot product, then measure how much variance survived.
# 🎯 YOUR TURN — project 2D points onto ONE axis to get 1D coordinates.
# Projecting = dot product of each point with a unit direction vector.
centered = [ # already mean-centred for you
[-3.0, -1.0],
[-1.0, 0.0],
[ 1.0, 1.0],
[ 3.0, 0.0],
]
axis = [1.0, 0.0] # project onto the x-axis → just keep the x value
projected = []
for p in centered:
coord = ___ # 👉 dot product: p[0]*axis[0] + p[1]*axis[1]
projected.append(coord)
print("1D coordinates:", projected)
# Variance kept (spread of the 1D numbers around their mean).
m = sum(projected) / len(projected)
var = ___ # 👉 average of (c - m)**2 over every c in projected
print("Variance kept:", round(var, 2))
# ✅ Expected output:
# 1D coordinates: [-3.0, -1.0, 1.0, 3.0]
# Variance kept: 5.06 Common Errors (And How to Fix Them)
These four mistakes bite almost everyone learning dimensionality reduction:
❌ Not scaling before PCA
A feature in thousands (salary) drowns out a feature in fractions (a ratio) just because its numbers are bigger, so PC1 ends up being "salary" by accident.
PCA().fit(X) # ❌ raw units — big-scale feature dominatesX_scaled = StandardScaler().fit_transform(X)
PCA().fit(X_scaled) # ✅ every feature on equal footing❌ Reading meaning into t-SNE distances and cluster sizes
"These two clusters are far apart, so they're very different" and "this cluster is huge, so it has more points" are both wrong. t-SNE distorts gaps and inflates dense regions on purpose.
✅ Fix: read t-SNE for which points group together only. Try several perplexity values, and confirm any real claim with PCA or the raw distances.
❌ Keeping too few components
Jumping straight to n_components=2 because it plots nicely can throw away most of your signal if those two components only explain, say, 40% of the variance.
pca = PCA(n_components=2) # ❌ chosen blind✅ Fix: let the variance decide.
pca = PCA(n_components=0.95) # ✅ keep 95% of the variance
print(pca.n_components_) # how many that turned out to be❌ Data leakage — fitting on the test set
Calling fit (or fit_transform) on the whole dataset before splitting lets the scaler and PCA peek at the test data. Your scores look great, then collapse in production.
X_all = scaler.fit_transform(X) # ❌ test rows leaked into the fit
X_train, X_test = split(X_all)✅ Fix: fit on train only, then transform test — a Pipeline does this safely inside cross-validation.
from sklearn.pipeline import make_pipeline
pipe = make_pipeline(StandardScaler(), PCA(0.95), model)
pipe.fit(X_train, y_train) # ✅ fit only sees training data📋 Quick Reference
| Method | Type | Use for output in a model? | Best for |
|---|---|---|---|
| PCA | Linear, extraction | Yes | Compress correlated features, decorrelate, denoise |
| t-SNE | Non-linear | No — viz only | Seeing local clusters in 2D/3D |
| UMAP | Non-linear | No — viz only | Faster t-SNE with better global layout |
| Feature selection | Keeps originals | Yes | Interpretable results, dropping redundant columns |
| explained_variance_ratio_ | Diagnostic | — | Choosing how many components to keep (90–95%) |
🎯 Mini-Challenge: Selection vs Extraction by Hand
Time to fly with less scaffolding. The starter has a tiny 3-feature dataset and only a comment outline — write the logic yourself. One feature is a near-constant (it carries no signal); two carry the real pattern.
# 🎯 MINI-CHALLENGE: feature SELECTION vs feature EXTRACTION
#
# You have a tiny dataset with 3 features. Two are near-duplicates.
# 1. Mean-centre all 3 columns (subtract each column's mean).
# 2. SELECTION: drop the column whose values barely change (lowest variance).
# 3. EXTRACTION: build ONE new feature = average of the two informative columns.
# 4. Print the variance of each centred column, then your single new feature.
#
# ✅ Expected: the constant-ish column has ~0 variance and should be dropped;
# your extracted feature should still vary a lot (it carries the signal).
data = [
[1.0, 2.0, 5.0],
[3.0, 6.0, 5.0],
[5.0, 10.0, 5.0],
[7.0, 14.0, 5.0],
]
# your code hereLesson complete — you can now shrink data without losing the plot!
You understand the curse of dimensionality, you can mean-centre and project data by hand, run PCA and read its explained variance, reach for t-SNE or UMAP to see clusters, and tell feature selection from feature extraction. You also know the four traps: not scaling, over-reading t-SNE, keeping too few components, and leaking test data.
Practice quiz
What is the main goal of dimensionality reduction?
- To add more features to a dataset
- To label data automatically
- To reduce the number of features while keeping the useful information
- To increase the number of training rows
Answer: To reduce the number of features while keeping the useful information. Dimensionality reduction squeezes many features down to fewer while preserving as much signal as possible.
What does PCA find?
- The directions of greatest variance in the data
- The class boundaries between groups
- The nearest neighbours of each point
- The missing values to impute
Answer: The directions of greatest variance in the data. PCA finds the principal components — the orthogonal directions along which the data varies most.
Why is variance important in PCA?
- Low variance means more information
- Variance measures the number of classes
- Variance is ignored by PCA
- Variance is treated as information — spread-out directions separate points
Answer: Variance is treated as information — spread-out directions separate points. PCA chases variance because a direction with high spread carries the most information about the data.
Which is the first step of PCA before projecting?
- One-hot encoding the labels
- Mean-centring the data
- Splitting into train and test
- Adding polynomial features
Answer: Mean-centring the data. PCA mean-centres the data first so variance is measured from the origin.
What is t-SNE primarily used for?
- Visualising high-dimensional data in 2D or 3D
- Feeding compressed features into a model
- Speeding up gradient descent
- Encoding categorical variables
Answer: Visualising high-dimensional data in 2D or 3D. t-SNE is a non-linear method built for visualisation — never feed its output into a downstream model.
What is a key warning when reading a t-SNE plot?
- It preserves exact global distances
- It is fully reversible
- Cluster distances and sizes are not meaningful
- It always keeps 95% of the variance
Answer: Cluster distances and sizes are not meaningful. t-SNE distorts gaps and cluster sizes on purpose; only which points group together is reliable.
How does UMAP typically compare to t-SNE?
- It is slower and ignores local structure
- It is usually faster and keeps more global layout
- It only works on images
- It is a linear method like PCA
Answer: It is usually faster and keeps more global layout. UMAP does a similar job to t-SNE, usually runs faster, and tends to preserve more global structure.
What is the difference between feature selection and feature extraction?
- Selection builds new features; extraction keeps originals
- They are the same thing
- Selection only works for images
- Selection keeps a subset of original columns; extraction builds new combined features
Answer: Selection keeps a subset of original columns; extraction builds new combined features. Feature selection keeps a subset of the original interpretable columns; extraction (like PCA) builds new combined features.
Why should you scale data before running PCA?
- Because PCA cannot handle decimals
- Because variance depends on units, so a big-numbered feature would dominate
- Because scaling adds more components
- Scaling is never needed before PCA
Answer: Because variance depends on units, so a big-numbered feature would dominate. PCA chases variance, which depends on units; standardising puts every feature on equal footing first.
A common rule for choosing how many components to keep is to retain about:
- 10-20% of the variance
- exactly 50% of the variance
- 90-95% of the cumulative variance
- 100% always
Answer: 90-95% of the cumulative variance. A common rule is to keep enough components to retain roughly 90-95% of the cumulative explained variance.
Continue this course
- Previous: Unsupervised Learning
- Next: Ensemble Methods — Bagging, boosting and stacking — why combining weak models beats one strong one
- Quick reference: AI & Machine Learning cheat sheet
Frequently asked questions
What is dimensionality reduction in machine learning?
It is the process of squeezing data with many features (columns) down to fewer features while keeping as much of the useful information as possible. You do it to visualise data in 2D or 3D, to speed up and stabilise models, and to strip out noise and redundancy.
What is the curse of dimensionality?
As you add more features, the space grows so fast that data becomes sparse and almost every point looks equally far from every other point. Distances stop being meaningful, models need exponentially more data to learn, and overfitting gets easier. Reducing dimensions pushes back against this.
What is the difference between PCA and t-SNE?
PCA is a linear method that finds the directions of greatest variance; it is fast, reversible, and preserves global structure, so it is good for compressing features before modelling. t-SNE is a non-linear method built only for visualisation: it preserves local neighbourhoods so clusters pop out in 2D, but its distances and cluster sizes are not meaningful and you should never feed its output into a model.
What is the difference between feature selection and feature extraction?
Feature selection keeps a subset of your original columns and throws the rest away, so the surviving features stay interpretable. Feature extraction (like PCA) builds brand-new features by combining the originals; you usually keep more information per feature, but the new features are harder to read.
Why do I have to scale my data before PCA?
PCA chases variance, and variance depends on units. A feature measured in thousands (like salary) will swamp a feature measured in fractions (like a ratio) purely because its numbers are bigger. Standardising every feature to mean 0 and standard deviation 1 with StandardScaler puts them on equal footing first.
How many components should I keep?
A common rule is to keep enough components to retain about 90-95% of the cumulative explained variance, which you read straight off explained_variance_ratio_. You can also look for the 'elbow' in a scree plot. Keeping too few throws away real signal; keeping too many defeats the purpose.