Ensemble Methods

Reviewed & published by Brayan K

By the end of this lesson you'll be able to combine several models into one stronger predictor — and explain exactly why a crowd of models beats the best individual model.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

🧠 Real-World Analogy: A Panel of Experts

Imagine a tricky medical case. You don't bet everything on one doctor — you convene a panel of specialists. The cardiologist, the radiologist, and the surgeon each see the problem from a different angle. Individually each can be wrong; but when most of the panel agrees on a diagnosis, you trust the consensus far more than any single voice.

An ensemble is exactly that panel, made of machine-learning models. Each model is an "expert" with its own blind spots. Because they make different mistakes, pooling their answers cancels out the random errors and keeps the shared signal. That is the whole idea — everything below is just three ways to assemble and run the panel.

🔑 The four strategies you'll meet

1 Why Ensembles Beat Single Models

A single model has a fixed weakness. A deep decision tree memorises noise (high variance — it swings wildly with the data). A shallow one is too simple to capture the pattern (high bias — it underfits). You can't escape both at once with one model.

Ensembles sidestep this. If three models each get a question right 70% of the time and their mistakes are independent, a majority vote is right far more than 70% of the time — because for the vote to be wrong, two models must fail on the same sample at once, which is much rarer. The key word is diverse: models that make the same mistakes give you nothing.

Run the worked example below. Three flawed models, none of them perfect, combine into a perfect ensemble on this set — purely because they err on different samples. This uses plain Python, no libraries.

# Majority vote of three classifiers — PLAIN Python, no libraries needed.
# Each "model" already gave its prediction (0 or 1) for 8 test samples.
# The ensemble takes the prediction that the MOST models agree on.

true_labels = [1, 0, 1, 1, 0, 1, 0, 1]   # the correct answers

# Three different models' predictions (each one is imperfect)
model_a    = [1, 0, 1, 0, 0, 1, 1, 1]   # wrong on samples 4 and 7
model_b    = [1, 1, 1, 1, 0, 0, 0, 1]   # wrong on samples 2 and 6
model_c    = [0, 0, 1, 1, 0, 1, 0, 1]   # wrong on sample 1

def majority_vote(a, b, c):
    """Return the label that at least 2 of the 3 models chose."""
    votes = a + b + c          # each is 0 or 1, so the sum is 0..3
    return 1 if votes >= 2 else 0   # 2 or 3 ones → predict 1

# Build the ensemble prediction for every sample
ensemble = []
for i in range(len(true_labels)):
    ensemble.append(majority_vote(model_a[i], model_b[i], model_c[i]))

# Count how many each model (and the ensemble) got right
def accuracy(preds):
    correct = sum(1 for p, t in zip(preds, true_labels) if p == t)
    return correct / len(true_labels)

print("Sample-by-sample:")
print("idx true  A  B  C  vote")
for i in range(len(true_labels)):
    mark = "OK" if ensemble[i] == true_labels[i] else "X"
    print(f"  {i+1}    {true_labels[i]}   {model_a[i]}  {model_b[i]}  {model_c[i]}   {ensemble[i]}  {mark}")

print()
print(f"Model A accuracy : {accuracy(model_a):.0%}")
print(f"Model B accuracy : {accuracy(model_b):.0%}")
print(f"Model C accuracy : {accuracy(model_c):.0%}")
print(f"Ensemble accuracy: {accuracy(ensemble):.0%}   <-- the crowd wins!")

# Expected output:
#   No single model is perfect (each is wrong on 1-2 samples), but because
#   they make DIFFERENT mistakes, the majority vote reaches 100% on this set.

2 Bagging & Random Forests — Cutting Variance

Bagging (short for Bootstrap Aggregating) trains many copies of the same model type, each on a different random sample of the training data drawn with replacement (a "bootstrap" sample, where some rows appear twice and others not at all). You then average their predictions (or vote, for classification).

Because each model sees slightly different data, their over-fitting quirks point in different directions and cancel out when averaged. Bagging attacks the variance half of the bias–variance tradeoff — it makes an unstable, overfit-prone model far more stable.

A Random Forest is bagging applied to decision trees, with one extra twist: at each split, each tree may only consider a random subset of features. That forces the trees to be different from one another, boosting diversity even more. It's the go-to low-effort, high-accuracy baseline for tabular data.

Here is the same idea using scikit-learn's RandomForestClassifier inside a soft-voting ensemble:

from sklearn.ensemble import RandomForestClassifier, VotingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.svm import SVC
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=42)

# Three DIFFERENT model types — diversity is what makes voting work
clf1 = LogisticRegression(max_iter=1000)
clf2 = RandomForestClassifier(n_estimators=100, random_state=42)
clf3 = SVC(probability=True, random_state=42)

# 'soft' voting averages predicted probabilities (usually beats 'hard')
ensemble = VotingClassifier(
    estimators=[("lr", clf1), ("rf", clf2), ("svc", clf3)],
    voting="soft",
)
ensemble.fit(X_tr, y_tr)
print(f"Voting accuracy: {ensemble.score(X_te, y_te):.2f}")

# Expected output:
# Voting accuracy: 1.00
# A simple AVERAGING ensemble for regression — PLAIN Python.
# Three models each predict a house price. We average them.
# Averaging cancels out each model's random over/under-shoot.

actual_price = 300_000   # the true price we are trying to predict

# Each model's guess (in dollars). None is perfect.
model_1 = 280_000   # too low  by 20k
model_2 = 330_000   # too high by 30k
model_3 = 305_000   # close, high by 5k

predictions = [model_1, model_2, model_3]

def average(values):
    return sum(values) / len(values)

ensemble_price = average(predictions)

print(f"Actual price   : ${actual_price:,.0f}")
print("-" * 30)
for i, p in enumerate(predictions, start=1):
    error = abs(p - actual_price)
    print(f"Model {i} guess  : ${p:,.0f}  (off by ${error:,.0f})")
print("-" * 30)
ens_error = abs(ensemble_price - actual_price)
print(f"Ensemble (avg) : ${ensemble_price:,.0f}  (off by ${ens_error:,.0f})")

# Expected output:
#   Ensemble (avg) : $305,000  (off by $5,000)
#   The averaged error is smaller than the worst single model — the
#   high and low guesses partly cancel each other out.

3 Boosting — Cutting Bias, One Correction at a Time

Boosting flips bagging on its head. Instead of training models independently and in parallel, it trains them sequentially. Each new model focuses on the examples the previous models got wrong, so the team steadily chips away at the errors. Boosting attacks the bias half of the tradeoff: it turns a pile of "weak learners" (models barely better than guessing) into one strong learner.

The price of boosting's accuracy: it can overfit if you use too many rounds or too high a learning rate, and it's sequential, so it's harder to parallelise than bagging. Here's gradient boosting in scikit-learn:

from sklearn.ensemble import GradientBoostingClassifier
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split

X, y = load_breast_cancer(return_X_y=True)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.3, random_state=42)

# Boosting: 100 small trees, each fixing the previous trees' errors.
# A SMALL learning_rate + MORE trees usually generalises better.
model = GradientBoostingClassifier(
    n_estimators=100,
    learning_rate=0.1,
    max_depth=3,
    random_state=42,
)
model.fit(X_tr, y_tr)
print(f"Train accuracy: {model.score(X_tr, y_tr):.3f}")
print(f"Test  accuracy: {model.score(X_te, y_te):.3f}")

# Expected output:
# Train accuracy: 1.000
# Test  accuracy: 0.965
# (Train far above test would signal the boosting is overfitting.)

4 Voting & Stacking — Combining Different Model Types

Bagging and boosting build an ensemble from one kind of model. Voting and stacking instead combine different model types — say a logistic regression, a random forest, and an SVM — to maximise diversity.

Stacking is the most powerful but the most prone to leakage: the meta-model must be trained on held-out predictions (via cross-validation), never on predictions the base models made for data they were trained on — see the Common Errors below.

🎯 Your Turn #1: Weighted Average Ensemble

A plain average treats every model equally. But a stronger model should count for more. Fill in the three blanks to compute a weighted average, then run it and check against the expected output in the comments.

# 🎯 YOUR TURN — a WEIGHTED averaging ensemble.
# Better models should count for more. Give each prediction a weight.

actual = 300_000

# (prediction, weight) — a trusted model gets a bigger weight
model_1 = (280_000, 1)   # weak  model, weight 1
model_2 = (330_000, 1)   # weak  model, weight 1
model_3 = (305_000, 3)   # strong model, weight 3

models = [model_1, model_2, model_3]

# 👉 1) Sum up (prediction * weight) for every model into 'weighted_sum'
weighted_sum = ___

# 👉 2) Sum up just the weights into 'total_weight'
total_weight = ___

# 👉 3) The weighted average = weighted_sum / total_weight
prediction = ___

print(f"Weighted ensemble: ${prediction:,.0f}")
print(f"Off by: ${abs(prediction - actual):,.0f}")

# ✅ Expected output:
#   Weighted ensemble: $305,000
#   Off by: $5,000
# (weighted_sum = 280000*1 + 330000*1 + 305000*3 = 1,525,000;
#  total_weight = 5; 1,525,000 / 5 = 305,000)

🎯 Your Turn #2: Weighted Majority Vote

Now do the classification version. Each model votes 0 or 1, but a more accurate model's vote carries more weight. Fill in the blanks to tally the weighted votes and pick the winner.

# 🎯 YOUR TURN — a WEIGHTED majority vote for classification.
# Each model votes 0 or 1, but a more accurate model's vote counts more.

# (vote, weight) for ONE sample whose true label is 1
votes = [
    (1, 2),   # model A votes 1, weight 2
    (0, 1),   # model B votes 0, weight 1
    (1, 1),   # model C votes 1, weight 1
]

# 👉 1) Add up the weights of every model that voted 1 → 'score_for_1'
score_for_1 = ___

# 👉 2) Add up the weights of every model that voted 0 → 'score_for_0'
score_for_0 = ___

# 👉 3) Predict 1 if score_for_1 is bigger, else 0
final = ___

print(f"Score for 1: {score_for_1}")
print(f"Score for 0: {score_for_0}")
print(f"Final prediction: {final}")

# ✅ Expected output:
#   Score for 1: 3
#   Score for 0: 1
#   Final prediction: 1
# (Models A and C voted 1 with weights 2+1=3, beating B's weight of 1.)

🧗 Mini-Challenge: A 5-Model Ensemble (support faded)

No blanks this time — just a comment outline. Build the whole majority-vote loop yourself for five models, then print the ensemble accuracy. The expected result is in the comments so you can self-check.

# 🎯 MINI-CHALLENGE: build a 5-model majority-vote ensemble
#
# 1. You are given true_labels and a list of 5 models' predictions (preds).
# 2. For each sample, count how many models voted 1.
#    - Predict 1 if 3 or more models agree (majority of 5), else 0.
# 3. Print the ensemble accuracy as a percentage.
#
# ✅ Expected: the ensemble should score 100% even though no single
#    model is perfect (each model is wrong on one sample).

true_labels = [1, 0, 1, 1, 0]
preds = [
    [1, 0, 1, 0, 0],   # model 1 (wrong on sample 4)
    [1, 1, 1, 1, 0],   # model 2 (wrong on sample 2)
    [0, 0, 1, 1, 0],   # model 3 (wrong on sample 1)
    [1, 0, 0, 1, 0],   # model 4 (wrong on sample 3)
    [1, 0, 1, 1, 1],   # model 5 (wrong on sample 5)
]

# your code here

5 Common Errors (And How to Fix Them)

❌ Correlated models give no boost

You ensemble five models and accuracy barely moves. They're all the same algorithm on the same data, so they make the same mistakes — averaging identical errors changes nothing.

✅ Fix: maximise diversity. Use different algorithms, different feature subsets (Random Forest does this), or different data samples (bagging). Diversity is the fuel of every ensemble.

Your gradient-boosting model scores 100% on training but slumps on the test set. Too many rounds or too high a learning_rate let it memorise noise.

✅ Fix: lower the learning_rate, cap n_estimators, use early_stopping on a validation set, and keep trees shallow (max_depth=3). More small trees beat fewer big ones.

❌ Data leakage in stacking

Your stacked model looks amazing in testing but flops in production. The meta-model was trained on predictions the base models made for their own training rows — it peeked at the answers.

✅ Fix: generate the meta-features with cross-validation (out-of-fold predictions only). Scikit-learn's StackingClassifier does this for you — don't hand-roll it unless you replicate the CV.

❌ Inference is too slow

A 500-tree forest or a giant stack is accurate but answers each request too slowly for real time.

✅ Fix: prune the ensemble (fewer estimators), pick a faster library (LightGBM over a deep forest), or distil the ensemble into one small model. Accuracy that misses the latency budget is worthless in production.

📋 Quick Reference: Bagging vs Boosting

AspectBaggingBoosting
How models trainIn parallel, independentlySequentially, each fixes the last
Data per modelRandom bootstrap sampleRe-weighted toward past errors
Mainly reducesVariance (overfitting)Bias (underfitting)
Overfit riskLow — very robustHigher — needs tuning
Combine byAveraging / majority voteWeighted sum of learners
Top algorithmsRandom ForestAdaBoost, XGBoost, LightGBM

💡 Pro Tip: For tabular data, start with Random Forest as a no-fuss baseline, then try XGBoost or LightGBM when you need the last few points of accuracy. Deep learning usually only wins on images, text, and audio — on structured data, gradient boosting is still king.

🎉 Lesson Complete!

You can now explain bagging (parallel, variance-cutting — Random Forest), boosting (sequential, bias-cutting — AdaBoost, Gradient Boosting, XGBoost/LightGBM), and voting & stacking for mixing model types. You built a working majority vote and an averaging ensemble in plain Python, and you know the bias–variance reason ensembles beat single models.

Practice quiz

What is an ensemble method?

  • A single very deep neural network
  • A way to clean missing data
  • Combining predictions of several models into one final answer
  • A method for reducing dimensions

Answer: Combining predictions of several models into one final answer. An ensemble combines several models' predictions — by averaging, voting, or stacking — to beat any single model.

How does bagging train its models?

  • Independently and in parallel on random bootstrap samples
  • Sequentially, each fixing the last one's errors
  • By splitting the model across GPUs
  • Using a single model on all the data

Answer: Independently and in parallel on random bootstrap samples. Bagging trains many models in parallel on random bootstrap samples, then averages them.

Bagging mainly reduces which part of the error?

  • Bias
  • The number of features
  • The learning rate
  • Variance

Answer: Variance. Bagging attacks variance — it stabilises an overfit-prone model by averaging diverse copies.

How does boosting train its models?

  • All at once on identical data
  • Sequentially, each new model fixing the previous models' mistakes
  • By averaging random subsets
  • By dropping low-variance columns

Answer: Sequentially, each new model fixing the previous models' mistakes. Boosting trains models one after another, each focusing on the examples the previous ones got wrong.

Boosting mainly reduces which part of the error?

  • Bias
  • Variance
  • Memory usage
  • The number of trees

Answer: Bias. Boosting attacks bias — it turns weak learners into one strong learner that underfits less.

What is a Random Forest?

  • Boosting applied to linear models
  • A single very deep decision tree
  • Bagging applied to decision trees with random feature subsets at each split
  • A stacking meta-model

Answer: Bagging applied to decision trees with random feature subsets at each split. A Random Forest is bagging of decision trees, with each split considering a random subset of features for extra diversity.

Why do diverse ensembles beat single models?

  • Because all models make the same mistakes
  • Because diverse models make different mistakes that cancel out
  • Because they use more memory
  • Because they always overfit

Answer: Because diverse models make different mistakes that cancel out. When models make independent errors, those errors tend to cancel while the correct signal reinforces.

What is the difference between hard voting and soft voting?

  • Hard voting averages probabilities; soft voting picks the majority class
  • They are identical
  • Soft voting only works for regression
  • Hard voting takes a majority class vote; soft voting averages predicted probabilities

Answer: Hard voting takes a majority class vote; soft voting averages predicted probabilities. Hard voting counts class votes; soft voting averages predicted probabilities and usually performs better.

What is stacking?

  • Stacking many copies of the same model
  • Training a meta-model that learns how to combine the base models' predictions
  • Adding more layers to a neural network
  • Removing correlated features

Answer: Training a meta-model that learns how to combine the base models' predictions. Stacking trains a meta-model on the base models' predictions to learn the best way to blend them.

Which library family typically dominates tabular data competitions?

  • Convolutional neural networks
  • k-means clustering
  • Gradient boosting (XGBoost, LightGBM)
  • PCA

Answer: Gradient boosting (XGBoost, LightGBM). Gradient-boosting libraries like XGBoost and LightGBM win most structured/tabular-data tasks.

Continue this course

Frequently asked questions

What is an ensemble method in machine learning?

An ensemble combines the predictions of several models into one final answer — by averaging, voting, or stacking — so the group is more accurate and more stable than any single model on its own.

What is the difference between bagging and boosting?

Bagging trains many independent models in parallel on random subsets of the data and averages them to reduce variance. Boosting trains models one after another, each one fixing the mistakes of the last, to reduce bias. Bagging fights overfitting; boosting fights underfitting.

Why do ensembles beat single models?

Because diverse models make different mistakes. When you combine them, the independent errors tend to cancel out while the correct signal reinforces, so the consensus is usually right even when individuals are wrong.

When should I use Random Forest vs XGBoost?

Random Forest (bagging) is a strong, low-tuning default that rarely overfits. XGBoost and LightGBM (boosting) usually score higher on tabular data but need more careful tuning of learning rate and tree count to avoid overfitting. On Kaggle, gradient boosting wins most structured-data tasks.

Do ensembles only work for classification?

No. For classification you take a majority vote (or average probabilities); for regression you average the numeric predictions. The same bagging, boosting, and stacking ideas apply to both.

Related lessons