Python for Machine Learning

Reviewed & published by Brayan K

Meet the four tools every ML project leans on — NumPy, pandas, scikit-learn, and matplotlib — and learn the one workflow (split → fit → predict → score) that ties them together.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

🛠️ Real-World Analogy: A Workshop and Its Tools

Picture building a piece of furniture in a workshop. You don't reach for one magic machine — you reach for the right tool at the right step. The Python ML stack works the same way: four specialised tools, each doing one job well, used in a fixed order.

📐 NumPy — the workbench

Fast arrays and matrix maths. Every other tool stacks its work on this surface.

📊 pandas — the parts bins

Labelled tables (DataFrames) to load, clean, filter, and sort your raw materials.

🧪 scikit-learn — the power tools

The algorithms that do the cutting and joining: fit, predict, score — same handles on every one.

📈 matplotlib — the tape measure

Charts to check your work before and after — it measures and inspects, it doesn't build.

Keep this picture in mind: pandas prepares the materials, NumPy holds the numbers, scikit-learn does the work, matplotlib checks it.

1 NumPy Arrays — The Container for the Numbers

A NumPy array is a grid of numbers that all share one type, stored together in memory. That layout is why maths on an array is fast: one operation runs across every element at once — called a vectorised operation — instead of you writing a Python loop.

Two words you'll meet constantly: a 1D array is a vector (a single row of values), and a 2D array is a matrix (a table where rows are samples and columns are features). The array's .shape tells you how many of each.

Worked example — read the comments, every line states its result:

import numpy as np

# A NumPy array is a fast, fixed-type grid of numbers. It is the
# container every ML library passes data around in.

prices = np.array([10.0, 20.0, 30.0, 40.0])   # 1D array (a "vector")
print(prices)            # Expected output: [10. 20. 30. 40.]
print(prices.shape)      # Expected output: (4,)   <- 4 elements, 1 axis

# Vectorised math: one operation applies to EVERY element, no loop.
print(prices * 1.2)      # Expected output: [12. 24. 36. 48.]
print(prices.mean())     # Expected output: 25.0

# A 2D array is a matrix: rows = samples, columns = features.
X = np.array([[1, 2],
              [3, 4],
              [5, 6]])
print(X.shape)           # Expected output: (3, 2)   <- 3 samples, 2 features
print(X.sum(axis=0))     # Expected output: [9 12]   <- sum down each column

2 pandas DataFrames — Your Data as a Table

Real data rarely arrives as a tidy grid of one type — it has named columns like age, city, and price, mixing numbers, text, and true/false. A DataFrame is pandas' answer: a labelled table you can program, like a spreadsheet with code.

The move you'll repeat in every project is splitting that table into X (the feature columns the model learns from) and y (the single target column it learns to predict). Everything after this lesson assumes you can do that split.

import pandas as pd

# A DataFrame is a labelled table: named columns + indexed rows.
# Think of it as a spreadsheet you can program.

df = pd.DataFrame({
    "city":   ["London", "Paris", "Berlin", "Rome"],
    "temp_c": [12, 16, 9, 18],
    "rain":   [True, False, True, False],
})

print(df.shape)              # Expected output: (4, 3)   <- 4 rows, 3 columns
print(df["temp_c"].mean())   # Expected output: 13.75

# Filter rows with a boolean condition (this is the workhorse of pandas):
warm = df[df["temp_c"] > 12]
print(warm["city"].tolist())  # Expected output: ['Paris', 'Rome']

# Split into features (X) and the label/target (y) for ML:
X = df[["temp_c", "rain"]]    # the inputs the model learns FROM
y = df["city"]                # the answer the model learns to predict
print(X.shape, y.shape)       # Expected output: (4, 2) (4,)

Notice df[df["temp_c"] > 12]: you build a column of True/False values, then use it to keep only the matching rows. That boolean-filter trick is the single most-used pandas operation.

3 The scikit-learn Pattern — fit / predict / score

scikit-learn's superpower is consistency: a decision tree, a linear model, and a support-vector machine all expose the same three methods. Learn them once and you can drive any model.

Two supporting pieces make those three honest. train_test_split holds back a slice of data the model never trains on, so the score is earned, not memorised. A Pipeline glues a transformer (something that reshapes the data, like a scaler that re-scales features) to the model, so the transformer is fitted on the training data only — never on the test data.

Worked example — the full pattern in one place:

from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

# Every scikit-learn model speaks the SAME three verbs:
#   .fit(X, y)   -> learn from labelled data
#   .predict(X)  -> guess labels for new data
#   .score(X, y) -> how often it was right (accuracy, 0.0-1.0)

# 1) Split FIRST so the test set stays unseen during training.
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

# 2) A Pipeline chains a transformer + a model into one object, so the
#    scaler is fitted on TRAIN ONLY and reused on test — no leakage.
model = make_pipeline(
    StandardScaler(),          # transformer: re-scales features
    LogisticRegression(),      # estimator: makes the prediction
)

model.fit(X_train, y_train)              # learn
preds = model.predict(X_test)            # predict
print(model.score(X_test, y_test))       # Expected output: 0.97  (example)

# .score() just compares preds to the true labels — exactly what you will
# compute by hand in the runnable exercise below.

4 Under the Hood — A Prediction Is Just Arithmetic

np.dot and .score sound advanced, but underneath they are plain arithmetic you can write without any library. Seeing that demystifies the whole stack — and it runs in the editor right now.

A single linear prediction is a dot product: multiply each feature by its weight, sum the results, and add a constant bias. That's all np.dot(features, weights) + bias does.

Worked example — a dot product by hand:

# What does numpy actually DO under the hood? A model prediction is mostly
# a "dot product": multiply each feature by its weight, then add them up.
# Here it is in plain Python so you can see there is no magic.

features = [2.0, 3.0, 1.0]    # one sample: 3 feature values
weights  = [0.5, 1.0, 2.0]    # what the model "learned" for each feature
bias     = 1.0                # a constant added at the end

total = bias
for f, w in zip(features, weights):
    total += f * w            # accumulate feature * weight

print("Prediction:", total)
# Expected output: Prediction: 7.0
# (1.0 + 2*0.5 + 3*1.0 + 1*2.0 = 1 + 1 + 3 + 2 = 7.0)
# In NumPy this is just:  np.dot(features, weights) + bias
# 🎯 YOUR TURN — compute a dot product by hand (this is what np.dot does)

features = [4.0, 2.0, 5.0]
weights  = [1.0, 3.0, 0.0]
bias     = 2.0

total = ___              # 👉 start the running total at the bias value

for f, w in zip(features, weights):
    total = ___          # 👉 add f * w to the running total each loop

print("Prediction:", total)

# ✅ Expected output: Prediction: 12.0
# (2.0 + 4*1.0 + 2*3.0 + 5*0.0 = 2 + 4 + 6 + 0 = 12.0)
# 🎯 YOUR TURN — score a model by hand (this is what .score() does)
# Accuracy = (number of correct guesses) / (total guesses)

predictions = ["cat", "dog", "cat", "bird", "dog"]   # what the model guessed
labels      = ["cat", "dog", "dog", "bird", "dog"]   # the true answers

correct = 0
for pred, true in zip(predictions, labels):
    if ___:                      # 👉 count it only when the guess matches the label
        correct += 1

accuracy = ___ / len(labels)     # 👉 divide correct guesses by the total count
print("Accuracy:", accuracy)

# ✅ Expected output: Accuracy: 0.8
# (4 of 5 correct — only the 3rd prediction was wrong)

📈 Where matplotlib Fits

matplotlib (and Seaborn, which is built on top of it) is for seeing your data — it never trains a model. You reach for it at two moments: before modelling to explore patterns, and after to inspect how the model did.

import matplotlib.pyplot as plt

# BEFORE modelling — explore the data
plt.scatter(df["temp_c"], df["rain"])   # do features relate to the target?
plt.hist(df["temp_c"])                  # what does the spread look like?

# AFTER modelling — inspect the results
plt.scatter(y_test, preds)              # predicted vs actual: close = good
plt.show()                              # render the figure on screen

Keep the boundary clear: pandas/NumPy hold the data, scikit-learn models it, matplotlib pictures it. A chart never changes a prediction — it changes your understanding.

! Common Errors (And How to Fix Them)

These four trip up almost every beginner. Spotting them early saves hours.

❌ Fitting on the test data (data leakage)

Scaling or fitting using the whole dataset before splitting:

scaler.fit(X)                    # ❌ sees the test rows too
X_train, X_test = train_test_split(X)

✅ Fix: split first, then fit on train only (a Pipeline does this for you):

X_train, X_test = train_test_split(X)
scaler.fit(X_train)              # ✅ test set stays unseen

❌ Not splitting at all

Scoring on the same rows the model trained on:

model.fit(X, y)
model.score(X, y)    # ❌ ~1.0 — it just memorised the answers

✅ Fix: always evaluate on held-out data:

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)
model.fit(X_train, y_train)
model.score(X_test, y_test)   # ✅ an honest score

ValueError: Found input variables with inconsistent numbers of samples — X and y have different lengths, or a 1D array was passed where 2D was expected:

model.fit(X, y)   # ❌ X has 100 rows, y has 90

✅ Fix: check shapes line up, and reshape a single feature to 2D:

print(X.shape, y.shape)   # must share the same first number
X = X.reshape(-1, 1)      # ✅ make one feature column 2D

❌ Scaling after the split, but fitting the scaler on test too

You split correctly, then re-fit the scaler on the test set:

scaler.fit(X_train); X_train = scaler.transform(X_train)
X_test = scaler.fit_transform(X_test)   # ❌ refits on test

✅ Fix: fit once on train; only transform the test set:

scaler.fit(X_train)
X_train = scaler.transform(X_train)
X_test  = scaler.transform(X_test)   # ✅ same parameters, no leakage

📋 Quick Reference

CallDoes
np.array([1,2,3])Make an array (vector / matrix)
arr.shapeRows & columns: (samples, features)
arr * 2, arr + arr2Vectorised maths, no loop
np.dot(a, b)Dot / matrix product
arr.mean(), arr.std()Summary statistics
CallDoes
pd.DataFrame(data)Build a labelled table
df["col"]Select one column (a Series)
df[df["x"] > 0]Filter rows by a condition
df.groupby("g").mean()Aggregate by group
df.describe()Quick stats for every column
CallDoes
train_test_split(X, y)Hold back a test set
model.fit(X_train, y_train)Train the model
model.predict(X_test)Guess labels for new data
model.score(X_test, y_test)Accuracy on held-out data
make_pipeline(scaler, model)Chain transformer + model, no leakage

🎯 Mini-Challenge: Train, Predict, Score (Plain Python)

Put the whole workflow together — no libraries. Your "model" is a simple threshold rule, and you'll score it by hand, exactly the way scikit-learn's .score() works. The starter below is a comment outline only.

# 🎯 MINI-CHALLENGE: a tiny train / predict / score loop in plain Python
#
# A "model" here is just a threshold: predict "pass" if score >= 50, else "fail".
#
# 1. Make a list 'scores' = [80, 45, 60, 30, 95] and a matching list of true
#    'labels' = ["pass", "fail", "pass", "fail", "pass"].
# 2. Loop through scores and build a 'predictions' list using the threshold 50.
# 3. Count how many predictions match the labels.
# 4. Print the accuracy (correct / total).
#
# ✅ Expected output: Accuracy: 1.0   (all 5 predicted correctly)

# your code here

Lesson complete — you know the ML toolkit and the workflow!

You can describe a NumPy array and its shape, split a DataFrame into X and y, drive any scikit-learn model with fit / predict / score, split data to avoid leakage, and place matplotlib correctly in the pipeline. You even computed a prediction and an accuracy by hand — so none of it is a black box.

Practice quiz

Why use a NumPy array instead of a plain Python list for maths?

  • It can store text but not numbers
  • It automatically trains a model
  • It runs vectorised operations in fast compiled code, often far faster than a Python loop
  • It never needs a shape

Answer: It runs vectorised operations in fast compiled code, often far faster than a Python loop. NumPy stores one type contiguously and runs math in compiled C, making vectorised ops much faster.

In a 2D NumPy array used for ML, what do rows and columns usually represent?

  • Rows are samples, columns are features
  • Rows are features, columns are samples
  • Both are labels
  • Rows are predictions, columns are errors

Answer: Rows are samples, columns are features. Conventionally each row is a sample and each column is a feature; .shape is (samples, features).

What is a pandas DataFrame?

  • A single number
  • A neural network layer
  • A type of plot
  • A labelled table with named columns and indexed rows

Answer: A labelled table with named columns and indexed rows. A DataFrame is a programmable, spreadsheet-like labelled table of mixed-type columns.

When splitting a DataFrame for ML, what are X and y?

  • X is the target column; y is the feature columns
  • X is the feature columns the model learns from; y is the target column it predicts
  • X and y are both labels
  • X is the index; y is the header

Answer: X is the feature columns the model learns from; y is the target column it predicts. X holds the input feature columns; y is the single target the model learns to predict.

What do scikit-learn's fit, predict, and score methods do?

  • fit trains on labelled data; predict guesses labels for new data; score measures accuracy
  • fit plots data; predict scales it; score deletes it
  • They all train the model
  • fit and predict are identical

Answer: fit trains on labelled data; predict guesses labels for new data; score measures accuracy. Every estimator shares fit (train), predict (guess), and score (evaluate) — for classifiers score is accuracy.

Why must you split data into train and test sets before evaluating?

  • To make training slower
  • Because models require exactly two datasets
  • So the score reflects performance on unseen data instead of memorised rows
  • To remove all missing values

Answer: So the score reflects performance on unseen data instead of memorised rows. Evaluating on held-out data gives an honest estimate; scoring on training rows just measures memorisation.

What is data leakage in this context?

  • Saving the model to the wrong folder
  • Information from the test set sneaking into training, e.g. fitting a scaler on the whole dataset before splitting
  • Using too few features
  • A bug in NumPy

Answer: Information from the test set sneaking into training, e.g. fitting a scaler on the whole dataset before splitting. Leakage gives over-optimistic scores; fit transformers on the training fold only.

How does a scikit-learn Pipeline help prevent leakage?

  • It deletes the test set
  • It skips the scaler entirely
  • It doubles the dataset
  • It fits every transformer on the training fold only and reuses those parameters on the test fold

Answer: It fits every transformer on the training fold only and reuses those parameters on the test fold. A Pipeline ties transformer + model so the scaler is fit on train only, never on test.

Underneath, a single linear prediction is mostly which operation?

  • Sorting the features
  • A dot product: multiply each feature by its weight, sum, then add a bias
  • Counting the rows
  • Removing punctuation

Answer: A dot product: multiply each feature by its weight, sum, then add a bias. np.dot(features, weights) + bias is the arithmetic behind a linear prediction.

Where does matplotlib fit in the ML workflow?

  • It trains the model
  • It replaces scikit-learn
  • It visualises data before and after modelling, but never trains the model itself
  • It cleans the data automatically

Answer: It visualises data before and after modelling, but never trains the model itself. matplotlib is for seeing data and results; pandas/NumPy hold it, scikit-learn models it.

Continue this course

Frequently asked questions

Why use NumPy arrays instead of plain Python lists?

NumPy stores numbers in a contiguous, single-type block of memory and runs math in fast compiled C, so vectorised operations like arr * 2 are often 10-50x faster than a Python loop. It also adds the shape, broadcasting, and matrix maths that pandas and scikit-learn are built on top of.

What is the difference between fit, predict, and score in scikit-learn?

fit(X, y) trains the model by learning patterns from labelled data. predict(X) uses the trained model to guess labels for new, unseen inputs. score(X, y) measures how good those guesses are — for classifiers it returns accuracy (the fraction of correct predictions), a number from 0.0 to 1.0.

Why do I have to split data into train and test sets?

If you evaluate a model on the same rows it learned from, it can simply memorise them and look perfect while failing on real data. train_test_split holds back a portion (commonly 20%) that the model never sees during training, so the score reflects how it performs on genuinely new examples.

What is data leakage and how does a Pipeline prevent it?

Leakage is when information from the test set sneaks into training — most commonly by fitting a scaler on the whole dataset before splitting. A scikit-learn Pipeline fits every transformer on the training fold only and reuses those exact parameters on the test fold, so the test set stays truly unseen.

Where does matplotlib fit into the machine-learning workflow?

Matplotlib (and Seaborn, built on it) is for visualising, not modelling. You use it before training to explore distributions and relationships, and after training to inspect predictions, errors, and feature importance. It never touches the model itself — it just helps you understand the data and results.