Feature Engineering & Selection

Reviewed & published by Brayan K

Better inputs beat fancier models. Learn to scale, encode, combine, and bin features — then select the ones that matter while avoiding the trap of data leakage.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

🎯 Real-World Analogy: Prepping Ingredients

A great chef spends most of the time on prep — washing, chopping, and measuring — before anything hits the pan. The cooking method matters, but raw, badly prepped ingredients ruin even a perfect recipe. Feature engineering is the prep work of machine learning.

Scaling is measuring to the same units, encoding is translating words into numbers the model can taste, and feature selection is throwing out the ingredients that don't belong. Get the prep right and a simple model cooks beautifully.

1 Scaling & Encoding

Scaling puts numeric features on a comparable footing so no single large-range feature dominates. Two common choices: standardisation (mean 0, std 1) and min-max scaling (squash to 0..1). Distance- and gradient-based models need it; trees do not.

Encoding turns categories into numbers. Use one-hot encoding for unordered categories (one 0/1 column each) so you never imply a false order. Plain label encoding is fine only for genuinely ordered categories (small < medium < large).

2 Scale & Encode by Hand

Let's implement the two most common transforms: standardising a numeric feature and one-hot encoding a category. Run it and watch the outputs.

# Two core transforms by hand (plain Python): scaling and one-hot encoding.

# 1) STANDARDISE a numeric feature to mean 0, standard deviation 1.
def standardize(values):
    n = len(values)
    mean = sum(values) / n
    var = sum((v - mean) ** 2 for v in values) / n
    std = var ** 0.5
    return [round((v - mean) / std, 2) for v in values]

ages = [20, 30, 40, 50]
print(standardize(ages))         # centered around 0

# 2) ONE-HOT encode an unordered category (no false ordering!).
def one_hot(value, categories):
    return [1 if value == c else 0 for c in categories]

colors = ["red", "green", "blue"]
print(one_hot("green", colors))  # [0, 1, 0]
print(one_hot("blue", colors))   # [0, 0, 1]

# Expected output:
# [-1.34, -0.45, 0.45, 1.34]
# [0, 1, 0]
# [0, 0, 1]

3 Interactions, Polynomials & Binning

Sometimes the raw features hide the pattern. Three ways to surface it:

4 Feature Selection: Filter, Wrapper, Embedded

More features isn't always better — irrelevant ones add noise and overfitting. Three families of feature selection help you keep the useful ones:

🌍 Worked Example: The Real Tool (scikit-learn)

In practice scikit-learn provides ready transforms. Study this structural reference (it isn't runnable here) and notice the golden rule: every transform is fit on the training set only, then applied to the test set.

# WORKED EXAMPLE — what the real tool looks like (scikit-learn).
# You don't run this here; study it so the library code feels familiar.
from sklearn.preprocessing import StandardScaler, OneHotEncoder, PolynomialFeatures
from sklearn.feature_selection import SelectKBest, f_classif

# Numeric features -> scale (fit on TRAIN ONLY to avoid leakage)
scaler = StandardScaler().fit(X_train_num)
X_train_num = scaler.transform(X_train_num)
X_test_num  = scaler.transform(X_test_num)   # reuse train statistics

# Categorical features -> one-hot (no false ordering)
ohe = OneHotEncoder(handle_unknown="ignore").fit(X_train_cat)

# Interaction / polynomial features so a linear model can fit curves
poly = PolynomialFeatures(degree=2, interaction_only=False)

# Filter feature selection: keep the k best by an ANOVA F-test
selector = SelectKBest(score_func=f_classif, k=10)

# Expected: each transform is FIT on train, then applied to test.
# No output here — this is a structural reference.

Wrapping these steps in a Pipeline makes the fit-on-train-only rule automatic and leakage-proof.

🎯 Your Turn 1: Min-Max Scaling

Fill in the blanks so min_max() squashes values into the 0..1 range. The expected output is in the comments.

# 🎯 YOUR TURN — finish min-max scaling (squash values into 0..1).
# formula: (v - min) / (max - min)

def min_max(values):
    lo = min(values)
    hi = max(values)
    return [round((v - ___) / (hi - ___), 2) for v in values]   # 👉 fill both blanks with lo

print(min_max([10, 20, 30, 40]))

# ✅ Expected output:
# [0.0, 0.33, 0.67, 1.0]

🎯 Your Turn 2: Bin an Age

Group a continuous age into categories. Fill in the two blanks to return the right bucket.

# 🎯 YOUR TURN — bin a continuous age into categories.
# child: < 13, teen: 13-17, adult: >= 18

def age_bin(age):
    if age < 13:
        return ___        # 👉 return "child"
    elif age < 18:
        return "teen"
    else:
        return ___        # 👉 return "adult"

for a in [8, 15, 40]:
    print(a, "->", age_bin(a))

# ✅ Expected output:
# 8 -> child
# 15 -> teen
# 40 -> adult

🎯 Mini-Challenge: Build an Interaction Feature

Create an interaction feature and show it perfectly predicts the target. Only a comment outline is provided.

# 🎯 MINI-CHALLENGE: build an interaction feature and rank by correlation
# A simple FILTER selection: keep the feature most correlated with the target.
#
# 1. price = [2, 4, 6, 8]   qty = [1, 2, 3, 4]   target = [2, 8, 18, 32]
# 2. Create interaction = [p * q for p, q in zip(price, qty)]  (=> 2,8,18,32)
# 3. Notice interaction matches target EXACTLY -> perfect signal
# 4. print() the interaction list and a message that it is the best feature
#
# ✅ Expected: interaction = [2, 8, 18, 32], identical to target

# your code here

5 Common Errors (And How to Fix Them)

These feature-engineering mistakes silently inflate your scores. Watch for them.

❌ Fitting the scaler on all the data

Computing scaling statistics over train + test leaks test information into training.

scaler.fit(X_all)   # ❌ then split -> leakage

✅ Fix: fit on train only, transform test:

scaler.fit(X_train); scaler.transform(X_test)

Mapping red/green/blue to 0/1/2 invents an order the model will misuse.

color = {"red": 0, "green": 1, "blue": 2}   # ❌ implies blue > red

✅ Fix: one-hot encode unordered categories:

OneHotEncoder().fit_transform(colors)

❌ A feature that secretly contains the answer

Including a column that is derived from the label gives unbeatable validation scores that vanish in production.

X = df[["age", "income", "was_refunded"]]   # ❌ "was_refunded" leaks the target

✅ Fix: drop anything unknown at prediction time:

X = df[["age", "income"]]   # only legitimately available features

📋 Quick Reference

ConceptWhat It IsKey Point
ScalingStandardise or min-maxNeeded for distance/gradient models
One-hot encodingOne 0/1 column per categoryFor unordered categories
Interaction featureCombine two featuresCaptures joint effects
Polynomial featuresPowers and productsLets linear models curve
BinningContinuous into rangesSoftens outliers
Filter / Wrapper / EmbeddedSelection strategiesStat / search / in-training
Data leakageTest/future info in featuresFit transforms on train only

🎉 Lesson Complete!

You can now scale and encode features correctly, craft interaction, polynomial, and binned features, choose between filter, wrapper, and embedded selection, and keep data leakage out of your pipeline.

Practice quiz

What is feature engineering?

  • Creating and transforming input features to help a model learn
  • Picking a model
  • Labelling the data
  • Tuning the learning rate

Answer: Creating and transforming input features to help a model learn. Feature engineering crafts better inputs — transforming, combining, or creating features so the model can find patterns more easily.

Which encoding turns a categorical column into one binary column per category?

  • Standard scaling
  • One-hot encoding
  • Log transform
  • Binning

Answer: One-hot encoding. One-hot encoding creates a 0/1 column for each category, which models can use without implying a false ordering.

Why is label (ordinal) encoding risky for unordered categories like colour?

  • It is too slow
  • It implies a false numeric order (e.g. red < blue < green)
  • It removes the column
  • It needs scaling

Answer: It implies a false numeric order (e.g. red < blue < green). Mapping unordered categories to 1, 2, 3 invents an order that linear and distance-based models will wrongly exploit.

What is an interaction feature?

  • A scaled feature
  • A new feature combining two others (e.g. their product) to capture joint effects
  • A removed feature
  • A binned feature

Answer: A new feature combining two others (e.g. their product) to capture joint effects. Interaction features (like price multiplied by quantity) let a model capture effects that depend on two features together.

What does binning (discretization) do?

  • Scales features to mean 0
  • Encodes text
  • Groups a continuous feature into ranges/buckets
  • Removes outliers automatically

Answer: Groups a continuous feature into ranges/buckets. Binning converts a continuous value (like age) into ranges (child/adult/senior), which can capture non-linear effects simply.

Filter, wrapper, and embedded are three families of what?

  • Scaling methods
  • Feature selection methods
  • Encoding methods
  • Loss functions

Answer: Feature selection methods. They are the three feature-selection strategies: filter (statistics), wrapper (search with a model), embedded (selection during training).

A wrapper feature-selection method (like recursive feature elimination) works by:

  • Ranking features by a quick statistic
  • Repeatedly training a model on feature subsets to see which help
  • Using L1 penalties during training
  • One-hot encoding everything

Answer: Repeatedly training a model on feature subsets to see which help. Wrapper methods search over feature subsets, training the model each time — accurate but computationally expensive.

An embedded method such as Lasso (L1) selects features by:

  • Ranking with correlation only
  • Shrinking some coefficients to exactly zero during training
  • Trying every subset
  • Binning the target

Answer: Shrinking some coefficients to exactly zero during training. Lasso's L1 penalty drives weak features' coefficients to zero as part of fitting, doing selection inside the model.

What is data leakage in feature engineering?

  • Losing rows of data
  • Information from the test set or future sneaking into the features
  • Too many features
  • Slow training

Answer: Information from the test set or future sneaking into the features. Leakage is when info you wouldn't have at prediction time (or test data) bleeds into training, giving falsely high scores.

To avoid leakage when scaling, you should fit the scaler on:

  • The whole dataset before splitting
  • The test set
  • Nothing — never scale
  • The training set only, then transform the test set

Answer: The training set only, then transform the test set. Fit the scaler (and other transforms) on training data only; applying training statistics to the test set prevents leakage.

Continue this course

Frequently asked questions

What is feature engineering and why does it matter so much?

Feature engineering is the craft of turning raw data into inputs a model can learn from well. That includes scaling numbers, encoding categories, combining columns, and creating new signals. It often matters more than the choice of algorithm: good features can make a simple model shine, while bad features hold back even the fanciest one.

When do I need to scale features, and how?

Distance- and gradient-based models (k-NN, SVM, linear/logistic regression, neural networks, PCA, k-means) need scaling so no single large-range feature dominates. Standardisation (mean 0, standard deviation 1) and min-max scaling (squash to 0..1) are the two common choices. Tree-based models (decision trees, random forests, gradient boosting) do not need scaling.

How should I encode categorical features?

For unordered categories (colour, city), use one-hot encoding — one 0/1 column per category — so the model never assumes a false order. For genuinely ordered categories (small < medium < large) ordinal encoding is fine. Avoid plain label encoding on unordered categories, because mapping them to 1, 2, 3 invents an order the model will misuse.

What are interaction, polynomial, and binning features?

An interaction feature combines two columns (often by multiplying them) to capture joint effects, like price times quantity. Polynomial features add powers and products of features so a linear model can fit curves. Binning groups a continuous feature into ranges (age into child/adult/senior), which can capture non-linear effects and reduce the impact of outliers.

What is the difference between filter, wrapper, and embedded feature selection?

Filter methods rank features with a quick statistic (correlation, mutual information) independent of any model — fast but ignores feature interactions. Wrapper methods (like recursive feature elimination) repeatedly train a model on different subsets to see which help — accurate but slow. Embedded methods do selection during training, like Lasso (L1) shrinking weak coefficients to zero or a tree's feature importances.

What is data leakage and how do I prevent it?

Leakage is when information you would not have at prediction time — including anything computed from the test set — sneaks into your features. It produces great validation scores that collapse in production. Prevent it by splitting the data first, fitting every transform (scalers, encoders, imputers) on the training set only, and dropping any feature that secretly encodes the target.