Ethical AI & Bias Mitigation

Reviewed & published by Brayan K

By the end of this lesson you'll be able to measure whether a model treats groups fairly, explain its decisions, protect people's privacy, and write the audit code that catches bias before it ships.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

Real-World Analogy: a fair referee

Think of an AI model as a referee in a football match. A fair referee applies the same rules to both teams: a foul is a foul whoever commits it, and a goal counts no matter which side scores. The crowd can see every decision, so the referee can be questioned and overruled.

A biased referee blows the whistle more often on one team — maybe without even realising it, because that's how they were trained. AI models do exactly this: they absorb bias from historical data and apply it at massive scale, silently. Everything in this lesson is a tool for being a fair referee: measure the calls per team (fairness metrics), show your reasoning (explainability), respect the players (privacy), and let yourself be reviewed (accountability).

1 Where Does Bias Enter?

Bias is a systematic error that pushes a model's outputs in one direction for one group. It rarely comes from one villain — it leaks in at several points, and your job is to know all of them.

Real systems have caused real harm by missing these:

CaseWhat HappenedWhere bias entered
Amazon hiringRésumé screener penalised womenBiased historical labels
COMPASRecidivism model harsher on Black defendantsProxy variables + labels
Healthcare algorithmUsed cost as a proxy for needBad objective / proxy
Facial recognition35% error for dark-skinned women vs 1% for light-skinned menUnbalanced sampling

2 Measuring Fairness — Parity, the 4/5ths Rule & Equalised Odds

You can't fix what you don't measure. The simplest fairness metric is demographic parity: the selection rate (the share of a group that gets a positive outcome) should be roughly equal across groups.

Regulators turn that into a number with the 4/5ths rule: divide the lowest group's selection rate by the highest to get the disparate-impact ratio. If it falls below 0.8, that is treated as adverse impact you must investigate.

Parity alone can be misleading, so also check equalised odds: among people who actually deserved a positive outcome, the model's success rate should be the same for every group. The worked example below computes selection rates per group and the disparate-impact ratio from a tiny dataset — read every comment, then run it.

# ============================================
# WORKED EXAMPLE: where does bias hide?
# Pure Python — no libraries needed.
# ============================================

# A loan model already made decisions on 12 applicants.
# Each row: (group, approved?)  approved = 1, denied = 0
applicants = [
    ("A", 1), ("A", 1), ("A", 1), ("A", 0),
    ("A", 1), ("A", 1),                        # Group A: 5 of 6 approved
    ("B", 1), ("B", 0), ("B", 0), ("B", 0),
    ("B", 1), ("B", 0),                        # Group B: 2 of 6 approved
]

# Selection rate = approvals / total, computed PER GROUP.
def selection_rate(rows, group):
    members  = [r for r in rows if r[0] == group]
    approved = sum(approved for _, approved in members)
    return approved / len(members)      # a fraction between 0 and 1

rate_a = selection_rate(applicants, "A")
rate_b = selection_rate(applicants, "B")

print(f"Group A approval rate: {rate_a:.0%}")   # 83%
print(f"Group B approval rate: {rate_b:.0%}")   # 33%

# Demographic parity asks: are these two rates equal?
# Disparate-impact ratio = lower rate / higher rate.
ratio = min(rate_a, rate_b) / max(rate_a, rate_b)
print(f"Disparate-impact ratio: {ratio:.2f}")   # 0.40

# The 4/5ths rule (US EEOC): a ratio below 0.8 is adverse impact.
if ratio < 0.8:
    print("FAIL: ratio < 0.8 -> adverse impact, investigate the model")
else:
    print("PASS: groups are selected at similar rates")

# Expected output:
# Group A approval rate: 83%
# Group B approval rate: 33%
# Disparate-impact ratio: 0.40
# FAIL: ratio < 0.8 -> adverse impact, investigate the model
# 🎯 YOUR TURN — measure fairness on a hiring model.
# Fill in every ___ then run it.

# Each row: (group, hired?)   hired = 1, not hired = 0
people = [
    ("women", 0), ("women", 1), ("women", 0), ("women", 0), ("women", 1),
    ("men",   1), ("men",   1), ("men",   1), ("men",   0), ("men",   1),
]

def selection_rate(rows, group):
    members = [r for r in rows if r[0] == group]
    hired   = sum(h for _, h in members)
    # 👉 return the fraction hired (hired divided by number of members)
    return ___

women_rate = selection_rate(people, "women")
men_rate   = selection_rate(people, "men")
print(f"Women hired: {women_rate:.0%}")
print(f"Men hired:   {men_rate:.0%}")

# 👉 disparate-impact ratio = lower rate / higher rate
ratio = ___ / ___
print(f"Disparate-impact ratio: {ratio:.2f}")

# 👉 the 4/5ths rule: flag if the ratio is below 0.8
if ___:
    print("FAIL: adverse impact")
else:
    print("PASS")

# ✅ Expected output:
# Women hired: 40%
# Men hired:   80%
# Disparate-impact ratio: 0.50
# FAIL: adverse impact

3 Transparency & Explainability (SHAP / LIME)

If a model denies someone a loan, they deserve to know why. Explainability tools answer that by attributing one prediction to the features that drove it: "+0.30 from income, −0.90 from late payments". That turns a black box into a sentence a human can challenge.

SHAP uses game theory to split a prediction fairly among its features and is consistent across the whole model. LIME builds a quick, simple approximation around a single prediction. The worked example shows the core idea — every feature's contribution is its value times its weight — and points to the real libraries.

# ============================================
# WORKED EXAMPLE: explainability (the SHAP / LIME idea)
# Why did the model deny THIS person? Attribute the score
# to each feature. Pure Python — the real tools just do this
# rigorously across the whole model.
# ============================================

# A simple, fully transparent scoring model.
# Each feature has a weight; score = sum(feature * weight).
weights = {"income": 0.6, "age": -0.2, "late_payments": -0.9}

applicant = {"income": 4, "age": 2, "late_payments": 3}

# Contribution of each feature = value * weight (this is the SHAP idea).
contributions = {f: applicant[f] * w for f, w in weights.items()}
score = sum(contributions.values())

print(f"Final score: {score:+.2f}")
for feature, contrib in contributions.items():
    arrow = "pushes APPROVE" if contrib > 0 else "pushes DENY"
    print(f"  {feature:<13}: {contrib:+.2f}  ({arrow})")

# Reading this: late_payments dominated the decision, not age.
# That is exactly what SHAP / LIME report for real black-box models,
# turning an opaque prediction into a sentence a human can challenge.

# In production you would import the real libraries:
#   import shap                 # game-theory feature attributions
#   explainer = shap.Explainer(model)
#   shap_values = explainer(X)  # per-prediction contributions
# Expected output:
# Final score: -0.70
#   income       : +2.40  (pushes APPROVE)
#   age          : -0.40  (pushes DENY)
#   late_payments: -2.70  (pushes DENY)

4 Privacy — PII & Differential Privacy

PII (personally identifiable information) is anything that pins data to a real person — name, email, exact address. The first rule is simple: strip direct identifiers before the data reaches the model.

But redaction is not enough — clever attackers can re-identify people from "anonymous" patterns. Differential privacy fixes this by adding calibrated random noise so that adding or removing any single person barely changes the result. The knob is the privacy budget epsilon (ε): lower ε means more privacy and less accuracy. The worked example redacts PII and applies the differential-privacy idea to a salary average.

# ============================================
# WORKED EXAMPLE: privacy — PII and differential privacy
# Pure Python. The real tools (below) automate the maths.
# ============================================
import random
random.seed(0)

# Step 1: never train on raw PII (personally identifiable info).
record = {"name": "Jordan Lee", "email": "[email protected]", "salary": 52000}
# Redact direct identifiers BEFORE the data reaches the model.
safe = {k: v for k, v in record.items() if k not in ("name", "email")}
print("Redacted record:", safe)        # {'salary': 52000}

# Step 2: differential privacy — answer questions about a group
# without exposing any one person. Add calibrated random noise so
# that adding/removing one record barely changes the answer.
salaries = [52000, 61000, 48000, 75000, 50000]
true_mean = sum(salaries) / len(salaries)

epsilon = 1.0                          # privacy budget: lower = more private
sensitivity = 30000 / len(salaries)    # how much one person can shift the mean
noise = random.gauss(0, sensitivity / epsilon)
private_mean = true_mean + noise

print(f"True mean salary:    {true_mean:,.0f}")
print(f"Private mean salary: {private_mean:,.0f}")
print("The reported number is close, but no single salary can be recovered.")

# In production you would lean on real libraries:
#   from opacus import PrivacyEngine        # DP-SGD for PyTorch
#   from diffprivlib.models import LogisticRegression  # DP scikit-learn
# Expected output (values vary because of the random noise):
# Redacted record: {'salary': 52000}
# True mean salary:    57,200
# Private mean salary: 5X,XXX
# The reported number is close, but no single salary can be recovered.
# 🎯 YOUR TURN — catch a proxy variable.
# We removed "gender", but a proxy may still leak it.
# Fill in every ___ then run it.

applicants = [
    # (club_member, gender, hired?)  "club_member" is the suspect proxy
    (1, "M", 1), (1, "M", 1), (1, "F", 1), (1, "F", 0),
    (0, "M", 0), (0, "M", 1), (0, "F", 0), (0, "F", 0),
]

# A proxy is dangerous when it strongly predicts a protected attribute.
# Check: what share of club members are male?
members = [a for a in applicants if a[0] == 1]
# 👉 count how many members have gender "M"
male_members = sum(1 for a in members if a[1] == ___)
share_male = male_members / len(members)
print(f"Share of club members who are male: {share_male:.0%}")

# 👉 a feature is a risky proxy if this share is far from 50% (say > 0.7)
is_proxy = ___ > 0.7
print(f"'club_member' is a proxy for gender: {is_proxy}")

# ✅ Expected output:
# Share of club members who are male: 50%
# 'club_member' is a proxy for gender: False

5 Accountability & Harmful-Misuse Risk

Accountability means a specific, named human owns the model's outcomes — not "the algorithm". When a decision is wrong, there must be a person to appeal to, a log of what the model did, and a clear path to override it. "The model said so" is never an acceptable answer.

Concretely, an accountable system has:

You also have to think about misuse — harm caused by using the model outside its intended purpose. A face-matcher built for unlocking phones could be repurposed for mass surveillance; a text generator could mass-produce scams or disinformation. Mitigation means restricting access, rate-limiting, watermarking outputs, refusing dangerous requests, and red-teaming the system before release.

6 Common Mistakes (And How to Fix Them)

❌ Training on biased data and trusting it

Feeding in years of skewed historical decisions teaches the model to copy them.

✅ Fix: audit the data first — check per-group base rates and rebalance or reweight before training.

❌ Optimising only for accuracy

A 95%-accurate model can still fail a minority group completely, because accuracy averages over everyone.

✅ Fix: track per-group metrics (selection rate, true-positive rate) alongside overall accuracy.

❌ Shipping with no audit

"It passed validation" is not the same as "it's fair and safe in the real world".

✅ Fix: run a fairness audit (the 4/5ths rule), write a model card, and re-audit on a schedule after launch.

❌ Dropping a protected column and assuming you're done

Removing "gender" does nothing if a proxy variable (name, ZIP code, club membership) still encodes it.

✅ Fix: test remaining features for correlation with the protected attribute, then drop or decorrelate the proxies too.

📋 Quick Reference — Ethical AI

ConceptWhat it means
Demographic parityEqual selection rates across groups
Equalised oddsEqual true-positive (and false-positive) rates across groups
Disparate impact / 4/5ths rulemin(rate) / max(rate); below 0.8 = adverse impact
Proxy variableA feature that secretly encodes a protected attribute
SHAP / LIMEAttribute one prediction to its features (explainability)
PIIData that identifies a real person; redact before training
Differential privacy (ε)Add noise so no individual is recoverable; lower ε = more privacy
AccountabilityA named owner, audit trail, and human override

🎯 Mini-Challenge: Build a Fairness Auditor

Time to fly solo. Using only what you've learned, write a small program that audits a model's decisions for two regions and flags adverse impact. The starter block has the brief and the expected output — the logic is up to you.

# 🎯 MINI-CHALLENGE: build a fairness auditor
#
# You are given decisions from a model. Audit it for fairness.
#
# decisions = [
#     ("north", 1), ("north", 1), ("north", 0), ("north", 1),
#     ("south", 0), ("south", 0), ("south", 1), ("south", 0),
# ]
#
# 1. Write selection_rate(rows, group) -> approvals / total for that group
# 2. Compute the rate for "north" and "south"
# 3. Compute the disparate-impact ratio = lower rate / higher rate
# 4. Print each rate as a percentage and the ratio to 2 decimals
# 5. Print "ADVERSE IMPACT" if the ratio is below 0.8, else "OK"
#
# ✅ Expected output:
# north: 75%
# south: 25%
# ratio: 0.33
# ADVERSE IMPACT

# your code here

🎉 Lesson Complete!

You can now spot where bias enters a system, measure it with demographic parity, equalised odds, and the disparate-impact ratio, explain a single prediction the SHAP/LIME way, protect privacy with PII redaction and differential privacy, and name who's accountable when things go wrong. That is the toolkit of a fair referee.

Practice quiz

What is bias in an AI model?

  • A systematic error that pushes outputs in one direction for a group
  • A random error that averages out
  • The model's learning rate
  • The number of features used

Answer: A systematic error that pushes outputs in one direction for a group. Bias is a systematic error — for example consistently scoring one group lower than another.

What does demographic parity require?

  • Equal accuracy on the training set
  • Roughly equal selection rates across groups
  • Equal numbers of features per group
  • That protected attributes are used directly

Answer: Roughly equal selection rates across groups. Demographic parity asks that the selection rate (share given a positive outcome) be roughly equal across groups.

How is the disparate-impact ratio computed?

  • Highest selection rate divided by lowest
  • Lowest selection rate divided by highest
  • Total approvals divided by total applicants
  • Accuracy divided by precision

Answer: Lowest selection rate divided by highest. The disparate-impact ratio is the lowest group's selection rate divided by the highest group's rate.

Under the four-fifths (80%) rule, what ratio signals adverse impact?

  • Below 0.8
  • Above 0.8
  • Exactly 1.0
  • Above 0.5

Answer: Below 0.8. A disparate-impact ratio below 0.8 is treated as evidence of adverse impact that must be justified or fixed.

What is a proxy variable?

  • A feature that secretly stands in for a protected attribute
  • A variable that holds the model's accuracy
  • A backup copy of the dataset
  • A feature with no predictive power

Answer: A feature that secretly stands in for a protected attribute. A proxy (like ZIP code for race) secretly encodes a protected attribute, so dropping the protected column alone does not remove the bias.

What do SHAP and LIME do?

  • Speed up training on GPUs
  • Attribute a single prediction to the features that drove it
  • Add noise for privacy
  • Balance the dataset across groups

Answer: Attribute a single prediction to the features that drove it. SHAP and LIME are explainability tools that attribute a prediction to its features, turning a black box into something auditable.

What is differential privacy?

  • Encrypting the model weights
  • Adding calibrated noise so no single person's data can be re-identified
  • Deleting all numeric columns
  • Training a separate model per user

Answer: Adding calibrated noise so no single person's data can be re-identified. Differential privacy adds calibrated random noise so adding or removing any one person barely changes the output.

In differential privacy, what does a lower epsilon (ε) mean?

  • More privacy and less accuracy
  • Less privacy and more accuracy
  • No effect on the trade-off
  • Higher learning rate

Answer: More privacy and less accuracy. The privacy budget epsilon controls the trade-off: lower ε means more privacy but noisier, less accurate results.

Why is optimising only for accuracy a problem for fairness?

  • Accuracy is impossible to measure
  • Accuracy averages over everyone and can hide poor performance on small groups
  • Accuracy always favours the minority group
  • Accuracy requires protected attributes

Answer: Accuracy averages over everyone and can hide poor performance on small groups. A high overall accuracy can hide that a minority group is served badly, so you must track per-group metrics too.

What does accountability for an AI system require?

  • That the model is fully autonomous with no oversight
  • A named owner, an audit trail, and a human who can override decisions
  • That you delete all logs after deployment
  • That accuracy is above 99%

Answer: A named owner, an audit trail, and a human who can override decisions. Accountability means a specific person owns outcomes, with logs and a human-in-the-loop able to review and reverse decisions.

Continue this course

Frequently asked questions

What is the difference between bias and fairness in AI?

Bias is a systematic error that pushes a model's outputs in one direction — for example, consistently scoring one group lower. Fairness is the goal: making sure those errors do not fall harder on some groups than others. You measure fairness with metrics like demographic parity (equal selection rates) and equalised odds (equal accuracy for qualified people across groups).

What is the four-fifths (80%) rule?

It is a US EEOC guideline for spotting adverse impact. Divide the selection rate of the lowest-selected group by the highest. If that disparate-impact ratio is below 0.8, regulators treat it as evidence of discrimination that you must justify or fix. It is a screening test, not proof on its own.

What is a proxy variable and why is it dangerous?

A proxy is a feature that secretly stands in for a protected attribute. ZIP code can encode race; first name can encode gender. Dropping the protected column does not help if proxies remain — the model just learns the same bias through the back door. You have to test for it, not assume removal worked.

What do SHAP and LIME do?

Both are explainability tools that answer 'why did the model decide this?'. They attribute a single prediction to the features that pushed it up or down (for example, '+0.30 from income, -0.15 from age'). SHAP is grounded in game theory and is consistent across the whole model; LIME builds a quick local approximation around one prediction. They turn a black box into something you can audit and challenge.

What is differential privacy in one sentence?

It adds carefully calibrated random noise to results so that adding or removing any single person's data barely changes the output — meaning no individual can be re-identified from what the model releases. The privacy budget epsilon (ε) controls the trade-off: lower ε means more privacy and less accuracy.

Why is optimising only for accuracy a problem?

A model can hit 95% accuracy while being deeply unfair, because accuracy averages over everyone and hides what happens to small groups. If 90% of your data is Group A, the model can ignore Group B entirely and still look great. Always track per-group metrics alongside overall accuracy.