Reinforcement Learning

Reviewed & published by Brayan K

Teach an agent to make good decisions through trial and error. By the end you'll understand the RL loop, write an epsilon-greedy agent in plain Python, and know where RL powers real systems — from game AI to robots.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

🐶 Real-World Analogy: Training a Dog with Treats

You can't hand a puppy a textbook. Instead you say "sit", wait, and the moment it sits you give a treat. The treat is a reward. The puppy doesn't know why the treat came at first — but after many tries it links the situation ("I heard 'sit'") and the action ("I sat") to the reward.

That is reinforcement learning exactly. The puppy is the agent. Your living room is the environment. "I just heard the word sit" is the state. Sitting, lying down, or running off are the actions. The treat (or no treat) is the reward. Over time the puppy forms a policy: "when I hear 'sit', sitting pays off."

Notice there were no labelled examples — nobody showed the puppy ten thousand photos of correct sitting. It learned purely from interaction and feedback. That's what separates RL from supervised learning.

1 The Five Pieces: Agent, Environment, State, Action, Reward

Every RL problem is built from the same five pieces. Learn these names once and the rest of the field reads easily.

Jargon check: a policy is the agent's strategy — a rule mapping each state to an action. The agent's job is to improve its policy until it reliably collects high reward.

2 The RL Loop (Worked Example)

RL runs in a loop: the agent observes a state, picks an action, the environment returns a new state and a reward, and round it goes. One full run from start to finish is called an episode.

Read this fully-commented example. It builds a tiny 5-cell corridor by hand and walks an agent to the goal so you can see every part of the loop with your own eyes. Run it and match the output.

# The reinforcement learning loop, spelled out step by step.
# No frameworks — just plain Python so you can SEE every part.

# The "environment": a tiny 1D corridor of 5 cells, 0..4.
# The agent starts at cell 0. The treasure is at cell 4.
GOAL = 4

def reward_for(state):
    # +10 when you reach the treasure, -1 for every other step
    # (the -1 nudges the agent toward the SHORTEST path).
    return 10 if state == GOAL else -1

def step(state, action):
    # An action is "right" (+1) or "left" (-1). Walls stop you.
    new_state = state + action
    if new_state < 0 or new_state > GOAL:
        new_state = state            # bumped a wall, stay put
    return new_state

# One episode: agent always walks right (a fixed, hand-written policy).
state = 0                            # the current STATE
total = 0                            # the cumulative REWARD
print("start at cell", state)

for t in range(5):
    action = +1                      # the ACTION the policy chose
    new_state = step(state, action)  # the ENVIRONMENT responds
    r = reward_for(new_state)        # the REWARD signal
    total += r
    print(f"step {t}: action=right -> cell {new_state}, reward {r}")
    state = new_state
    if state == GOAL:                # episode ends at the goal
        break

print("total reward:", total)

# Expected output:
# start at cell 0
# step 0: action=right -> cell 1, reward -1
# step 1: action=right -> cell 2, reward -1
# step 2: action=right -> cell 3, reward -1
# step 3: action=right -> cell 4, reward 10
# total reward: 7

The -1 per step plus +10 at the goal is called reward shaping — designing the numbers so the agent prefers short paths. Get the shaping wrong and the agent learns the wrong thing.

3 Policy vs Value — Two Things the Agent Learns

There are two quantities an RL agent can learn, and beginners mix them up constantly:

Policy — "what should I do?"

A rule that maps a state to an action. Ask it "I'm in state X, what now?" and it answers with an action. The puppy's "hear sit → sit" is a policy.

Value — "how good is this?"

A number estimating the total future reward you expect from a state (or a state-action pair). It doesn't tell you what to do directly — it scores options.

Many algorithms learn values first, then act greedily on them: "estimate how good each action is, then pick the highest." That is exactly what the bandit below does — it keeps a value estimate per arm and exploits the best one.

4 Exploration vs Exploitation (Worked Example)

Here is the central dilemma of RL. Exploitation means choosing the action you currently think is best. Exploration means trying a different action to find out if something better exists. Lean too far either way and you lose.

The classic toy problem is the multi-armed bandit: several slot machine arms, each paying out with a hidden probability. The epsilon-greedy rule solves it simply — with probability epsilon (say 0.1) explore a random arm, otherwise exploit the arm with the highest current estimate.

import random

# Multi-armed bandit: 3 slot machines ("arms").
# Each arm pays 1 with a hidden probability. The agent must DISCOVER
# the best arm using only the rewards it sees.
random.seed(7)                       # deterministic so the output is fixed

true_probs = [0.2, 0.5, 0.8]         # hidden! the agent never reads this
n_arms = len(true_probs)

def pull(arm):
    # Returns 1 (win) or 0 (lose) for the chosen arm.
    return 1 if random.random() < true_probs[arm] else 0

# What the agent KNOWS: how good it currently estimates each arm to be,
# and how many times it has tried each arm.
estimates = [0.0, 0.0, 0.0]          # the VALUE estimate per arm
counts    = [0, 0, 0]

epsilon = 0.1                        # 10% of the time: EXPLORE at random
total_reward = 0

for t in range(2000):
    if random.random() < epsilon:
        arm = random.randrange(n_arms)        # EXPLORE: try something
    else:
        best = max(estimates)                 # EXPLOIT: pick current best
        arm = estimates.index(best)

    reward = pull(arm)
    total_reward += reward

    # Update the running average estimate for that arm (incremental mean).
    counts[arm] += 1
    estimates[arm] += (reward - estimates[arm]) / counts[arm]

print("estimated arm values:", [round(e, 2) for e in estimates])
print("times each arm was pulled:", counts)
print("agent's best arm:", estimates.index(max(estimates)))
print("total reward:", total_reward)

# Expected output:
# estimated arm values: [0.31, 0.52, 0.8]
# times each arm was pulled: [71, 67, 1862]
# agent's best arm: 2
# total reward: 1545

Because most pulls exploit the best-known arm, arm 2 (the true best) gets pulled thousands of times — but the occasional 10% exploration is what let the agent find arm 2 in the first place. That balance is the whole game.

5 The Markov Decision Process (MDP)

The Markov Decision Process is the formal frame underneath every RL problem. Don't let the name scare you — it's just the five pieces written precisely:

The "Markov" property is the key simplifying assumption: the next state depends only on the current state and action — not on the entire history of how you got there. That makes the maths tractable. Your corridor example was a small MDP, and so was the bandit (a one-state MDP).

🌍 In the Real World: Gymnasium

You built tiny environments by hand to learn the mechanics. In real projects you reach for Gymnasium (the maintained successor to OpenAI Gym), which provides ready-made environments behind a standard reset() / step() interface — the very same loop you wrote.

This block needs pip install gymnasium, so run it on your own machine rather than in the sandbox above. Read it to see how the hand-built loop maps onto a real library.

# In a real project you don't hand-build the environment — you use
# Gymnasium (the maintained successor to OpenAI Gym). It gives you
# ready-made worlds with a standard reset()/step() interface.
#
#   pip install gymnasium
import gymnasium as gym

env = gym.make("FrozenLake-v1", is_slippery=False)

state, info = env.reset(seed=0)      # start a new episode
done = False
total = 0

while not done:
    action = env.action_space.sample()        # random policy (placeholder)
    state, reward, terminated, truncated, info = env.step(action)
    total += reward
    done = terminated or truncated

print("episode finished, total reward:", total)

# Expected output (random actions rarely reach the goal):
# episode finished, total reward: 0.0
#
# A trained agent would learn a policy that reaches the goal for reward 1.0.
# Gymnasium gives you the SAME state/action/reward loop you built by hand
# above — it just supplies the environment for you.

🎯 Your Turn #1: Exploit the Best Arm

The explore branch is written for you. Fill in the exploit line so the agent picks the arm with the highest estimate. Match the expected output in the comment.

# 🎯 YOUR TURN — fill in the blanks marked with ___
import random
random.seed(1)

# Hidden true probabilities for 3 arms (the agent can't read these).
true_probs = [0.3, 0.9, 0.6]

def pull(arm):
    return 1 if random.random() < true_probs[arm] else 0

estimates = [0.0, 0.0, 0.0]
counts    = [0, 0, 0]
epsilon   = 0.1

for t in range(1500):
    if random.random() < epsilon:
        arm = random.randrange(3)              # explore
    else:
        # 👉 EXPLOIT: choose the arm with the HIGHEST estimate.
        #    Hint: max(estimates) is the best value;
        #    estimates.index(...) gives its position.
        arm = ___                               # 👉 replace ___
    reward = pull(arm)
    counts[arm] += 1
    estimates[arm] += (reward - estimates[arm]) / counts[arm]

print("best arm:", estimates.index(max(estimates)))

# ✅ Expected output:
# best arm: 1

🎯 Your Turn #2: Make the Epsilon-Greedy Choice

Now write both halves of the decision: the explore return (a random arm) and the exploit return (the best arm). The hints are right there in the comments.

# 🎯 YOUR TURN — fill in the blanks marked with ___
import random
random.seed(2)

estimates = [0.1, 0.7, 0.4]     # pretend the agent already learned these
epsilon   = 0.2                 # explore 20% of the time

def choose_action():
    if random.random() < epsilon:
        # 👉 EXPLORE: return a RANDOM arm number from 0, 1, 2.
        return ___              # 👉 replace ___  (hint: random.randrange(3))
    else:
        # 👉 EXPLOIT: return the arm with the highest estimate.
        return ___              # 👉 replace ___  (hint: estimates.index(max(estimates)))

picks = [choose_action() for _ in range(10)]
print("actions chosen:", picks)
print("mostly the best arm (1)?", picks.count(1) >= 7)

# ✅ Expected output:
# actions chosen: [1, 1, 0, 1, 2, 1, 1, 1, 2, 1]
# mostly the best arm (1)? True

🏆 Mini-Challenge: A 4-Arm Bandit from Scratch

Support is fading now. You get only a comment outline — write the epsilon-greedy agent yourself. Use the two worked examples above as your reference.

import random
random.seed(0)

# 🎯 MINI-CHALLENGE: epsilon-greedy on a 4-arm bandit
# 1. true_probs = [0.1, 0.25, 0.6, 0.45]  (hidden payouts; arm 2 is best)
# 2. Write pull(arm): return 1 if random.random() < true_probs[arm] else 0
# 3. Keep estimates = [0, 0, 0, 0] and counts = [0, 0, 0, 0]
# 4. Loop 3000 times with epsilon = 0.1:
#       - explore (random arm) with probability epsilon, else exploit best
#       - update counts and estimates with the incremental mean
# 5. Print the agent's best arm and the total reward
#
# ✅ Expected: best arm is 2, and total reward is roughly 1600+

# your code here

Common Mistakes (And How to Fix Them)

❌ No exploration (epsilon = 0)

A pure-greedy agent locks onto whichever arm happened to win first and never discovers the truly best one.

✅ Fix: keep a small exploration rate (e.g. epsilon = 0.1), or decay it from high to low over time so you explore early and exploit later.

If you reward the agent for the wrong thing, it optimises the wrong thing. A vacuum bot rewarded for "dust collected" learns to dump dust and re-vacuum it.

✅ Fix: reward the outcome you actually want, not a proxy. Add small penalties (like -1 per step) to discourage stalling.

If reward only arrives at a distant goal and is zero everywhere else, the agent wanders randomly and almost never stumbles onto it, so it never learns.

✅ Fix: add intermediate rewards (shaping), shrink the environment while learning, or use exploration bonuses to encourage reaching new states.

If the environment changes over time (payouts drift, an opponent adapts), an estimate built as a plain average of all past rewards reacts too slowly.

✅ Fix: use a fixed step size — estimate += alpha * (reward - estimate) — so recent experience counts more than ancient experience.

📋 Quick Reference

TermWhat It MeansIn the Examples
AgentThe learner that chooses actionsThe bandit player
EnvironmentReturns next state + rewardThe corridor / slot machines
StateThe current situationstate = the cell number
ActionA choice available nowMove right / pull an arm
RewardNumeric feedback to maximise+10 goal, -1 step, win/lose
PolicyState → action strategy"pick the best-estimated arm"
ValueExpected future rewardestimates[arm]
Epsilon-greedyExplore ε of the time, else exploitrandom.random() < epsilon
MDPStates, actions, transitions, rewards, γThe whole formal frame

Pro tip: ChatGPT and other chat models are fine-tuned with Reinforcement Learning from Human Feedback (RLHF). Humans rank responses, and RL nudges the model toward the answers people prefer — the exact same reward-maximising loop you learned here, just with a language model as the agent.

🎉 Lesson Complete!

You now know the language of reinforcement learning: an agent acts in an environment, observing states, choosing actions, and collecting rewards in a loop. You can explain policy vs value, balance exploration vs exploitation with epsilon-greedy, frame a problem as a Markov Decision Process, and name where RL is used.

Practice quiz

What is reinforcement learning in simple terms?

  • Learning from labelled examples
  • Memorising a fixed dataset
  • Learning by trial and error: an agent acts, receives rewards, and learns a policy that maximises total reward
  • Compressing data

Answer: Learning by trial and error: an agent acts, receives rewards, and learns a policy that maximises total reward. RL learns from interaction and a reward signal, not from labelled correct answers.

How does RL differ from supervised learning?

  • RL has no labels — it only gets a reward signal after acting and must discover good actions itself
  • RL uses labelled examples; supervised does not
  • They are the same
  • RL never uses feedback

Answer: RL has no labels — it only gets a reward signal after acting and must discover good actions itself. Supervised learning is told the right answer; RL only sees rewards and figures out which actions are good.

Which are the five core pieces of an RL problem?

  • Input, hidden, output, loss, optimizer
  • Train, test, validate, score, deploy
  • Chunk, embed, store, retrieve, generate
  • Agent, environment, state, action, reward

Answer: Agent, environment, state, action, reward. Every RL problem is built from agent, environment, state, action, and reward.

What is a policy?

  • A number scoring a state
  • The agent's strategy — a rule mapping each state to an action
  • The reward signal
  • The environment's transition table

Answer: The agent's strategy — a rule mapping each state to an action. A policy tells the agent what action to take in each state.

What is the difference between a policy and a value?

  • A policy says what to do; a value estimates the expected future reward of a state or action
  • A value says what to do; a policy is a reward
  • They are identical
  • A policy is always a single number

Answer: A policy says what to do; a value estimates the expected future reward of a state or action. Policy answers 'what should I do?'; value answers 'how good is this?' (expected future reward).

What is the exploration vs exploitation trade-off?

  • Choosing between two models
  • Splitting data into train and test
  • Exploitation picks the action you think is best; exploration tries others to find something better
  • Tuning the learning rate

Answer: Exploitation picks the action you think is best; exploration tries others to find something better. Too little exploration gets stuck on a mediocre choice; too much wastes reward. Epsilon-greedy balances them.

How does the epsilon-greedy rule choose actions?

  • Always pick the best-estimated action
  • With probability epsilon pick a random action, otherwise pick the highest-estimated action
  • Always pick randomly
  • Pick the lowest-estimated action

Answer: With probability epsilon pick a random action, otherwise pick the highest-estimated action. Epsilon controls how often the agent explores instead of exploiting the current best estimate.

In a multi-armed bandit, what is the agent trying to discover?

  • The number of arms
  • The colour of each arm
  • The exact payout formula given to it
  • Which arm has the highest hidden payout probability, using only the rewards it observes

Answer: Which arm has the highest hidden payout probability, using only the rewards it observes. The agent never reads the true probabilities; it estimates each arm's value from observed rewards.

What is a Markov Decision Process (MDP)?

  • A neural network architecture
  • The formal frame for RL: states, actions, transition probabilities, rewards, and a discount
  • A type of reward only
  • A clustering algorithm

Answer: The formal frame for RL: states, actions, transition probabilities, rewards, and a discount. An MDP defines states, actions, transitions, rewards, and gamma — the formal structure under every RL problem.

What does the 'Markov' property mean?

  • The agent remembers every past step
  • Rewards are always positive
  • The next state depends only on the current state and action, not the full history
  • The environment never changes

Answer: The next state depends only on the current state and action, not the full history. The Markov property says the future depends only on the present state and action, which makes the maths tractable.

Continue this course

Frequently asked questions

What is reinforcement learning in simple terms?

It is learning by trial and error. An agent takes actions in an environment, receives rewards or penalties, and gradually learns a policy — a strategy for choosing actions — that maximises its total reward over time.

How is reinforcement learning different from supervised learning?

Supervised learning trains on labelled examples that tell it the correct answer. Reinforcement learning has no labels — the agent only gets a reward signal after acting, and must figure out for itself which actions lead to good outcomes.

What is the exploration vs exploitation trade-off?

Exploitation means picking the action you currently believe is best; exploration means trying other actions to discover something better. If you never explore you can get stuck on a mediocre choice; if you always explore you waste reward. Epsilon-greedy balances the two.

What is a Markov Decision Process (MDP)?

An MDP is the mathematical frame for RL. It is defined by states, actions, a reward for each transition, and transition probabilities. The 'Markov' part means the next state depends only on the current state and action — not the full history.

Where is reinforcement learning actually used?

Game-playing AI (AlphaGo, Atari, chess), robotics and locomotion control, recommendation and ad systems, traffic and energy optimisation, and fine-tuning large language models with human feedback (RLHF).