Policy Gradient Methods
Reviewed & published by Brayan K
By the end of this lesson you'll understand how an agent can learn what to do directly — turning action preferences into probabilities, sampling actions, and nudging the policy toward higher reward.
Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
What You'll Learn in This Lesson
- How policy-based RL differs from value-based RL like Q-learning
- How softmax turns action preferences into a probability distribution
- How to sample an action from a policy with Python's random module
- The intuition behind the REINFORCE algorithm and its update rule
- Why advantages and baselines slash the variance of the gradient
- How actor-critic methods (and PPO) build on these ideas
🎯 Real-World Analogy: Learning Habits vs Rating Every Option
Imagine two tennis players. The first, a value-based player, mentally scores every possible shot before each swing — "a cross-court forehand is worth 7, a drop shot is worth 4" — then picks the highest score. It's accurate but slow, and it falls apart when there are infinitely many shot angles to rate.
The second, a policy-based player, just builds instincts: "when the ball comes here, I usually go down the line." They don't rate options — they directly adjust how often they play each shot based on whether it tends to win the point. That direct adjustment of action probabilities is exactly what a policy gradient does.
1 Policy-Based vs Value-Based RL
A policy (written π, "pi") is the agent's rule for choosing actions: given a state, it says how likely each action is. Written formally it's π(a|s) — "the probability of action a in state s."
Value-based methods like Q-learning never store a policy directly. They learn a value for each action, then act greedily on those values. Policy-based methods flip that around: they store the policy itself and adjust it to earn more reward, never bothering to rate actions on an absolute scale.
- Learns Q(s, a) — how good each action is
- Acts greedily on the values
- Struggles with continuous actions
- Usually a deterministic policy
- Learns π(a|s) — action probabilities directly
- Samples actions from the policy
- Handles continuous actions naturally
- Naturally stochastic — keeps exploring
2 A Policy Is Just Preferences Turned Into Probabilities
The simplest policy stores one number per action — a preference. To turn raw preferences into a valid probability distribution (all positive, summing to 1), you run them through the softmax function. Bigger preference means bigger probability, but every action keeps a non-zero chance, so the agent still explores.
Run this worked example and watch the preferences become probabilities:
import math
# A policy is just a rule for PICKING actions.
# We store one "preference" number per action, then turn those
# preferences into probabilities with the softmax function.
# Higher preference -> higher probability (but never exactly 0 or 1).
preferences = [2.0, 1.0, 0.1] # 3 actions: Left, Stay, Right
def softmax(prefs):
# Subtract the max for numerical stability (avoids huge exp values).
biggest = max(prefs)
exps = [math.exp(p - biggest) for p in prefs] # all > 0
total = sum(exps)
return [e / total for e in exps] # each in (0, 1), sums to 1
probs = softmax(preferences)
print("Action preferences:", preferences)
print("Action probabilities:")
names = ["Left", "Stay", "Right"]
for name, p in zip(names, probs):
print(f" P({name:<5}) = {p:.3f}") # P(Left) = 0.659 ...
print(f"Probabilities sum to {sum(probs):.3f}") # 1.000 (always)
# Notice: the BEST action (Left) gets the most probability,
# but every action keeps SOME chance -> the agent still explores.3 Sampling an Action From the Policy
Once you have probabilities, the agent doesn't always pick the top one — it samples. Python's random.choices does weighted sampling: pass the probabilities as weights and it returns higher-probability actions more often, but not every time. That built-in randomness is how policy gradient agents explore.
This example samples 1,000 actions and shows the frequencies matching the target probabilities:
import math
import random
random.seed(7) # makes the random draws repeatable for this lesson
def softmax(prefs):
biggest = max(prefs)
exps = [math.exp(p - biggest) for p in prefs]
total = sum(exps)
return [e / total for e in exps]
def sample_action(probs):
# random.choices does weighted sampling: actions with higher
# probability are picked more often, but not always.
actions = list(range(len(probs)))
return random.choices(actions, weights=probs, k=1)[0]
preferences = [2.0, 1.0, 0.1]
probs = softmax(preferences)
names = ["Left", "Stay", "Right"]
# Sample 1000 actions and count how often each one shows up.
counts = [0, 0, 0]
for _ in range(1000):
a = sample_action(probs)
counts[a] += 1
print("Target probabilities vs sampled frequency:")
for i, name in enumerate(names):
freq = counts[i] / 1000
print(f" {name:<5} target={probs[i]:.3f} sampled={freq:.3f}")
# The sampled frequencies closely match the target probabilities.
# That is what "sampling from a policy" means in reinforcement learning.4 The REINFORCE Algorithm and the Policy Gradient Objective
REINFORCE is the simplest policy gradient algorithm, and the rule is intuitive: take an action, see the reward, then make that action more likely if the reward was good and less likely if it was bad. Repeat millions of times and the policy drifts toward whatever earns reward.
The objective being maximised is the expected reward, written J(θ) = E[ R ], where θ ("theta") are the policy's parameters (your preferences). The famous policy gradient theorem says you can climb this objective with the update:
In plain English: nudge each parameter in the direction that increases the log-probability of the action you took, scaled by the reward R. Positive reward pushes the action up; negative reward pushes it down. The worked code below does exactly this for one action, and you'll do an update by hand in the next exercise.
import math
# 🎯 YOUR TURN — fill in the blanks marked with ___
def softmax(prefs):
biggest = max(prefs)
# 1) Exponentiate each (pref - biggest) so all values are positive
exps = [math.exp(p - biggest) for p in prefs]
# 2) Divide each exp by the total so they sum to 1
total = ___ # 👉 add up everything in 'exps' with sum(...)
return [e / total for e in exps]
# An agent in a maze with 3 moves and these learned preferences:
preferences = [0.5, 3.0, 0.5] # Up, Right, Down
probs = softmax(preferences)
names = ["Up", "Right", "Down"]
for name, p in zip(names, probs):
print(f"P({name}) = {p:.3f}")
# 3) Print which action is most likely (highest probability)
best_index = ___ # 👉 use probs.index(max(probs))
print("Most likely move:", names[best_index])
# ✅ Expected output:
# P(Up) = 0.067
# P(Right) = 0.866
# P(Down) = 0.067
# Most likely move: Rightimport math
import random
random.seed(1)
# 🎯 YOUR TURN — fill in the blanks marked with ___
# Do ONE policy-gradient update: nudge the chosen action's preference
# UP when the reward is positive, and DOWN when it is negative.
def softmax(prefs):
biggest = max(prefs)
exps = [math.exp(p - biggest) for p in prefs]
total = sum(exps)
return [e / total for e in exps]
prefs = [0.0, 0.0] # 2 actions: Left, Right (start equal)
learning_rate = 0.5
probs = softmax(prefs)
action = random.choices([0, 1], weights=probs, k=1)[0]
reward = 1 if action == 1 else -1 # pretend "Right" is the good action
print(f"Chose action {action}, reward {reward}")
print(f"Before: P(Left)={probs[0]:.3f}, P(Right)={probs[1]:.3f}")
# REINFORCE-style update for the action we actually took:
# new_pref = old_pref + learning_rate * reward * (1 - prob_of_that_action)
# 1) Compute the nudge for the chosen action
nudge = learning_rate * reward * (1 - probs[action]) # 👉 keep this line
# 2) Apply the nudge to prefs[action]
prefs[action] = prefs[action] + ___ # 👉 add 'nudge' to the chosen preference
probs = softmax(prefs)
print(f"After: P(Left)={probs[0]:.3f}, P(Right)={probs[1]:.3f}")
# ✅ Expected output:
# Chose action 1, reward 1
# Before: P(Left)=0.500, P(Right)=0.500
# After: P(Left)=0.438, P(Right)=0.562📉 Advantages, Baselines, and Actor-Critic
Plain REINFORCE works, but it's noisy. Because it uses the raw reward of a whole episode, the learning signal swings wildly — two nearly identical episodes can give very different totals. That high variance makes training slow and unstable.
The fix is the advantage: instead of using the raw reward, subtract a baseline (a reference reward, often the average) and use the difference:
Now an action is judged as better or worse than expected, not on its raw, noisy total. This slashes variance without biasing the gradient — actions that beat the baseline go up, those below it go down.
Actor-critic methods take this one step further. They run two pieces side by side: the actor is the policy that chooses actions, and the critic is a learned value function that estimates how good each state is. The critic's estimate becomes the baseline for the actor's update. This is the foundation of A2C and PPO — the algorithm used to fine-tune ChatGPT with RLHF — which is why those methods learn far more stably than plain REINFORCE.
5 Common Errors (And How to Fix Them)
These are the four mistakes that derail almost every first policy gradient implementation:
❌ Training is unstable and never converges (high variance)
Plain REINFORCE uses the full episode return, which swings wildly from run to run. The gradient estimate is so noisy the policy can't settle.
# Subtract a baseline to get an advantage, and average over a batch
advantage = reward - baseline # not the raw reward
prefs[action] += lr * advantage * (1 - probs[action])❌ Forgetting the baseline entirely
If every reward in your task is positive (say 0 to 100), every action gets pushed up — the policy can't tell good actions from merely-less-good ones.
baseline = running_average_of_rewards # so advantages are both + and -
advantage = reward - baseline # below-average actions now go DOWN❌ Learning rate too high — the policy collapses
A large learning rate lets one lucky episode shove all the probability onto a single action. The policy stops exploring and gets stuck forever.
learning_rate = 0.01 # start small; PPO commonly uses 3e-4
# If P(one action) rockets to ~1.0 in a few steps, lower it.❌ Reward scaling is off
Rewards of +1000 produce giant gradients; rewards of 0.0001 produce no learning. Either way the update size is wrong.
# Normalise advantages to roughly mean 0, std 1 before the update
mean = sum(advs) / len(advs)
std = (sum((a - mean) ** 2 for a in advs) / len(advs)) ** 0.5
advs = [(a - mean) / (std + 1e-8) for a in advs]🎯 Mini-Challenge: Train a Policy With a Baseline
Time to put it all together with the support faded. The starter below is just a comment outline — write the training loop yourself. Use softmax for the policy, random.choices to sample, and an advantage (reward minus a running baseline) for the update.
import math
import random
random.seed(42)
# 🎯 MINI-CHALLENGE: Train a tiny policy with REINFORCE + a baseline
#
# One state, two actions: 0 = "Left", 1 = "Right".
# "Right" pays reward +1, "Left" pays reward 0.
# Over many episodes the policy should learn to almost always pick Right.
#
# Steps:
# 1. Start with prefs = [0.0, 0.0] and learning_rate = 0.1
# 2. Keep a running 'baseline' = average reward seen so far (start at 0.0)
# 3. For 300 episodes:
# - softmax(prefs) -> probs
# - sample an action with random.choices([0,1], weights=probs)
# - reward = 1 if action == 1 else 0
# - advantage = reward - baseline # subtracting the baseline cuts variance
# - prefs[action] += learning_rate * advantage * (1 - probs[action])
# - update baseline toward reward: baseline = baseline + 0.05 * (reward - baseline)
# 4. Print the final P(Left) and P(Right)
#
# ✅ Expected (roughly): P(Right) should end up well above 0.9
# your code here📋 Quick Reference
| Term | What It Means | In Code / Maths |
|---|---|---|
| Policy π(a|s) | Probability of each action in a state | softmax(prefs) |
| Softmax | Turns preferences into probabilities | exp(x) / sum(exp(x)) |
| Sampling | Pick an action by its probability | random.choices(a, weights=p) |
| REINFORCE | Simplest policy gradient update | ∇J ≈ R · ∇log π(a|s) |
| Advantage | Reward relative to a baseline | reward − baseline |
| Actor-Critic | Policy (actor) + value baseline (critic) | A2C, PPO |
❓ Frequently Asked Questions
Lesson complete — you can think in policies now!
You learned how policy-based RL differs from value-based RL, how softmax turns preferences into probabilities, how to sample an action, the REINFORCE update rule and its objective, and why advantages, baselines, and actor-critic methods make training stable. These ideas scale all the way up to PPO and the RLHF that aligns modern language models.
Practice quiz
What does a policy-based RL method learn directly?
- A value for each state only
- The reward function
- The policy itself — the probability of each action in each state
- The transition probabilities
Answer: The policy itself — the probability of each action in each state. Policy-based methods adjust action probabilities directly, rather than learning action values first.
How does value-based RL (like Q-learning) differ from policy-based RL?
- Value-based learns action values then acts greedily; policy-based learns the policy directly
- Value-based has no rewards
- Policy-based cannot explore
- They are the same approach
Answer: Value-based learns action values then acts greedily; policy-based learns the policy directly. Q-learning learns Q(s,a) and acts greedily; policy gradients tune action probabilities directly.
What does the softmax function do to action preferences?
- Picks only the largest preference
- Sets every probability to zero
- Sorts them in order
- Turns them into a probability distribution that is all positive and sums to 1
Answer: Turns them into a probability distribution that is all positive and sums to 1. Softmax exponentiates and normalizes preferences so they form a valid probability distribution.
Why does a softmax policy keep exploring?
- It always picks a random action
- Every action keeps a non-zero probability, so the agent still tries others
- It ignores preferences
- It never updates
Answer: Every action keeps a non-zero probability, so the agent still tries others. Even the best action does not get probability 1, so lower-preference actions still get chosen sometimes.
What is the core idea of the REINFORCE algorithm?
- Make actions that earned high reward more likely and low-reward actions less likely
- Always pick the action with the highest Q-value
- Never change the policy
- Memorize the entire reward table
Answer: Make actions that earned high reward more likely and low-reward actions less likely. REINFORCE nudges the policy so high-reward actions become more probable and low-reward ones less probable.
The policy gradient update is proportional to which quantity?
- The number of states
- The learning rate alone
- Reward times the gradient of the log-probability of the chosen action
- The discount factor squared
Answer: Reward times the gradient of the log-probability of the chosen action. The policy gradient theorem gives an update of R times the gradient of log pi(a|s).
Why does plain REINFORCE have high variance?
- It uses too small a learning rate
- It uses the raw return of a whole episode, which swings wildly between runs
- It never receives rewards
- It only uses one state
Answer: It uses the raw return of a whole episode, which swings wildly between runs. Episode returns are noisy, so the gradient estimate is noisy and learning is slow and unstable.
What is the advantage in policy gradient methods?
- The raw reward times two
- The number of episodes run
- The maximum Q-value
- Reward minus a baseline, judging an action as better or worse than expected
Answer: Reward minus a baseline, judging an action as better or worse than expected. advantage = reward - baseline; subtracting a baseline cuts variance without biasing the gradient.
Why does subtracting a baseline help training?
- It makes all rewards positive
- It dramatically lowers variance without biasing the gradient
- It removes the need for a policy
- It speeds up the environment
Answer: It dramatically lowers variance without biasing the gradient. A baseline reframes actions as better-or-worse-than-expected, reducing noise in the gradient estimate.
In actor-critic methods, what role does the critic play?
- It chooses the actions
- It deletes the policy
- It estimates state values, serving as the baseline for the actor's policy update
- It stores the replay buffer
Answer: It estimates state values, serving as the baseline for the actor's policy update. The actor is the policy; the critic's value estimate is the baseline that cuts variance (used in A2C and PPO).
Continue this course
- Previous: Q-Learning & Deep Q-Networks (DQN)
- Next: Computer Vision Pipelines with OpenCV & PyTorch/TensorFlow — Build end-to-end vision pipelines for classification, detection, and segmentation
- Quick reference: AI & Machine Learning cheat sheet
Frequently asked questions
What is the difference between policy-based and value-based reinforcement learning?
Value-based methods like Q-learning learn how good each action is (a value), then act greedily on those values. Policy-based methods skip the value step and directly learn the policy — the probability of each action in each state — adjusting those probabilities to earn more reward. Policy-based methods naturally output stochastic policies and handle continuous action spaces, which value-based methods struggle with.
What is the REINFORCE algorithm in simple terms?
REINFORCE runs a full episode, then nudges the policy so actions that led to high reward become more likely and actions that led to low reward become less likely. The size of the nudge is the reward multiplied by the gradient of the log-probability of the chosen action. It is the simplest policy gradient method and the foundation everything else builds on.
Why do policy gradient methods have such high variance?
REINFORCE uses the raw return of a whole episode as its learning signal, and returns swing wildly from one episode to the next because of randomness in both the policy and the environment. Two near-identical episodes can produce very different totals, so the gradient estimate is noisy and learning is slow and unstable.
What is a baseline and why does subtracting it help?
A baseline is a reference reward you subtract from the actual reward to get the advantage (reward minus baseline). A common baseline is the average reward, or a learned value estimate. Subtracting it dramatically lowers variance without biasing the gradient, because actions are now judged as better-or-worse-than-expected rather than on their raw, noisy totals.
What is actor-critic and how does it relate to policy gradients?
Actor-critic combines both worlds: the actor is a policy that chooses actions, and the critic is a value function that estimates how good states are. The critic's estimate serves as the baseline for the actor's policy gradient update, cutting variance. PPO and A2C are actor-critic methods, which is why they learn far more stably than plain REINFORCE.