Diffusion Models Explained

Reviewed & published by Brayan K

By the end of this lesson you'll be able to explain — and run code for — how DALL-E and Stable Diffusion turn random noise into images by learning to denoise, step by step.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

🌍 Real-World Analogy: Sculpting from Static

Picture an old TV tuned to a dead channel — a screen full of fuzzy static. Now imagine a sculptor who can look at any patch of static and carefully chip away the randomness until a clear picture appears underneath. That is exactly what a diffusion model does: it sculpts an image out of noise.

During training the model watches the opposite happen thousands of times — real photos slowly dissolving into static — so it learns precisely what "noise" looks like at every stage. To generate something new you hand it a fresh screen of static and your text description, and it removes the noise a little at a time, guided by your words, until a finished image emerges.

Diffusion models power DALL-E, Stable Diffusion, Midjourney, and Imagen. They produce higher-quality images than older GANs with far more stable training, and the same denoising recipe also generates audio, video, and 3D shapes.

1 The Forward Process — adding noise on purpose

The forward process takes clean data and adds a small amount of Gaussian noise (bell-curve random values) at each of many timesteps. Repeat it enough and the data becomes indistinguishable from pure noise. Crucially this process is fixed, not learned — it follows a preset noise schedule (called beta) that decides how much noise to add at each step.

Why deliberately destroy data? Because it creates perfect training pairs: at every step you know exactly how much noise you added, so you can teach a network to undo it. Run the worked example below to watch a tiny 5-value signal dissolve into noise.

# FORWARD PROCESS — turning a clean signal into noise, one step at a time.
# A diffusion model first DESTROYS data so it can later learn to REBUILD it.
# We use a tiny 1-D "signal" (think of it as 5 pixels) instead of an image.

import random
random.seed(0)                       # reproducible noise for the demo

signal = [0.9, 0.7, 0.5, 0.3, 0.1]   # the clean data we start from
steps = 4                            # how many times we add noise
noise_amount = 0.15                  # std-dev of noise added per step

print("step 0 (clean):", [round(v, 2) for v in signal])

x = list(signal)                     # copy so we keep the original
for step in range(1, steps + 1):
    # Add a little Gaussian (bell-curve) noise to EVERY value.
    x = [v + random.gauss(0, noise_amount) for v in x]
    print(f"step {step}        :", [round(v, 2) for v in x])

print()
print("Each step nudges the numbers by random amounts.")
print("After enough steps the signal is unrecognisable — almost pure noise.")

# Expected output (numbers vary slightly, shape is the same):
# step 0 (clean): [0.9, 0.7, 0.5, 0.3, 0.1]
# step 4         : [a noisy, scrambled version of the line above]

2 The Reverse Process — denoising one step

The reverse process is the part the model actually learns. Given a noisy sample, the network predicts the noise that was added, and you subtract that prediction to get a cleaner sample. The network is usually a U-Net — an encoder-decoder shape with skip connections that's great at processing images — but for the idea you only need one move: predict the noise, then subtract it.

The example below fakes a decent noise prediction so you can see a single denoise step pull a noisy signal back toward the clean one.

# REVERSE PROCESS — one denoising step.
# The model's whole job is to PREDICT the noise that was added.
# Once we have that prediction, we SUBTRACT it to recover cleaner data.

import random
random.seed(1)

clean = [0.9, 0.7, 0.5, 0.3, 0.1]                 # the data we want back
true_noise = [random.gauss(0, 0.2) for _ in clean] # the noise reality added
noisy = [c + n for c, n in zip(clean, true_noise)] # what the model receives

# A REAL model is a neural net (a U-Net). Here we fake a decent guess so the
# idea is clear: predicted_noise ≈ the noise that was added.
predicted_noise = [n * 0.9 for n in true_noise]    # 90% accurate guess

# ONE denoise step = subtract the predicted noise from the noisy input.
denoised = [x - p for x, p in zip(noisy, predicted_noise)]

print("clean target   :", [round(v, 2) for v in clean])
print("noisy input    :", [round(v, 2) for v in noisy])
print("predicted noise:", [round(v, 2) for v in predicted_noise])
print("after 1 denoise:", [round(v, 2) for v in denoised])
print()
print("Notice 'after 1 denoise' is CLOSER to the clean target than the input.")
print("Sampling just repeats this subtract-the-noise step many times.")

# Expected output (values vary, the pattern does not):
# 'after 1 denoise' sits between the noisy input and the clean target.

3 Sampling — generating from pure noise

Sampling is how you create something new. You start from a screen of pure random noise and apply the denoise step over and over, walking backwards through the timesteps until a clean sample appears. More steps generally means higher quality but slower generation — early models used 1,000 steps, while DDIM and modern samplers get great results in 20-50.

The toy loop below starts from random noise and converges toward a target shape. A real model never sees the target; it learned that shape from millions of training images.

# SAMPLING — generate data from scratch by denoising over many steps.
# Start from PURE NOISE, then repeatedly predict-and-subtract toward clean data.

import random
random.seed(7)

target = [0.9, 0.7, 0.5, 0.3, 0.1]      # the "shape" we are trying to recover
x = [random.gauss(0, 1) for _ in target] # step 0 = pure random noise

print("start (pure noise):", [round(v, 2) for v in x])

total_steps = 8
for step in range(total_steps, 0, -1):
    # Toy "model": guess how far each value is from the target shape.
    # A real model never sees the target — it learned the shape during training.
    predicted_noise = [xi - ti for xi, ti in zip(x, target)]
    rate = 1 / step                      # take a partial step each time
    x = [xi - rate * p for xi, p in zip(x, predicted_noise)]

print("end (denoised)    :", [round(v, 2) for v in x])
print()
print("More steps = smoother, higher-quality result (but slower).")

# Expected output:
# 'end (denoised)' is very close to [0.9, 0.7, 0.5, 0.3, 0.1].

4 Training — learning to predict the noise

Training is surprisingly simple. Take a clean sample, pick a random timestep, add the matching amount of Gaussian noise, and ask the network to predict that noise. The loss is the mean squared error between the predicted noise and the true noise. Because you generated the noise yourself, you always have a perfect answer to grade against — no human labels needed. This is self-supervised learning.

Run the example to compare an untrained model (predicts nothing) against a trained one (predicts the noise accurately) and see the loss drop.

# TRAINING — the loss a diffusion model actually minimises.
# Goal: given a noisy sample, the model should PREDICT the added noise.
# Loss = mean squared error between predicted noise and true noise.

import random
random.seed(3)

clean = [0.5, 0.2, 0.8, 0.4]
true_noise = [random.gauss(0, 1) for _ in clean]
noisy = [c + n for c, n in zip(clean, true_noise)]

def mse(predicted, actual):
    diffs = [(p - a) ** 2 for p, a in zip(predicted, actual)]
    return sum(diffs) / len(diffs)

bad_guess  = [0.0, 0.0, 0.0, 0.0]            # an untrained model: predicts nothing
good_guess = [n * 0.95 for n in true_noise]  # a trained model: close to the truth

print("loss (untrained):", round(mse(bad_guess, true_noise), 3))
print("loss (trained)  :", round(mse(good_guess, true_noise), 3))
print()
print("Training drives this loss DOWN: better noise prediction = better samples.")

# Expected output:
# loss (untrained): a fairly large number
# loss (trained)  : close to 0  (good noise prediction)
# 🎯 YOUR TURN — run the FORWARD process: add Gaussian noise to a signal.
# Fill in every ___ . Run it and compare against the expected output.

import random
random.seed(0)

signal = [1.0, 0.6, 0.2]      # a tiny clean signal
steps = 3

x = list(signal)
for step in range(steps):
    # 1) Add Gaussian noise (mean 0, std 0.1) to EACH value.
    #    👉 use random.gauss(0, 0.1)
    x = [v + ___ for v in x]            # 👉 replace ___

    # 2) 👉 print the step number and the rounded values
    print("step", step + 1, [round(v, 2) for v in ___])   # 👉 replace ___

# ✅ Expected output (numbers vary, 3 lines, values drift away from the signal):
# step 1 [...]
# step 2 [...]
# step 3 [...]
# 🎯 YOUR TURN — run ONE reverse step: subtract the predicted noise.
# Fill in every ___ . The model's prediction is given to you.

noisy           = [1.1, 0.4, 0.9]   # what the model receives
predicted_noise = [0.2, 0.1, 0.3]   # what the model thinks the noise is

# 1) Denoise = noisy value MINUS predicted noise, value by value.
#    👉 subtract p from x for each pair
denoised = [x - ___ for x, p in zip(noisy, predicted_noise)]   # 👉 replace ___

# 2) 👉 print the denoised list
print("denoised:", ___)             # 👉 replace ___

# ✅ Expected output:
# denoised: [0.9, 0.3, 0.6]

5 Conditioning & Classifier-Free Guidance

So far the model denoises toward any realistic image. To get the image you want, you condition the denoising on a text prompt. The prompt is encoded (often by a model like CLIP) and fed into the U-Net through cross-attention, so every denoise step is nudged toward "a cosy cabin in a snowy forest" rather than just anything.

Classifier-free guidance (CFG) makes the prompt count more. At each step the model predicts noise twice — once with the prompt and once with no prompt — then exaggerates the difference between them. The guidance scale controls how much: around 7-8 is the usual sweet spot.

How the guidance scale behaves:

6 Latent Diffusion — how Stable Diffusion stays fast

Denoising every pixel of a 512×512 image, hundreds of times, is slow. Latent diffusion fixes this by doing the whole process in a small latent space. An autoencoder first compresses the image down to a tiny representation (e.g. 64×64), diffusion runs there, and then a decoder expands the cleaned latent back to a full image.

This is the trick behind Stable Diffusion: it denoises something roughly 48× smaller, so it runs on a consumer GPU instead of a data centre. The text conditioning and guidance you just learned happen inside that latent space.

The worked example below shows the real Hugging Face diffusers API. It needs diffusers, torch, and a GPU, so treat it as a preview of the real code — the expected output is shown in a comment.

# REAL diffusion in practice: Stable Diffusion via Hugging Face 'diffusers'.
# (Needs the diffusers + torch libraries and a GPU; shown to illustrate the API.)

from diffusers import StableDiffusionPipeline
import torch

# Load a pretrained LATENT diffusion model. "Latent" means it denoises in a
# small compressed space, not full-resolution pixels — that is why it's fast.
pipe = StableDiffusionPipeline.from_pretrained(
    "runwayml/stable-diffusion-v1-5",
    torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")

# The text prompt CONDITIONS the denoising via cross-attention.
# guidance_scale is classifier-free guidance: higher = follow the prompt harder.
image = pipe(
    "a cosy cabin in a snowy forest, digital art",
    num_inference_steps=30,   # sampling steps (more = higher quality, slower)
    guidance_scale=7.5,       # 7-8 is the usual sweet spot
).images[0]

image.save("cabin.png")
print("Saved a 512x512 image to cabin.png")

# Expected output:
# Saved a 512x512 image to cabin.png
# (diffusers also drives audio gen — see AudioLDM / MusicGen pipelines.)

🎨 Beyond Images: What Diffusion Can Generate

The predict-the-noise recipe works on anything you can add noise to. Swap the data and the network shape, keep the forward/reverse idea, and you get a generator for a new domain:

! Common Errors (And How to Fix Them)

❌ Output is blurry or noisy — too few sampling steps

You set num_inference_steps very low, so denoising didn't finish.

# Give the sampler enough steps to fully denoise.
pipe(prompt, num_inference_steps=30)   # 20-50 is a good range
# DDIM/DPM samplers look great at ~25-30; raw DDPM may need more.

❌ Image ignores the prompt OR looks "fried" — wrong guidance scale

Guidance too low ignores your words; too high oversaturates and distorts.

# Stay near the sweet spot.
pipe(prompt, guidance_scale=7.5)   # 7-8 = balanced
# guidance_scale=1   -> prompt mostly ignored
# guidance_scale=20  -> oversaturated, distorted output

❌ Garbage samples — train/sample noise-schedule mismatch

You sampled with a different number of steps or schedule than you trained on.

# Sampling MUST match the noise schedule used in training.
# Use the same scheduler the model was trained/released with:
from diffusers import DDIMScheduler
pipe.scheduler = DDIMScheduler.from_config(pipe.scheduler.config)

❌ "CUDA out of memory"

Pixel-space or large batches won't fit; latent diffusion plus half precision helps.

# Use fp16 and generate one image at a time.
pipe = pipe.to("cuda")
pipe.enable_attention_slicing()   # trades a little speed for less memory

📋 Quick Reference

TermWhat It MeansIn One Line
Forward processAdds noise to data over many stepsFixed, not learned
Reverse processRemoves noise step by stepWhat the network learns
Noise schedule (beta)How much noise per timestepLinear or cosine
Training targetPredict the added noiseMSE loss, self-supervised
U-NetEncoder-decoder networkPredicts the noise
Sampling stepsDenoise iterations to generate20-50 (DDIM) vs 1000
ConditioningGuide output with a promptCross-attention on text
Classifier-free guidanceStrengthen the promptScale ≈ 7-8
Latent diffusionDiffuse in compressed spaceWhy SD is fast

🎯 Mini Challenge: Noise It, Then Denoise It

Time to fade the scaffolding. Add noise to a clean signal over a few steps, then take one denoise step using a noise prediction you compute yourself, and show that you land back near the original. The starter gives you only a comment outline — write the logic yourself.

# 🎯 MINI-CHALLENGE: noise then denoise a signal
# Brief: add noise to a clean signal over a few steps, then take ONE denoise
# step using a "predicted noise" you compute, and print how close you got.
#
# 1. import random; random.seed(0)
# 2. clean = [0.8, 0.5, 0.2]  and an empty list noise_added = []
# 3. Loop 3 times: pick n = random.gauss(0, 0.1), append n to noise_added,
#    and add n to each value of a working copy 'noisy'
# 4. predicted = the SUM of noise added to each position (a perfect guess)
#    -> denoised = noisy value minus predicted, position by position
# 5. print clean, noisy (rounded), and denoised (rounded)
#
# ✅ Expected (example): 'denoised' should land back very close to 'clean'
#    clean    : [0.8, 0.5, 0.2]
#    denoised : [0.8, 0.5, 0.2]  (give or take rounding)

# your code here

Lesson complete — you can now explain how diffusion models work!

You learned that the forward process adds Gaussian noise to data, that the network is trained to predict that noise with an MSE loss, and that sampling starts from pure noise and denoises step by step. You also saw how text conditioning plus classifier-free guidance steer the output, and how latent diffusion makes Stable Diffusion fast enough for a consumer GPU.

Practice quiz

In a diffusion model, what does the forward process do?

  • Gradually adds Gaussian noise to data over many steps
  • Removes noise from a noisy image
  • Classifies the image into a category
  • Compresses the image into a latent code

Answer: Gradually adds Gaussian noise to data over many steps. The forward (noising) process is fixed and adds a little Gaussian noise at each timestep until the data becomes pure noise.

Is the forward (noising) process learned or fixed?

  • Learned by gradient descent
  • Fixed — it follows a preset noise schedule
  • Learned by the discriminator
  • Chosen randomly at inference time

Answer: Fixed — it follows a preset noise schedule. The forward process is fixed: it follows a preset noise schedule (beta). Only the reverse process is learned.

What does the neural network actually predict during diffusion training?

  • The class label of the image
  • The next pixel value
  • The noise that was added at a given timestep
  • The number of denoising steps needed

Answer: The noise that was added at a given timestep. The network is trained to predict the noise added at a random timestep; the loss is the MSE between predicted and true noise.

What is the reverse (denoising) process used for?

  • Destroying data into noise
  • Generating data by predicting and subtracting noise step by step
  • Measuring how blurry an image is
  • Encoding text into embeddings

Answer: Generating data by predicting and subtracting noise step by step. The reverse process is the learned part: starting from noise, predict the noise and subtract it repeatedly to rebuild clean data.

How does sampling generate a brand-new image?

  • It starts from a clean image and adds noise
  • It copies the nearest training image
  • It starts from pure noise and denoises repeatedly
  • It averages all training images

Answer: It starts from pure noise and denoises repeatedly. Sampling starts from pure random noise and applies the denoise step many times, walking backward through the timesteps.

What network architecture is typically used to predict the noise?

  • A U-Net
  • A decision tree
  • A support vector machine
  • A k-nearest-neighbours model

Answer: A U-Net. Diffusion models usually use a U-Net — an encoder-decoder with skip connections that is well suited to images.

What loss function does diffusion training minimise?

  • Cross-entropy on class labels
  • Mean squared error between predicted and true noise
  • Hinge loss
  • Adversarial loss from a discriminator

Answer: Mean squared error between predicted and true noise. The objective is the MSE between the predicted noise and the true noise added during the forward process.

What does classifier-free guidance control with its guidance scale?

  • The learning rate of training
  • How strongly generation follows the text prompt
  • The number of layers in the U-Net
  • The size of the latent space

Answer: How strongly generation follows the text prompt. Classifier-free guidance strengthens the prompt's influence; a scale around 7-8 is the usual sweet spot.

Why is latent diffusion (e.g. Stable Diffusion) faster than pixel-space diffusion?

  • It uses fewer training images
  • It denoises in a small compressed latent space instead of full-resolution pixels
  • It skips the reverse process entirely
  • It only generates black-and-white images

Answer: It denoises in a small compressed latent space instead of full-resolution pixels. Latent diffusion runs the noising/denoising inside a compressed latent space, so there is far less data to process.

Roughly how many sampling steps do modern samplers like DDIM need for good quality?

  • About 20-50 steps
  • Exactly 1 step
  • Always 10,000 steps
  • About 1 million steps

Answer: About 20-50 steps. Early models used ~1,000 steps, but DDIM and modern samplers get great results in about 20-50 steps.

Continue this course

Frequently asked questions

What is a diffusion model in simple terms?

A diffusion model is a generative AI that learns to create data by reversing a noising process. During training it watches clean data slowly turn into random noise, and learns to predict the noise at each step. To generate something new it starts from pure noise and removes the predicted noise step by step until a clean image, audio clip, or other sample emerges.

What is the difference between the forward and reverse process?

The forward (noising) process is fixed, not learned — it just adds a little Gaussian noise to the data over many steps until nothing is left but noise. The reverse (denoising) process is what the neural network learns: starting from noise, it predicts and subtracts the noise step by step to rebuild realistic data. Forward destroys; reverse creates.

What does the model actually predict during training?

It predicts the noise. You take a clean sample, add a known amount of Gaussian noise at a random timestep, and ask the network to output that noise. The training loss is the mean squared error between the predicted noise and the true noise. Because the noise is known, you have a perfect target — no labels required.

What is classifier-free guidance and the guidance scale?

Classifier-free guidance steers generation toward your text prompt without a separate classifier. The model runs twice — once with the prompt and once unconditionally — and pushes the result in the direction the prompt adds. The guidance scale controls how hard it pushes: around 7-8 is the sweet spot, too low ignores the prompt, and too high causes oversaturated, distorted images.

What is latent diffusion and why is Stable Diffusion fast?

Latent diffusion runs the whole noising and denoising process inside a small compressed 'latent' space instead of on full-resolution pixels. An autoencoder shrinks the image first, diffusion happens on that tiny representation, then a decoder expands it back to a full image. This is why Stable Diffusion runs on a consumer GPU while pixel-space diffusion is far slower.

What can diffusion models generate besides images?

The same predict-the-noise recipe works on any data you can add noise to. Beyond images (DALL-E, Stable Diffusion, Midjourney) it powers audio and music generation (AudioLDM, MusicGen), video, 3D shapes, and even molecule design. You change the data and the network architecture, but the forward/reverse diffusion idea stays the same.

Related lessons