How Large Language Models Work

Reviewed & published by Brayan K

By the end of this lesson you'll be able to explain, in plain English and with runnable code, how an LLM turns your prompt into text — token by token.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

🌍 Real-World Analogy: Autocomplete on Steroids

You already use a tiny language model every day: phone keyboard autocomplete. Type "I'll be there in five" and it suggests "minutes". An LLM is that exact idea — autocomplete on steroids.

Imagine someone who has read almost everything ever written and remembers the patterns. When you type "The cat sat on the", they don't understand cats — they just know that "mat" is the most likely next word. Scale that predictor up to billions of internal settings and trillions of words of practice, and you get something that can write essays, code, and answers. That's all an LLM is doing: predicting the next token, over and over, very well.

1 Next-Token Prediction — the one job an LLM has

An LLM has exactly one job: given some text, output a probability for every possible next token. A token is a chunk of text (often a word piece). The model assigns, say, 62% to "mat" and 21% to "floor", picks one, appends it, and runs again. Generating a paragraph is just this loop repeated hundreds of times — that's called autoregressive generation.

The simplest way to pick is greedy decoding: always take the highest-probability token. Run the worked example below — it takes a probability dictionary and grabs the most likely next word.

# How an LLM picks the next word: NEXT-TOKEN PREDICTION
# An LLM reads your text and outputs a probability for every possible
# next token. "Greedy" decoding just grabs the most likely one.

# Pretend the model just read: "The cat sat on the"
# These are the probabilities it assigned to possible next tokens.
next_token_probs = {
    "mat":   0.62,   # most likely
    "floor": 0.21,
    "roof":  0.10,
    "couch": 0.05,
    "moon":  0.02,
}

# Greedy pick = the token with the highest probability.
best_token = max(next_token_probs, key=next_token_probs.get)

print("Context: 'The cat sat on the'")
print()
for token, prob in next_token_probs.items():
    bar = "#" * int(prob * 40)          # tiny text bar chart
    print(f"  {token:<6} {prob:>4.2f} {bar}")

print()
print("Greedy next token ->", best_token)
print("Generated so far  -> 'The cat sat on the " + best_token + "'")

# Expected output:
# Greedy next token -> mat
# An LLM repeats this loop: pick a token, append it, predict the next.

2 Decoder-Only Transformers & Causal Attention

GPT, LLaMA, and Mistral all use a decoder-only transformer. The key mechanism is self-attention: at every position, the model looks back at earlier tokens and decides which ones matter for predicting what comes next. "Causal" (or masked) attention means a token can only see tokens before it — never the future. That's what makes left-to-right generation possible.

Reading "The cat sat on" with causal attention:

Stack dozens of these attention-plus-feed-forward layers, and the model can capture grammar, facts, and style. You don't need the matrix maths to use one — but knowing that each token only attends to the past explains why an LLM writes one token at a time.

3 Tokens & the Context Window

Models don't read characters or whole words — they read tokens. A rough rule for English is one token per four characters (about ¾ of a word). The context window is the maximum number of tokens — your prompt plus the reply — the model can hold at once. Go over it and the oldest text falls out of view, so the model effectively "forgets" the start of a long conversation.

# TOKENS & CONTEXT WINDOW
# LLMs don't read characters or whole words — they read TOKENS (word pieces).
# A rough rule of thumb in English: ~1 token per 4 characters (~0.75 words).

text = "Large language models predict the next token."

# Real tokenizers (like tiktoken) split smarter, but we'll estimate.
approx_tokens = max(1, len(text) // 4)
words = len(text.split())

print("Text:", text)
print("Characters:", len(text))
print("Words:", words)
print("Estimated tokens:", approx_tokens)
print()

# The CONTEXT WINDOW is the max tokens a model can consider at once
# (prompt + answer combined). Anything older gets dropped.
context_window = 8192
prompt_tokens = 6000
reply_budget = context_window - prompt_tokens

print(f"Context window : {context_window} tokens")
print(f"Prompt uses    : {prompt_tokens} tokens")
print(f"Room for reply : {reply_budget} tokens")

# Expected output:
# Estimated tokens: ~11
# Room for reply : 2192 tokens
# If the prompt fills the window, the model 'forgets' the earliest text.

4 Parameters & Weights — where the "knowledge" lives

A model's parameters (also called weights) are just numbers — billions of them — that the model adjusts during training. They're the dials that turn an input of tokens into an output of probabilities. "GPT-3 has 175 billion parameters" means 175 billion of these numbers.

There's no database of facts inside; the "knowledge" is encoded across all those weights. More parameters generally means more capacity to capture patterns — but, as you'll see in the scaling section, size alone isn't everything.

5 Temperature & Sampling — controlling randomness

Greedy decoding always picks the top token, which can feel robotic and repetitive. Instead, models usually sample from the probabilities. The raw scores the model produces are called logits; softmax turns them into probabilities that sum to 1.

Temperature is the knob. You divide the logits by the temperature before softmax. A low temperature (e.g. 0.5) sharpens the distribution — the top token dominates, output is safe and consistent. A high temperature (e.g. 2.0) flattens it — output gets diverse and creative, but also more likely to go off the rails. Run the example to watch the probabilities shift.

# TEMPERATURE controls randomness. The model outputs raw scores called
# "logits". Temperature divides the logits, then softmax turns them into
# probabilities. Low temp = confident/repetitive. High temp = creative/risky.

import math

# Raw scores (logits) for 4 candidate next tokens.
tokens = ["mat", "floor", "roof", "moon"]
logits = [3.0, 1.5, 0.8, -1.0]

def softmax(scores):
    m = max(scores)                       # subtract max for numerical safety
    exps = [math.exp(s - m) for s in scores]
    total = sum(exps)
    return [e / total for e in exps]

def apply_temperature(logits, temperature):
    # Divide every logit by the temperature BEFORE softmax.
    scaled = [l / temperature for l in logits]
    return softmax(scaled)

print(f"{'token':<8}{'T=0.5':>8}{'T=1.0':>8}{'T=2.0':>8}")
for t in (0.5, 1.0, 2.0):
    probs = apply_temperature(logits, t)
    if t == 0.5:
        cold = probs
    if t == 2.0:
        hot = probs

for i, tok in enumerate(tokens):
    p_cold = apply_temperature(logits, 0.5)[i]
    p_mid  = apply_temperature(logits, 1.0)[i]
    p_hot  = apply_temperature(logits, 2.0)[i]
    print(f"{tok:<8}{p_cold:>8.3f}{p_mid:>8.3f}{p_hot:>8.3f}")

print()
print("Low temperature  -> 'mat' dominates (safe, predictable).")
print("High temperature -> probabilities flatten (creative, more random).")

# Expected output (probabilities for 'mat'):
# T=0.5 ~ 0.94   T=1.0 ~ 0.74   T=2.0 ~ 0.52
# Lower temperature concentrates probability on the top token.
# 🎯 YOUR TURN — make the model pick its next token (greedy decoding)
# Fill in every ___ . Run it and compare against the expected output.

next_token_probs = {
    "pizza":  0.55,
    "salad":  0.30,
    "rocks":  0.15,
}

# 1) Greedy decoding = the token with the HIGHEST probability.
#    👉 use max(...) with key=next_token_probs.get
best_token = ___          # 👉 replace ___

# 2) 👉 print the winning token
print("Next token:", ___)

# 3) 👉 build the sentence "I want to eat <best_token>"
print(f"I want to eat {___}")

# ✅ Expected output:
# Next token: pizza
# I want to eat pizza
# 🎯 YOUR TURN — turn raw logits into probabilities, with temperature
# Fill in every ___ . No numpy — just the math module.

import math

logits = [2.0, 1.0, 0.1]      # raw scores for 3 candidate tokens

def softmax(scores):
    m = max(scores)
    exps = [math.exp(s - m) for s in scores]
    total = sum(exps)
    # 👉 each probability = its exp divided by the total
    return [e / ___ for e in exps]      # 👉 replace ___

def apply_temperature(logits, temperature):
    # 👉 divide every logit by temperature, THEN softmax
    scaled = [l / ___ for l in logits]  # 👉 replace ___
    return softmax(scaled)

probs = apply_temperature(logits, 1.0)
print("Probabilities:", [round(p, 3) for p in probs])
print("They sum to:", round(sum(probs), 3))

# ✅ Expected output:
# Probabilities: [0.659, 0.242, 0.099]
# They sum to: 1.0

6 Pretraining vs Fine-Tuning

Pretraining is the expensive part: the model learns general language by predicting the next token across trillions of tokens of internet text. This takes months on huge GPU clusters and can cost millions of dollars. The result is a model that knows language but isn't specialised.

Fine-tuning takes that pretrained model and cheaply adapts it to a specific task or style using a much smaller dataset — sometimes just thousands of examples. Techniques like LoRA make this affordable on a single GPU. The rule of thumb: almost everyone fine-tunes; almost no one pretrains from scratch.

The worked examples below use the real Hugging Face transformers library. They need transformers and torch installed locally, so treat them as a preview of the real API — the expected output is shown in comments.

# REAL LLMs in practice: Hugging Face 'transformers'
# (This needs the transformers + torch libraries installed locally;
#  it is shown here to illustrate the real API and its output.)

from transformers import pipeline

# A text-generation pipeline downloads a pretrained decoder-only model.
generator = pipeline("text-generation", model="gpt2")

result = generator(
    "The future of artificial intelligence is",
    max_new_tokens=20,
    temperature=0.7,     # same temperature idea you coded above
    do_sample=True,
)

print(result[0]["generated_text"])

# Expected output (varies because do_sample=True):
# The future of artificial intelligence is bright, with new tools
# helping people learn faster and build smarter applications every day.
# PRETRAINING vs FINE-TUNING
# Pretraining: learn language from trillions of tokens (months, $millions).
# Fine-tuning: cheaply adapt that pretrained model to YOUR task.

from transformers import AutoModelForCausalLM, AutoTokenizer

# 1) Load a model that was already PRETRAINED on the open internet.
tokenizer = AutoTokenizer.from_pretrained("distilgpt2")
model = AutoModelForCausalLM.from_pretrained("distilgpt2")

print("Loaded a pretrained model:", model.config.model_type)
print("Vocabulary size (tokens):", model.config.vocab_size)
print("Parameters (weights):", sum(p.numel() for p in model.parameters()))

# 2) FINE-TUNING would continue training on your own examples so the model
#    speaks in your style/domain. You almost never pretrain from scratch.

# Expected output:
# Loaded a pretrained model: gpt2
# Vocabulary size (tokens): 50257
# Parameters (weights): 82176768   (~82 million)

7 Emergent Abilities & Scaling Laws

Scaling laws are the surprising finding that a model's prediction error falls in a smooth, predictable way as you add more parameters, more data, and more compute. The Chinchilla result added an important twist: for a fixed compute budget, you often want a smaller model trained on more data — roughly 20 tokens of training data per parameter.

Emergent abilities are skills that barely exist in small models but suddenly appear once a model crosses a size threshold — things like following instructions, doing multi-step arithmetic, or in-context learning. Nobody programmed these in; they emerge from scale. This is why bigger models can feel qualitatively, not just quantitatively, smarter.

⚠️ Limitations: Hallucination & What LLMs Can't Do

Because an LLM optimises for plausible next tokens — not for truth — it will sometimes produce confident, fluent statements that are simply wrong. This is called hallucination. There's no fact-checker inside; if a made-up citation or API method "sounds right", the model may emit it.

The practical fix: verify important facts, ground the model with retrieval (RAG) or tools, and lower the temperature when you need consistency.

! Common Errors (And How to Fix Them)

❌ Output cut off / "maximum context length exceeded"

Your prompt plus reply went past the context window, so the model truncated or errored.

# Trim old messages or summarise them before sending.
# Leave room for the reply: reply_budget = context_window - prompt_tokens
# If reply_budget is small, shorten the prompt.

❌ The model confidently states something false (hallucination)

It generated a plausible-sounding but wrong fact, citation, or function name.

# Don't trust facts blindly. Ground the model:
#  - lower the temperature (e.g. 0.2) for factual tasks
#  - give it the source text (retrieval / RAG)
#  - ask it to say "I don't know" when unsure

❌ Output is random / repetitive — wrong temperature

Too-high temperature = nonsense; temperature of 0 = identical, repetitive text.

# Use ~0.2 for facts/code, ~0.7-1.0 for creative writing.
generator(prompt, temperature=0.7, do_sample=True)  # balanced
generator(prompt, temperature=0.0)                  # deterministic

❌ Same task, different wording, different answer (prompt sensitivity)

A tiny rephrase changed the result because the model reacts to exact tokens.

# Be explicit and consistent. Spell out the format you want:
# "Answer in one sentence." / "Return valid JSON only."
# Provide 1-2 examples (few-shot) to anchor the style.

📋 Quick Reference

TermWhat It MeansIn One Line
TokenA chunk of text (word piece)~4 chars of English
Next-token predictionOutput a probability per tokenPick, append, repeat
Decoder-only transformerArchitecture behind GPT/LLaMACausal (past-only) attention
Context windowMax tokens held at oncePrompt + reply combined
Parameters / weightsThe trained numbersWhere capacity lives
LogitsRaw pre-softmax scoressoftmax → probabilities
TemperatureRandomness knobLow = safe, high = creative
PretrainingLearn language from scratchMonths, $millions
Fine-tuningAdapt a pretrained modelCheap, task-specific
HallucinationConfident but false outputPlausible ≠ true

🎯 Mini Challenge: Build a Tiny Next-Token Generator

Time to fade the scaffolding. Starting from the word "the", greedily generate three more tokens and print the finished sentence. The starter below gives you only a comment outline — write the logic yourself.

# 🎯 MINI-CHALLENGE: build a tiny next-token generator
# Brief: starting from "the", generate 3 more tokens by GREEDILY picking
# the most likely next token at each step, then print the final sentence.
#
# 1. Make a dict 'transitions' mapping a word -> {nextWord: probability}
#    e.g. "the": {"cat": 0.7, "dog": 0.3}, "cat": {"sat": 0.9, "ran": 0.1}, ...
# 2. Start with current = "the" and an output list ["the"]
# 3. Loop 3 times: pick the highest-probability next word (max + key=.get),
#    append it to output, then set current to that word
# 4. print the sentence with " ".join(output)
#
# ✅ Expected (example): if greedy path is the -> cat -> sat -> down
#    Output: the cat sat down

# your code here

Lesson complete — you can now explain how an LLM works!

You learned that an LLM is a decoder-only transformer doing next-token prediction over tokens within a context window, that its knowledge lives in billions of parameters, and that temperature, softmax, pretraining vs fine-tuning, scaling laws, and hallucination shape what it produces.

Practice quiz

What is the core task an LLM is trained to do?

  • Classify whole documents at once
  • Translate between fixed language pairs only
  • Predict the next token given the preceding text
  • Compress text losslessly

Answer: Predict the next token given the preceding text. An LLM outputs a probability for every possible next token, picks one, appends it, and repeats — autoregressive generation.

Which architecture do GPT, LLaMA, and Mistral use?

  • Decoder-only transformer
  • Encoder-only transformer
  • Recurrent neural network
  • Convolutional network

Answer: Decoder-only transformer. These are decoder-only transformers that generate text left to right using causal self-attention.

What does causal (masked) self-attention enforce?

  • Each token can see all other tokens
  • Tokens attend only to the next token
  • Attention is disabled during generation
  • Each token can only attend to tokens before it

Answer: Each token can only attend to tokens before it. Causal attention masks the future so a token attends only to itself and earlier tokens, enabling left-to-right generation.

Roughly how much English text does one token represent?

  • About 1 character
  • About 4 characters (~0.75 of a word)
  • Exactly one word
  • About 20 characters

Answer: About 4 characters (~0.75 of a word). A rough rule of thumb is one token per four characters of English, about three-quarters of a word.

What is the context window?

  • The maximum number of tokens (prompt plus reply) the model holds at once
  • The number of layers in the model
  • The size of the vocabulary
  • The temperature setting

Answer: The maximum number of tokens (prompt plus reply) the model holds at once. The context window caps how many tokens the model can consider; older text falls out of view when exceeded.

What are a model's parameters (weights)?

  • A database of facts it queries
  • The list of allowed tokens
  • The numbers adjusted during training that map tokens to output probabilities
  • The prompts given by users

Answer: The numbers adjusted during training that map tokens to output probabilities. Parameters are the trained numbers; there is no fact database — knowledge is distributed across the weights.

What does softmax do to a model's logits?

  • Removes the lowest scores
  • Turns raw scores into probabilities that sum to 1
  • Sorts the tokens alphabetically
  • Divides them by the context length

Answer: Turns raw scores into probabilities that sum to 1. Softmax converts the raw logit scores into a valid probability distribution over the vocabulary.

How does a LOW temperature affect generated text?

  • Makes output more random and creative
  • Increases the context window
  • Adds more parameters
  • Sharpens the distribution so the top token dominates — safe and consistent

Answer: Sharpens the distribution so the top token dominates — safe and consistent. Low temperature concentrates probability on the highest-scoring token, giving consistent, less random output.

What is the difference between pretraining and fine-tuning?

  • Fine-tuning learns language from scratch; pretraining is cheap
  • Pretraining learns general language from trillions of tokens; fine-tuning cheaply adapts it to a task
  • They are the same process
  • Pretraining only runs at inference time

Answer: Pretraining learns general language from trillions of tokens; fine-tuning cheaply adapts it to a task. Pretraining is the expensive general-language phase; fine-tuning adapts that model to a specific task or style.

Why do LLMs sometimes hallucinate?

  • They run out of memory
  • Their temperature is always 0
  • They optimise for plausible next tokens, not for truth
  • They always cite their sources

Answer: They optimise for plausible next tokens, not for truth. An LLM has no internal fact-checker; when a plausible continuation is false, it emits a confident, fluent error.

Continue this course

Frequently asked questions

How does a large language model actually work?

An LLM is a decoder-only transformer trained to predict the next token. It turns your text into tokens, runs them through layers of self-attention and feed-forward weights, and outputs a probability for every possible next token. It picks one, appends it, and repeats — generating text one token at a time.

What is a token and what is the context window?

A token is a chunk of text — often a word piece — and roughly one token equals about four characters of English. The context window is the maximum number of tokens (prompt plus reply) the model can consider at once. Text beyond that limit is dropped, so the model effectively forgets the earliest content.

What does temperature do when generating text?

Temperature scales the model's raw scores (logits) before softmax. A low temperature (near 0) makes the model confident and repetitive — it almost always picks the top token. A high temperature flattens the probabilities, making output more random and creative but more error-prone.

What is the difference between pretraining and fine-tuning?

Pretraining teaches a model general language from trillions of tokens — it costs months of compute and millions of dollars. Fine-tuning cheaply adapts that already-pretrained model to your specific task or style using a small dataset. Almost everyone fine-tunes; almost no one pretrains from scratch.

Why do LLMs hallucinate or make things up?

An LLM optimises for plausible next tokens, not for truth. It has no built-in fact database, so when the most likely continuation sounds right but is wrong, it produces a confident, fluent falsehood — a hallucination. Verify important facts and prefer retrieval-augmented or grounded approaches for accuracy.

What are emergent abilities and scaling laws?

Scaling laws describe how loss falls predictably as you add parameters, data, and compute. Emergent abilities are skills (like multi-step reasoning or following instructions) that barely appear in small models but show up once a model crosses a size threshold — they emerge with scale rather than being explicitly programmed.

Related lessons