Fine-Tuning LLMs: LoRA & QLoRA
Reviewed & published by Brayan K
By the end of this lesson you'll know when to prompt, when to use RAG, and when to fine-tune — and how to prepare data and train an adapter without melting your GPU.
Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
What You'll Learn in This Lesson
- You'll be able to choose between prompting, RAG, and fine-tuning for any task
- You'll understand full fine-tuning vs parameter-efficient LoRA, QLoRA, and PEFT
- You'll explain instruction tuning and how RLHF and DPO align a model
- You'll prepare clean instruction/response training records yourself
- You'll read a real Hugging Face PEFT + Trainer fine-tuning script
- You'll spot and prevent overfitting and catastrophic forgetting
🌍 Real-World Analogy: Specialising a Generalist Employee
A pre-trained LLM is like a sharp new graduate who knows a bit about everything. You have three ways to get the work you need from them, in increasing cost:
- Prompting = leaving a clear sticky note on their desk. Instant, free, changeable — but they forget it the moment the task ends.
- RAG = giving them a filing cabinet of your company documents to look things up in. They stay a generalist, but now answer with your facts.
- Fine-tuning = sending them on a training course so the new skill becomes second nature. Expensive and slow, but afterwards they just do it without being told.
LoRA is the smart twist: instead of re-educating the whole employee, you teach them one focused playbook (a tiny adapter) they slot in for your task — and can pop out again for the next client.
1 Prompting vs RAG vs Fine-Tuning — Choose the Cheapest That Works
Fine-tuning is powerful, but it is the last tool you should reach for, not the first. Most problems are solved more cheaply by writing a better prompt or by retrieval. Use this rule of thumb:
- Prompting: change the model's output by changing the instruction. Zero training, instant to iterate. Try this first, always.
- RAG (retrieval-augmented generation): the model needs facts it doesn't have — your docs, today's prices, a private wiki. You fetch the relevant text and paste it into the prompt at query time. No training; facts can change every minute.
- Fine-tuning: the model needs a new behaviour — a consistent tone, a strict output format, a niche skill — that prompting can't make reliable. You change the weights with many examples.
The runnable example below builds a tiny decision helper and shows how few parameters LoRA actually trains. Run it and read every comment.
# WHEN should you fine-tune? Use a simple decision helper, then estimate cost.
def recommend(needs_private_facts, needs_new_behaviour, examples_available):
"""Return the cheapest approach that fits the need."""
if needs_private_facts and not needs_new_behaviour:
return "RAG" # inject documents at query time, no training
if needs_new_behaviour and examples_available >= 500:
return "Fine-tune" # teach a new style/skill with many examples
return "Prompting" # just write a better prompt — fastest, free
cases = [
("Answer questions about OUR docs", True, False, 0),
("Always reply in our brand voice", False, True, 2000),
("One-off summary of an article", False, False, 0),
]
for desc, facts, behaviour, n in cases:
print(f"{recommend(facts, behaviour, n):>10} <- {desc}")
print()
# LoRA only trains two small matrices instead of a full weight matrix.
# Full matrix has d*d numbers; LoRA trains 2*d*r where r (rank) is small.
def lora_savings(d, r):
full = d * d
lora = 2 * d * r
return full, lora, full / lora
print(f"{'dim':>6} {'rank':>5} {'full':>12} {'LoRA':>10} {'smaller by':>11}")
for d in (768, 4096):
for r in (8, 16):
full, lora, factor = lora_savings(d, r)
print(f"{d:>6} {r:>5} {full:>12,} {lora:>10,} {factor:>10.0f}x")
# Expected output:
# Prompting <- One-off summary of an article
# RAG <- Answer questions about OUR docs
# Fine-tune <- Always reply in our brand voice
#
# dim rank full LoRA smaller by
# 768 8 589,824 12,288 48x
# 768 16 589,824 24,576 24x
# 4096 8 16,777,216 65,536 256x
# 4096 16 16,777,216 131,072 128x2 Full Fine-Tuning vs Parameter-Efficient (LoRA, QLoRA, PEFT)
Full fine-tuning updates every weight in the model. It gives the best quality but is brutally expensive: training a 70B model in 32-bit needs hundreds of gigabytes of GPU memory just to hold the weights, gradients, and optimiser state.
PEFT (Parameter-Efficient Fine-Tuning) is the family of methods that avoid this. The most popular is LoRA (Low-Rank Adaptation): freeze the original weights and learn two small matrices whose product is added to a weight matrix. A full weight matrix holds d×d numbers; LoRA trains only 2×d×r, where the rank r (usually 8–16) is tiny. That is often under 1% of the parameters.
QLoRA stacks one more trick on top: load the frozen base model in 4-bit precision (NormalFloat4) instead of 16-bit. The adapters stay in higher precision and do the learning, while the giant frozen base barely takes up room — so a 70B model that needed four A100s can be fine-tuned on a single 48GB GPU.
Here is what a real LoRA setup looks like with Hugging Face peft and transformers. It's read-only (it needs a GPU and the libraries), but notice the trainable% line in the expected output:
And QLoRA is the same idea with one extra config object to quantise the base model to 4-bit:
3 Instruction Tuning, RLHF & DPO — Teaching Behaviour
Instruction tuning is just supervised fine-tuning on instruction → response pairs. It teaches the model to follow commands rather than only continue text. This is the workhorse you'll use most.
RLHF (Reinforcement Learning from Human Feedback) goes further to make answers preferred by people. Humans rank several model answers; a small reward model learns to predict those rankings; then the LLM is optimised to score highly under that reward model. It works well but has many moving parts.
DPO (Direct Preference Optimisation) reaches a similar place far more simply. You give it pairs of a chosen answer and a rejected answer, and it trains the model directly to prefer the chosen one — no separate reward model, no reinforcement-learning loop. For most teams today, DPO is the easier path to alignment.
4 Dataset Prep — The Part That Actually Decides Quality
A fine-tune is only as good as its data. You build a list of records, each with a prompt (the instruction, wrapped in a consistent template) and a completion (the ideal response). The template matters: the model learns to produce whatever follows your ### Response: marker, so it must be identical in training and at inference.
Good data is consistent (same format every row), clean (no empty or contradictory answers), and representative (covers the cases you'll actually see). The runnable example below formats raw pairs into records and then shows the core idea of training — a single weight learning by gradient descent. Run it:
# Fine-tuning starts with DATA: instruction -> response pairs.
# Here we format raw examples into clean training records — no libraries needed.
raw_examples = [
("Translate to French: Hello", "Bonjour"),
("Summarise: The cat sat on the mat.", "A cat sat on a mat."),
("Capital of Japan?", "Tokyo"),
]
# A simple chat-style template the model learns to complete.
def to_record(instruction, response):
return {
"prompt": "### Instruction:\n" + instruction + "\n\n### Response:\n",
"completion": response,
}
dataset = [to_record(i, r) for i, r in raw_examples]
print("Built", len(dataset), "training records\n")
for record in dataset:
# repr() shows the \n characters so you can see the exact text the model sees
print("PROMPT :", repr(record["prompt"]))
print("COMPLETION:", repr(record["completion"]))
print("-" * 40)
# --- A tiny "gradient step" toy update of a single weight -------------------
# Real training nudges millions of weights; the IDEA is just this loop:
# 1) predict, 2) measure error, 3) step the weight toward less error.
weight = 0.0 # the parameter we are "training"
target = 1.0 # what we want the model to output
learning_rate = 0.3 # how big each step is
print("\nToy gradient descent on one weight:")
for step in range(1, 6):
prediction = weight # trivial model: output = weight
error = prediction - target # how wrong we are
gradient = 2 * error # derivative of (prediction - target)**2
weight = weight - learning_rate * gradient # step downhill
print(f" step {step}: weight={weight:.4f} error={error:+.4f}")
print(f"\nFinal weight {weight:.4f} (target was {target}) — it learned!")
# Expected output:
# Built 3 training records
#
# PROMPT : '### Instruction:\nTranslate to French: Hello\n\n### Response:\n'
# COMPLETION: 'Bonjour'
# ----------------------------------------
# ...
# Toy gradient descent on one weight:
# step 1: weight=0.6000 error=-1.0000
# ...
# Final weight ~1.0000 (target was 1.0) — it learned!That five-line loop is the whole secret of training, scaled up: predict, measure the error, nudge each weight a little against the error, repeat. Fine-tuning runs this over billions of weights (or, with LoRA, just the adapter's few million).
5 Evaluation & Overfitting — Did It Actually Get Better?
Always hold back a validation set the model never trains on. During training you watch two numbers: training loss (error on data it sees) and validation loss (error on data it doesn't). When training loss keeps falling but validation loss starts rising, you are overfitting — the model is memorising your examples instead of learning the pattern.
- Train for few epochs (often 1–3). More passes mostly memorise.
- Use a small learning rate (e.g. 2e-4 for LoRA) so you nudge, not shove.
- Watch for catastrophic forgetting: if the model gets great at your task but worse at everything else, you trained too hard or on too narrow a dataset.
- Judge on real outputs, not just loss — run the model on held-out prompts and read the answers.
🎯 Your Turn 1: Build Training Records
Finish the to_record function so each FAQ pair becomes a clean instruction/response record. Fill in the blanks marked ___ and check your output against the expected result in the comments.
# 🎯 YOUR TURN — fill in the blanks marked with ___
# A FAQ bot needs training records. Format each (question, answer) pair into
# the chat template the model will learn to complete.
faq = [
("How do I reset my password?", "Click 'Forgot password' on the login page."),
("What are your hours?", "We are open 9am to 5pm, Monday to Friday."),
]
def to_record(question, answer):
# 👉 build the prompt: "### Instruction:\n" + the question + "\n\n### Response:\n"
prompt = ___
# 👉 the completion is simply the answer
completion = ___
return {"prompt": prompt, "completion": completion}
dataset = [to_record(q, a) for q, a in faq] # 👉 build a record for every pair
print("Records:", len(dataset))
print(dataset[0]["prompt"] + dataset[0]["completion"])
# ✅ Expected output:
# Records: 2
# ### Instruction:
# How do I reset my password?
#
# ### Response:
# Click 'Forgot password' on the login page.🎯 Your Turn 2: Prompt, RAG, or Fine-Tune?
Complete the choose function so it returns the right approach for each task. Fill in the blanks marked ___.
# 🎯 YOUR TURN — fill in the blanks marked with ___
# Pick the cheapest approach for each task.
# Rules: private facts only -> "RAG"
# new behaviour + lots of examples -> "Fine-tune"
# otherwise -> "Prompting"
def choose(private_facts, new_behaviour, n_examples):
if private_facts and not new_behaviour:
return ___ # 👉 return the string for document lookup
if new_behaviour and n_examples >= 500:
return ___ # 👉 return the string for training the model
return "Prompting"
print(choose(True, False, 0)) # querying your internal wiki
print(choose(False, True, 3000)) # teaching a consistent legal-writing style
print(choose(False, False, 0)) # rephrasing one email politely
# ✅ Expected output:
# RAG
# Fine-tune
# Prompting🎯 Mini-Challenge: Clean the Dataset (faded)
No blanks this time — just a brief and an outline. Drop the bad records and format the rest. Use the expected output to check yourself.
# 🎯 MINI-CHALLENGE: Spot the bad training records
#
# 1. Start with a list of (instruction, response) pairs — include at least one
# BAD pair where the response is an empty string "".
# 2. Write a function clean(pairs) that keeps only pairs where BOTH the
# instruction and the response are non-empty after .strip().
# 3. Format the survivors into records: {"prompt": ..., "completion": ...}
# using the "### Instruction:\n...\n\n### Response:\n" template.
# 4. Print how many were dropped and how many records remain.
#
# ✅ Expected (with 1 empty response out of 4 pairs):
# Dropped 1 bad pair(s), kept 3 records
# your code here! Common Errors (And How to Fix Them)
❌ Fine-tuning when prompting would do
You spin up GPUs and label data, but a one-line prompt tweak already solved it.
✅ Fix: Always try prompting first, then RAG. Only fine-tune when a careful prompt still can't make the behaviour reliable.
❌ Too little (or low-quality) data
A dozen examples can't teach a behaviour; the model barely shifts or learns the noise.
✅ Fix: Aim for a few hundred to a few thousand clean, consistent examples. Quality and consistency beat raw volume.
After training, the model nails your task but is suddenly worse at general questions.
✅ Fix: Use a small learning rate, fewer epochs, and LoRA (which leaves the base weights frozen). Mix in some general examples if needed.
Training loss keeps dropping while validation loss climbs — it's memorising, not learning.
✅ Fix: Stop early when validation loss turns up, hold out a validation set, and train for only 1–3 epochs.
Great validation scores, terrible real answers — because inference uses a different template than training.
✅ Fix: Keep the prompt template (markers, whitespace, newlines) byte-for-byte identical at train and inference time.
📋 Quick Reference: Prompt vs RAG vs Fine-Tune
| Approach | Best for | Changes weights? | Cost / Speed |
|---|---|---|---|
| Prompting | Most tasks; quick iteration | No | Free, instant |
| RAG | Private / changing facts | No | Low, fast |
| LoRA / QLoRA | New behaviour, <1% params trained | Adapter only | Moderate |
| Full fine-tune | Max quality, deep changes | All weights | High, slow |
| RLHF / DPO | Aligning to human preference | Yes (after instruction tuning) | High |
❓ Frequently Asked Questions
Lesson complete — you can now reason about fine-tuning like a practitioner!
You can pick prompting, RAG, or fine-tuning for a task; explain LoRA, QLoRA, and PEFT; describe instruction tuning, RLHF, and DPO; prepare clean instruction/response data; and guard against overfitting, catastrophic forgetting, and format mismatch.
Practice quiz
What is fine-tuning a large language model?
- Training a model from scratch on the whole internet
- Writing a longer prompt
- Continuing training of a pretrained model on your own examples to adapt its behaviour
- Compressing the model to fewer bits
Answer: Continuing training of a pretrained model on your own examples to adapt its behaviour. Fine-tuning starts from a capable base model and nudges its weights with a small, focused dataset.
When should you prefer RAG over fine-tuning?
- When the model needs private or changing FACTS
- When you need a new consistent behaviour or style
- When you have no GPU at all
- When you want to reduce model size
Answer: When the model needs private or changing FACTS. RAG injects documents at query time, so it adds knowledge that can change without retraining. Fine-tuning changes behaviour.
What does LoRA (Low-Rank Adaptation) do?
- Updates every weight in the model
- Quantises the model to 4-bit
- Removes attention layers
- Freezes the base weights and trains two small low-rank adapter matrices
Answer: Freezes the base weights and trains two small low-rank adapter matrices. LoRA freezes the base model and learns small low-rank matrices, often under 1% of the parameters.
What extra trick does QLoRA add on top of LoRA?
- It trains all weights in full precision
- It loads the frozen base model in 4-bit precision to save memory
- It removes the adapter matrices
- It uses a larger learning rate
Answer: It loads the frozen base model in 4-bit precision to save memory. QLoRA loads the frozen base in 4-bit (NormalFloat4), so even a 70B model can fine-tune on a single GPU.
What does PEFT stand for?
- Parameter-Efficient Fine-Tuning
- Pretrained Embedding Fine-Tuning
- Partial Encoder Feature Transfer
- Precision Enhanced Forward Training
Answer: Parameter-Efficient Fine-Tuning. PEFT is the family of Parameter-Efficient Fine-Tuning methods, of which LoRA is the most popular.
What is instruction tuning?
- Reinforcement learning with a reward model
- Quantising the model to int8
- Supervised fine-tuning on instruction/response pairs so the model follows commands
- Retrieving documents at query time
Answer: Supervised fine-tuning on instruction/response pairs so the model follows commands. Instruction tuning is supervised fine-tuning on instruction-response pairs, teaching the model to follow commands.
How does DPO differ from RLHF?
- DPO needs a separate reward model and an RL loop
- DPO trains directly on chosen vs rejected answer pairs, with no separate reward model
- DPO does not change the model weights
- DPO only works for image models
Answer: DPO trains directly on chosen vs rejected answer pairs, with no separate reward model. DPO (Direct Preference Optimisation) trains directly on preferred/rejected pairs, avoiding RLHF's reward model and RL loop.
How do you detect overfitting during fine-tuning?
- Training loss rises while validation loss falls
- Both losses fall together forever
- The model size increases
- Training loss keeps falling while validation loss starts rising
Answer: Training loss keeps falling while validation loss starts rising. When training loss drops but validation loss climbs, the model is memorising the training data rather than generalising.
What is catastrophic forgetting?
- The model crashes during training
- The model gets great at the new task but worse at general skills
- The optimiser loses the gradients
- The dataset is deleted after training
Answer: The model gets great at the new task but worse at general skills. Training too hard or on too narrow data can make the model lose its previously learned general abilities.
Why must the prompt template match between training and inference?
- To save disk space
- To reduce the number of parameters
- Because the model learns to respond to the exact cue it saw in training
- Because libraries require it
Answer: Because the model learns to respond to the exact cue it saw in training. If training uses '### Response:' but inference uses a different marker, the model never sees its learned cue and quality collapses.
Continue this course
- Previous: Tokenization Strategies (BPE, WordPiece, SentencePiece)
- Next: Reinforcement Learning Basics (MDP, Policies, Rewards) — Markov Decision Processes, value functions, policies, and the Bellman equation
- Quick reference: AI & Machine Learning cheat sheet
Frequently asked questions
What is fine-tuning a large language model?
Fine-tuning continues training a pre-trained model on your own instruction/response examples so it adapts its behaviour, style, or format. You are not training from scratch — you start from a capable base model and nudge its weights with a small, focused dataset.
When should I fine-tune instead of prompting or using RAG?
Prompt first — it is free and instant, and a good prompt solves most tasks. Use RAG (retrieval-augmented generation) when the model needs private or changing FACTS, because it injects documents at query time without training. Fine-tune only when you need a new consistent BEHAVIOUR or style and you have hundreds of high-quality examples that prompting cannot reliably produce.
What is the difference between full fine-tuning and LoRA/QLoRA?
Full fine-tuning updates every weight, which is expensive and needs huge GPU memory. LoRA (a parameter-efficient method, PEFT) freezes the base model and trains two tiny low-rank adapter matrices — often under 1% of the parameters. QLoRA goes further by loading the frozen base model in 4-bit, so even a 70B model can be fine-tuned on a single GPU.
What are instruction tuning, RLHF, and DPO?
Instruction tuning is supervised fine-tuning on instruction/response pairs so the model follows commands. RLHF (Reinforcement Learning from Human Feedback) then aligns the model to human preferences using a reward model. DPO (Direct Preference Optimisation) reaches a similar result more simply by training directly on pairs of preferred and rejected answers, without a separate reward model.
How much data do I need, and how do I avoid overfitting?
Quality beats quantity: a few hundred to a few thousand clean, consistent examples often beat tens of thousands of noisy ones. To avoid overfitting, hold out a validation set, watch for validation loss rising while training loss falls, train for only 1-3 epochs, and use a small learning rate so the model does not forget its general skills (catastrophic forgetting).