Fine-Tuning LLMs: LoRA & QLoRA

Reviewed & published by Brayan K

By the end of this lesson you'll know when to prompt, when to use RAG, and when to fine-tune — and how to prepare data and train an adapter without melting your GPU.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

🌍 Real-World Analogy: Specialising a Generalist Employee

A pre-trained LLM is like a sharp new graduate who knows a bit about everything. You have three ways to get the work you need from them, in increasing cost:

LoRA is the smart twist: instead of re-educating the whole employee, you teach them one focused playbook (a tiny adapter) they slot in for your task — and can pop out again for the next client.

1 Prompting vs RAG vs Fine-Tuning — Choose the Cheapest That Works

Fine-tuning is powerful, but it is the last tool you should reach for, not the first. Most problems are solved more cheaply by writing a better prompt or by retrieval. Use this rule of thumb:

The runnable example below builds a tiny decision helper and shows how few parameters LoRA actually trains. Run it and read every comment.

# WHEN should you fine-tune? Use a simple decision helper, then estimate cost.

def recommend(needs_private_facts, needs_new_behaviour, examples_available):
    """Return the cheapest approach that fits the need."""
    if needs_private_facts and not needs_new_behaviour:
        return "RAG"          # inject documents at query time, no training
    if needs_new_behaviour and examples_available >= 500:
        return "Fine-tune"    # teach a new style/skill with many examples
    return "Prompting"        # just write a better prompt — fastest, free

cases = [
    ("Answer questions about OUR docs", True,  False, 0),
    ("Always reply in our brand voice", False, True,  2000),
    ("One-off summary of an article",   False, False, 0),
]
for desc, facts, behaviour, n in cases:
    print(f"{recommend(facts, behaviour, n):>10}  <-  {desc}")

print()

# LoRA only trains two small matrices instead of a full weight matrix.
# Full matrix has d*d numbers; LoRA trains 2*d*r where r (rank) is small.
def lora_savings(d, r):
    full = d * d
    lora = 2 * d * r
    return full, lora, full / lora

print(f"{'dim':>6} {'rank':>5} {'full':>12} {'LoRA':>10} {'smaller by':>11}")
for d in (768, 4096):
    for r in (8, 16):
        full, lora, factor = lora_savings(d, r)
        print(f"{d:>6} {r:>5} {full:>12,} {lora:>10,} {factor:>10.0f}x")

# Expected output:
#  Prompting  <-  One-off summary of an article
#        RAG  <-  Answer questions about OUR docs
#  Fine-tune  <-  Always reply in our brand voice
#
#    dim  rank         full       LoRA  smaller by
#    768     8      589,824     12,288         48x
#    768    16      589,824     24,576         24x
#   4096     8   16,777,216     65,536        256x
#   4096    16   16,777,216    131,072        128x

2 Full Fine-Tuning vs Parameter-Efficient (LoRA, QLoRA, PEFT)

Full fine-tuning updates every weight in the model. It gives the best quality but is brutally expensive: training a 70B model in 32-bit needs hundreds of gigabytes of GPU memory just to hold the weights, gradients, and optimiser state.

PEFT (Parameter-Efficient Fine-Tuning) is the family of methods that avoid this. The most popular is LoRA (Low-Rank Adaptation): freeze the original weights and learn two small matrices whose product is added to a weight matrix. A full weight matrix holds d×d numbers; LoRA trains only 2×d×r, where the rank r (usually 8–16) is tiny. That is often under 1% of the parameters.

QLoRA stacks one more trick on top: load the frozen base model in 4-bit precision (NormalFloat4) instead of 16-bit. The adapters stay in higher precision and do the learning, while the giant frozen base barely takes up room — so a 70B model that needed four A100s can be fine-tuned on a single 48GB GPU.

Here is what a real LoRA setup looks like with Hugging Face peft and transformers. It's read-only (it needs a GPU and the libraries), but notice the trainable% line in the expected output:

And QLoRA is the same idea with one extra config object to quantise the base model to 4-bit:

3 Instruction Tuning, RLHF & DPO — Teaching Behaviour

Instruction tuning is just supervised fine-tuning on instruction → response pairs. It teaches the model to follow commands rather than only continue text. This is the workhorse you'll use most.

RLHF (Reinforcement Learning from Human Feedback) goes further to make answers preferred by people. Humans rank several model answers; a small reward model learns to predict those rankings; then the LLM is optimised to score highly under that reward model. It works well but has many moving parts.

DPO (Direct Preference Optimisation) reaches a similar place far more simply. You give it pairs of a chosen answer and a rejected answer, and it trains the model directly to prefer the chosen one — no separate reward model, no reinforcement-learning loop. For most teams today, DPO is the easier path to alignment.

4 Dataset Prep — The Part That Actually Decides Quality

A fine-tune is only as good as its data. You build a list of records, each with a prompt (the instruction, wrapped in a consistent template) and a completion (the ideal response). The template matters: the model learns to produce whatever follows your ### Response: marker, so it must be identical in training and at inference.

Good data is consistent (same format every row), clean (no empty or contradictory answers), and representative (covers the cases you'll actually see). The runnable example below formats raw pairs into records and then shows the core idea of training — a single weight learning by gradient descent. Run it:

# Fine-tuning starts with DATA: instruction -> response pairs.
# Here we format raw examples into clean training records — no libraries needed.

raw_examples = [
    ("Translate to French: Hello", "Bonjour"),
    ("Summarise: The cat sat on the mat.", "A cat sat on a mat."),
    ("Capital of Japan?", "Tokyo"),
]

# A simple chat-style template the model learns to complete.
def to_record(instruction, response):
    return {
        "prompt": "### Instruction:\n" + instruction + "\n\n### Response:\n",
        "completion": response,
    }

dataset = [to_record(i, r) for i, r in raw_examples]

print("Built", len(dataset), "training records\n")
for record in dataset:
    # repr() shows the \n characters so you can see the exact text the model sees
    print("PROMPT    :", repr(record["prompt"]))
    print("COMPLETION:", repr(record["completion"]))
    print("-" * 40)

# --- A tiny "gradient step" toy update of a single weight -------------------
# Real training nudges millions of weights; the IDEA is just this loop:
# 1) predict, 2) measure error, 3) step the weight toward less error.
weight = 0.0          # the parameter we are "training"
target = 1.0          # what we want the model to output
learning_rate = 0.3   # how big each step is

print("\nToy gradient descent on one weight:")
for step in range(1, 6):
    prediction = weight              # trivial model: output = weight
    error = prediction - target      # how wrong we are
    gradient = 2 * error             # derivative of (prediction - target)**2
    weight = weight - learning_rate * gradient   # step downhill
    print(f"  step {step}: weight={weight:.4f}  error={error:+.4f}")

print(f"\nFinal weight {weight:.4f} (target was {target}) — it learned!")

# Expected output:
# Built 3 training records
#
# PROMPT    : '### Instruction:\nTranslate to French: Hello\n\n### Response:\n'
# COMPLETION: 'Bonjour'
# ----------------------------------------
# ...
# Toy gradient descent on one weight:
#   step 1: weight=0.6000  error=-1.0000
#   ...
# Final weight ~1.0000 (target was 1.0) — it learned!

That five-line loop is the whole secret of training, scaled up: predict, measure the error, nudge each weight a little against the error, repeat. Fine-tuning runs this over billions of weights (or, with LoRA, just the adapter's few million).

5 Evaluation & Overfitting — Did It Actually Get Better?

Always hold back a validation set the model never trains on. During training you watch two numbers: training loss (error on data it sees) and validation loss (error on data it doesn't). When training loss keeps falling but validation loss starts rising, you are overfitting — the model is memorising your examples instead of learning the pattern.

🎯 Your Turn 1: Build Training Records

Finish the to_record function so each FAQ pair becomes a clean instruction/response record. Fill in the blanks marked ___ and check your output against the expected result in the comments.

# 🎯 YOUR TURN — fill in the blanks marked with ___

# A FAQ bot needs training records. Format each (question, answer) pair into
# the chat template the model will learn to complete.

faq = [
    ("How do I reset my password?", "Click 'Forgot password' on the login page."),
    ("What are your hours?", "We are open 9am to 5pm, Monday to Friday."),
]

def to_record(question, answer):
    # 👉 build the prompt: "### Instruction:\n" + the question + "\n\n### Response:\n"
    prompt = ___
    # 👉 the completion is simply the answer
    completion = ___
    return {"prompt": prompt, "completion": completion}

dataset = [to_record(q, a) for q, a in faq]   # 👉 build a record for every pair

print("Records:", len(dataset))
print(dataset[0]["prompt"] + dataset[0]["completion"])

# ✅ Expected output:
# Records: 2
# ### Instruction:
# How do I reset my password?
#
# ### Response:
# Click 'Forgot password' on the login page.

🎯 Your Turn 2: Prompt, RAG, or Fine-Tune?

Complete the choose function so it returns the right approach for each task. Fill in the blanks marked ___.

# 🎯 YOUR TURN — fill in the blanks marked with ___

# Pick the cheapest approach for each task.
# Rules: private facts only  -> "RAG"
#        new behaviour + lots of examples -> "Fine-tune"
#        otherwise            -> "Prompting"

def choose(private_facts, new_behaviour, n_examples):
    if private_facts and not new_behaviour:
        return ___                 # 👉 return the string for document lookup
    if new_behaviour and n_examples >= 500:
        return ___                 # 👉 return the string for training the model
    return "Prompting"

print(choose(True,  False, 0))     # querying your internal wiki
print(choose(False, True,  3000))  # teaching a consistent legal-writing style
print(choose(False, False, 0))     # rephrasing one email politely

# ✅ Expected output:
# RAG
# Fine-tune
# Prompting

🎯 Mini-Challenge: Clean the Dataset (faded)

No blanks this time — just a brief and an outline. Drop the bad records and format the rest. Use the expected output to check yourself.

# 🎯 MINI-CHALLENGE: Spot the bad training records
#
# 1. Start with a list of (instruction, response) pairs — include at least one
#    BAD pair where the response is an empty string "".
# 2. Write a function clean(pairs) that keeps only pairs where BOTH the
#    instruction and the response are non-empty after .strip().
# 3. Format the survivors into records: {"prompt": ..., "completion": ...}
#    using the "### Instruction:\n...\n\n### Response:\n" template.
# 4. Print how many were dropped and how many records remain.
#
# ✅ Expected (with 1 empty response out of 4 pairs):
#    Dropped 1 bad pair(s), kept 3 records

# your code here

! Common Errors (And How to Fix Them)

❌ Fine-tuning when prompting would do

You spin up GPUs and label data, but a one-line prompt tweak already solved it.

✅ Fix: Always try prompting first, then RAG. Only fine-tune when a careful prompt still can't make the behaviour reliable.

❌ Too little (or low-quality) data

A dozen examples can't teach a behaviour; the model barely shifts or learns the noise.

✅ Fix: Aim for a few hundred to a few thousand clean, consistent examples. Quality and consistency beat raw volume.

After training, the model nails your task but is suddenly worse at general questions.

✅ Fix: Use a small learning rate, fewer epochs, and LoRA (which leaves the base weights frozen). Mix in some general examples if needed.

Training loss keeps dropping while validation loss climbs — it's memorising, not learning.

✅ Fix: Stop early when validation loss turns up, hold out a validation set, and train for only 1–3 epochs.

Great validation scores, terrible real answers — because inference uses a different template than training.

✅ Fix: Keep the prompt template (markers, whitespace, newlines) byte-for-byte identical at train and inference time.

📋 Quick Reference: Prompt vs RAG vs Fine-Tune

ApproachBest forChanges weights?Cost / Speed
PromptingMost tasks; quick iterationNoFree, instant
RAGPrivate / changing factsNoLow, fast
LoRA / QLoRANew behaviour, <1% params trainedAdapter onlyModerate
Full fine-tuneMax quality, deep changesAll weightsHigh, slow
RLHF / DPOAligning to human preferenceYes (after instruction tuning)High

❓ Frequently Asked Questions

Lesson complete — you can now reason about fine-tuning like a practitioner!

You can pick prompting, RAG, or fine-tuning for a task; explain LoRA, QLoRA, and PEFT; describe instruction tuning, RLHF, and DPO; prepare clean instruction/response data; and guard against overfitting, catastrophic forgetting, and format mismatch.

Practice quiz

What is fine-tuning a large language model?

  • Training a model from scratch on the whole internet
  • Writing a longer prompt
  • Continuing training of a pretrained model on your own examples to adapt its behaviour
  • Compressing the model to fewer bits

Answer: Continuing training of a pretrained model on your own examples to adapt its behaviour. Fine-tuning starts from a capable base model and nudges its weights with a small, focused dataset.

When should you prefer RAG over fine-tuning?

  • When the model needs private or changing FACTS
  • When you need a new consistent behaviour or style
  • When you have no GPU at all
  • When you want to reduce model size

Answer: When the model needs private or changing FACTS. RAG injects documents at query time, so it adds knowledge that can change without retraining. Fine-tuning changes behaviour.

What does LoRA (Low-Rank Adaptation) do?

  • Updates every weight in the model
  • Quantises the model to 4-bit
  • Removes attention layers
  • Freezes the base weights and trains two small low-rank adapter matrices

Answer: Freezes the base weights and trains two small low-rank adapter matrices. LoRA freezes the base model and learns small low-rank matrices, often under 1% of the parameters.

What extra trick does QLoRA add on top of LoRA?

  • It trains all weights in full precision
  • It loads the frozen base model in 4-bit precision to save memory
  • It removes the adapter matrices
  • It uses a larger learning rate

Answer: It loads the frozen base model in 4-bit precision to save memory. QLoRA loads the frozen base in 4-bit (NormalFloat4), so even a 70B model can fine-tune on a single GPU.

What does PEFT stand for?

  • Parameter-Efficient Fine-Tuning
  • Pretrained Embedding Fine-Tuning
  • Partial Encoder Feature Transfer
  • Precision Enhanced Forward Training

Answer: Parameter-Efficient Fine-Tuning. PEFT is the family of Parameter-Efficient Fine-Tuning methods, of which LoRA is the most popular.

What is instruction tuning?

  • Reinforcement learning with a reward model
  • Quantising the model to int8
  • Supervised fine-tuning on instruction/response pairs so the model follows commands
  • Retrieving documents at query time

Answer: Supervised fine-tuning on instruction/response pairs so the model follows commands. Instruction tuning is supervised fine-tuning on instruction-response pairs, teaching the model to follow commands.

How does DPO differ from RLHF?

  • DPO needs a separate reward model and an RL loop
  • DPO trains directly on chosen vs rejected answer pairs, with no separate reward model
  • DPO does not change the model weights
  • DPO only works for image models

Answer: DPO trains directly on chosen vs rejected answer pairs, with no separate reward model. DPO (Direct Preference Optimisation) trains directly on preferred/rejected pairs, avoiding RLHF's reward model and RL loop.

How do you detect overfitting during fine-tuning?

  • Training loss rises while validation loss falls
  • Both losses fall together forever
  • The model size increases
  • Training loss keeps falling while validation loss starts rising

Answer: Training loss keeps falling while validation loss starts rising. When training loss drops but validation loss climbs, the model is memorising the training data rather than generalising.

What is catastrophic forgetting?

  • The model crashes during training
  • The model gets great at the new task but worse at general skills
  • The optimiser loses the gradients
  • The dataset is deleted after training

Answer: The model gets great at the new task but worse at general skills. Training too hard or on too narrow data can make the model lose its previously learned general abilities.

Why must the prompt template match between training and inference?

  • To save disk space
  • To reduce the number of parameters
  • Because the model learns to respond to the exact cue it saw in training
  • Because libraries require it

Answer: Because the model learns to respond to the exact cue it saw in training. If training uses '### Response:' but inference uses a different marker, the model never sees its learned cue and quality collapses.

Continue this course

Frequently asked questions

What is fine-tuning a large language model?

Fine-tuning continues training a pre-trained model on your own instruction/response examples so it adapts its behaviour, style, or format. You are not training from scratch — you start from a capable base model and nudge its weights with a small, focused dataset.

When should I fine-tune instead of prompting or using RAG?

Prompt first — it is free and instant, and a good prompt solves most tasks. Use RAG (retrieval-augmented generation) when the model needs private or changing FACTS, because it injects documents at query time without training. Fine-tune only when you need a new consistent BEHAVIOUR or style and you have hundreds of high-quality examples that prompting cannot reliably produce.

What is the difference between full fine-tuning and LoRA/QLoRA?

Full fine-tuning updates every weight, which is expensive and needs huge GPU memory. LoRA (a parameter-efficient method, PEFT) freezes the base model and trains two tiny low-rank adapter matrices — often under 1% of the parameters. QLoRA goes further by loading the frozen base model in 4-bit, so even a 70B model can be fine-tuned on a single GPU.

What are instruction tuning, RLHF, and DPO?

Instruction tuning is supervised fine-tuning on instruction/response pairs so the model follows commands. RLHF (Reinforcement Learning from Human Feedback) then aligns the model to human preferences using a reward model. DPO (Direct Preference Optimisation) reaches a similar result more simply by training directly on pairs of preferred and rejected answers, without a separate reward model.

How much data do I need, and how do I avoid overfitting?

Quality beats quantity: a few hundred to a few thousand clean, consistent examples often beat tens of thousands of noisy ones. To avoid overfitting, hold out a validation set, watch for validation loss rising while training loss falls, train for only 1-3 epochs, and use a small learning rate so the model does not forget its general skills (catastrophic forgetting).