Advanced NLP: BERT, T5 & LLaMA

Reviewed & published by Brayan K

By the end you'll explain how transformers read context, tell BERT (understanding) apart from GPT (generation), fine-tune a pretrained model, and run NER, question answering, and summarization with a single line of Hugging Face code.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

🌍 Real-World Analogy: Understanding Meaning, Not Just Words

Imagine two readers handed the sentence "I went to the bank." A dictionary-only reader sees the word "bank" and shrugs — it could be money or a river. A thoughtful reader looks at the words around it ("I fished off the …") and instantly knows which one you mean.

That thoughtful reader is a transformer. It doesn't store one fixed meaning per word — it builds a contextual embedding, a fresh vector for each word shaped by its neighbours. That is the whole reason modern NLP understands meaning, not just spelling. BERT is the detective reading the full report in both directions to understand; GPT is the storyteller writing one word at a time; T5 is the translator who reads everything, then writes a fresh version.

1 Contextual Embeddings — Words With a Meaning That Moves

An embedding is a list of numbers (a vector) that represents a piece of text so a computer can do maths with meaning. The breakthrough of transformers is that the embedding is contextual: the same word gets a different vector depending on the sentence.

You compare two vectors with cosine similarity — a score from 0 (unrelated) to 1 (same direction). This single tool powers semantic search: find the document whose vector points the same way as your question, even if it shares no exact words. Run the worked example below and read the comments — every result is stated inline.

# Why "context" is the heart of modern NLP
#
# Older NLP gave every word ONE fixed vector (word2vec, GloVe).
# So "bank" had the same meaning in "river bank" and "savings bank".
# Transformers (BERT, GPT) produce a CONTEXTUAL embedding: the vector
# for a word changes depending on the surrounding words.

# We fake "contextual" vectors here so you can see the idea without a GPU.
bank_money = [0.9, 0.1, 0.2]   # "bank" near: money, account, loan
bank_river = [0.1, 0.9, 0.8]   # "bank" near: river, water, fishing

def dot(a, b):
    return sum(x * y for x, y in zip(a, b))

def magnitude(v):
    return sum(x * x for x in v) ** 0.5

def cosine_similarity(a, b):
    # 1.0 = identical direction, 0.0 = unrelated
    return dot(a, b) / (magnitude(a) * magnitude(b))

print("Same word 'bank', two contexts:")
print("  finance-bank vs river-bank:", round(cosine_similarity(bank_money, bank_river), 2))
print()

# A query vector that clearly means "finance"
query = [0.85, 0.15, 0.25]
print("Query (finance) vs finance-bank:", round(cosine_similarity(query, bank_money), 2))
print("Query (finance) vs river-bank:  ", round(cosine_similarity(query, bank_river), 2))

# Expected output:
# Same word 'bank', two contexts:
#   finance-bank vs river-bank: 0.41
#
# Query (finance) vs finance-bank: 1.0
# Query (finance) vs river-bank:   0.45

2 BERT vs GPT — Encoder vs Decoder

BERT is an encoder. It is pretrained with masked language modeling (MLM): random words are hidden behind a [MASK] token and BERT must guess them using context from both sides. Because it reads in both directions, it is brilliant at understanding — classification, NER, similarity, and question answering.

GPT is a decoder. It is pretrained to predict the next token using only the words to its left. That left-to-right view makes it a natural generator — chat, writing, code. T5 combines both: an encoder reads the input, a decoder writes the output, which is the perfect shape for summarization and translation.

3 Transfer Learning & Fine-Tuning

You almost never train an NLP model from zero. Instead you use transfer learning: take a model already pretrained on billions of words, then fine-tune it on your much smaller, task-specific dataset. The model already "knows" language; you only teach it your specific job.

The recipe for fine-tuning BERT on a task like NER or sentiment is short:

# Fine-tuning recipe (conceptual)
# 1. Load a pretrained model      -> bert-base-uncased
# 2. Add a small task head        -> e.g. a linear classification layer
# 3. Train a few epochs (3-5)     -> on your labelled data only
# 4. Use a tiny learning rate     -> 2e-5 with warmup, so you don't wipe
#                                     out what the model already learned

With a few thousand labelled examples you can reach accuracy that would have needed millions of examples to train from scratch. That is the power — and the economy — of transfer learning.

4 NER, Question Answering & Summarization in One Line

Three of the most common NLP jobs each have a ready-made pretrained model:

Hugging Face's pipeline() downloads and wires up the right model for you. Study the worked example below — each call is followed by its # Expected output.

# In real projects you do NOT build transformers from scratch.
# Hugging Face's pipeline() loads a pretrained model in ONE line.
# (This needs 'pip install transformers' + a download, so study the
#  Expected output rather than running it in the browser sandbox.)

from transformers import pipeline

# 1) Named Entity Recognition — tag people, orgs, places
ner = pipeline("ner", grouped_entities=True)
print(ner("Tim Cook is the CEO of Apple in California."))
# Expected output:
# [{'entity_group': 'PER', 'word': 'Tim Cook', 'score': 0.99},
#  {'entity_group': 'ORG', 'word': 'Apple', 'score': 0.99},
#  {'entity_group': 'LOC', 'word': 'California', 'score': 0.99}]

# 2) Question Answering — extract the answer span from a context
qa = pipeline("question-answering")
print(qa(question="Who created Python?",
         context="Python was created by Guido van Rossum in 1991."))
# Expected output:
# {'answer': 'Guido van Rossum', 'score': 0.98, 'start': 25, 'end': 41}

# 3) Summarization — shorten a long passage
summarizer = pipeline("summarization")
print(summarizer("The flood damaged hundreds of homes. Emergency crews "
                  "worked through the night to rescue stranded families and "
                  "deliver food and clean water to the affected region.",
                  max_length=20, min_length=8))
# Expected output:
# [{'summary_text': 'Emergency crews rescued families after the flood.'}]

🎯 Your Turn 1: Build a Rule-Based NER Tagger

Real NER uses BERT, but the idea is simple enough to do by hand: look each word up in a dictionary of known entities and label it. Fill in the one blank, then run it and check your output against the comment.

# 🎯 YOUR TURN — a tiny Named Entity Recognition tagger
# Real NER uses BERT, but the IDEA is: look each word up and label it.
# Fill in the blanks marked with ___

# A dictionary mapping known words to their entity TYPE
known = {
    "Apple": "ORG",
    "Google": "ORG",
    "Paris": "LOC",
    "London": "LOC",
    "Alice": "PER",
    "Bob": "PER",
}

sentence = "Alice flew from London to Paris to visit Google"

tags = []
for word in sentence.split():
    # 👉 if the word is a key in 'known', use its type; else label it "O" (outside)
    label = known.get(word, ___)   # 👉 replace ___ with the default label "O"
    tags.append((word, label))

for word, label in tags:
    print(f"{word:>8} -> {label}")

# ✅ Expected output:
#    Alice -> PER
#     flew -> O
#     from -> O
#   London -> LOC
#       to -> O
#    Paris -> LOC
#       to -> O
#    visit -> O
#   Google -> ORG

🎯 Your Turn 2: Semantic Search With Cosine Similarity

This is how a search box "understands" a question: turn the query and each document into vectors, then rank documents by cosine similarity. Complete the cosine formula and confirm the password document scores higher than the weather one.

# 🎯 YOUR TURN — semantic search with cosine similarity
# Transformers turn a whole sentence into ONE vector. Sentences with
# similar MEANING point in a similar direction, even with different words.
# Fill in the blanks marked with ___

def dot(a, b):
    return sum(x * y for x, y in zip(a, b))

def magnitude(v):
    return sum(x * x for x in v) ** 0.5

def cosine_similarity(a, b):
    # 👉 cosine = dot product divided by (magnitude(a) * magnitude(b))
    return dot(a, b) / (___ * magnitude(b))   # 👉 replace ___ with magnitude(a)

# Toy "sentence vectors" (pretend a transformer produced these)
query        = [0.9, 0.1, 0.2]   # "How do I reset my password?"
doc_password = [0.8, 0.2, 0.1]   # "Steps to change your account password"
doc_weather  = [0.1, 0.9, 0.7]   # "Tomorrow will be sunny and warm"

print("query vs password doc:", round(cosine_similarity(query, doc_password), 2))
print("query vs weather doc: ", round(cosine_similarity(query, doc_weather), 2))

# ✅ Expected output:
# query vs password doc: 0.99
# query vs weather doc:  0.3

🎯 Mini-Challenge: Keyword Sentiment Tagger

Time to fade the scaffolding. Only a comment outline is provided — you write the logic. Build a tiny lexicon-based sentiment tagger that decides whether a review is positive, negative, or mixed.

# 🎯 MINI-CHALLENGE: keyword sentiment tagger
# 1. Make a dict 'lexicon' mapping words to "POS" or "NEG"
#    e.g. {"great": "POS", "love": "POS", "terrible": "NEG", "hate": "NEG"}
# 2. Split this review into words:
#       review = "I love this great phone but hate the terrible battery"
# 3. Count how many POS words and how many NEG words appear
# 4. Print "Overall: POSITIVE" if pos > neg, "Overall: NEGATIVE" if neg > pos,
#    otherwise "Overall: MIXED"
#
# ✅ Expected output (with the review above): Overall: MIXED
#    (2 positive words: love, great  —  2 negative words: hate, terrible)

# your code here

5 Common Errors (And How to Fix Them)

Treating a word as one fixed meaning (old word2vec thinking). "Apple" the company and "apple" the fruit are different — a model that ignores context will confuse them.

✅ Fix: Use a transformer model whose embeddings are contextual; never average away the sentence.

❌ Wrong model for the task

Using GPT for plain classification or NER. It works, but it is 10–100× slower and pricier than a fine-tuned BERT, and often less accurate on the label.

✅ Fix: Encoder (BERT) for understanding, decoder (GPT) for generation, encoder-decoder (T5) for rewriting. Use the simplest model that solves the task.

❌ Exceeding the token limit

BERT caps at 512 tokens. Feed a long document straight in and it is silently truncated — the model never sees the end, so your answer or summary is wrong.

✅ Fix: Chunk long text into overlapping windows, or use a long-context model designed for it.

❌ Poor fine-tuning data quality

Fine-tuning on noisy, mislabelled, or imbalanced data. The model faithfully learns your mistakes — garbage in, garbage out — and scores look fine on a flawed test set.

✅ Fix: Clean and balance labels, hold out a trustworthy validation set, and prefer fewer correct examples over many noisy ones.

📋 Quick Reference: BERT vs GPT

FeatureBERT (Encoder)GPT (Decoder)
AttentionBidirectional (both sides)Left-to-right only
PretrainingMasked LM (fill the blank)Next-token prediction
Best atUnderstandingGeneration
Typical tasksClassification, NER, QA, similarityChat, writing, code, reasoning
SeesAll tokens at oncePast tokens only
Examplebert-base-uncasedGPT-4, LLaMA, Mistral

Need to rewrite text (summarize/translate)? Use an encoder-decoder like T5 or BART — it has both.

❓ Frequently Asked Questions

Lesson complete — you understand the NLP model landscape!

You can explain contextual embeddings, distinguish BERT's encoder from GPT's decoder, describe transfer learning and fine-tuning, and run NER, question answering, and summarization with Hugging Face's pipeline(). You also built a rule-based NER tagger and a cosine-similarity search by hand.

Practice quiz

What is a contextual embedding?

  • One fixed vector per word, regardless of sentence
  • A one-hot vector for a word
  • A word vector whose values depend on the surrounding text
  • A list of all words in the vocabulary

Answer: A word vector whose values depend on the surrounding text. Transformers give a word a different vector depending on context — 'bank' differs in 'river bank' vs 'savings bank'.

Which older method gave every word ONE fixed vector?

  • word2vec / GloVe
  • BERT
  • GPT
  • T5

Answer: word2vec / GloVe. word2vec and GloVe produce a single static vector per word, ignoring the sentence context.

Cosine similarity of 1.0 between two vectors means:

  • They are unrelated
  • They are opposite
  • One is the zero vector
  • They point in the same direction

Answer: They point in the same direction. Cosine similarity ranges from 0 (unrelated) to 1 (same direction); 1.0 means identical direction.

BERT is an encoder pretrained with which objective?

  • Next-token prediction
  • Masked language modeling (fill the blank)
  • Image classification
  • Reinforcement learning

Answer: Masked language modeling (fill the blank). BERT masks random tokens and predicts them using context from both sides, making it strong at understanding.

GPT is a decoder pretrained to:

  • Predict the next token left-to-right
  • Fill in masked words
  • Translate between languages only
  • Cluster documents

Answer: Predict the next token left-to-right. GPT predicts the next token using only the words to its left, which makes it a natural text generator.

For a pure understanding task like classification or NER, which model type fits best?

  • A decoder (GPT)
  • A diffusion model
  • An encoder (BERT)
  • A k-NN classifier

Answer: An encoder (BERT). Encoders like BERT read the whole sentence bidirectionally, which is ideal for labelling or finding things in text.

Which architecture is best for summarization and translation?

  • Encoder-only (BERT)
  • Encoder-decoder (T5 / BART)
  • Decoder-only (GPT)
  • A bag-of-words model

Answer: Encoder-decoder (T5 / BART). Encoder-decoder models read the entire input then generate new text — exactly the shape of rewriting tasks.

What is fine-tuning in transfer learning for NLP?

  • Training a model from scratch on your data
  • Deleting layers from the model
  • Encoding text as one-hot vectors
  • Starting from a pretrained model and training it further on your task data

Answer: Starting from a pretrained model and training it further on your task data. Fine-tuning continues training a pretrained model on a smaller task-specific dataset, reaching high accuracy with less data.

What does Named Entity Recognition (NER) do?

  • Summarises a passage
  • Tags words as people, organisations, or locations
  • Translates text
  • Generates new sentences

Answer: Tags words as people, organisations, or locations. NER labels spans of text with entity types such as person (PER), organisation (ORG), and location (LOC).

Why can feeding a very long document into BERT silently fail?

  • BERT cannot read English
  • BERT only accepts images
  • BERT caps at 512 tokens and truncates the rest
  • BERT needs a GPU to read text

Answer: BERT caps at 512 tokens and truncates the rest. BERT has a 512-token limit, so extra text is silently cut off — chunk long inputs or use a long-context model.

Continue this course

Frequently asked questions

What is a contextual embedding in NLP?

It is a vector for a word or sentence whose values depend on the surrounding text. Unlike older fixed embeddings (word2vec, GloVe) where 'bank' always had one vector, a transformer gives 'bank' a different vector in 'river bank' versus 'savings bank', which is why modern NLP understands meaning, not just spelling.

What is the difference between BERT and GPT?

BERT is an encoder trained with masked language modeling: it reads the whole sentence in both directions, so it is best for understanding tasks like classification, NER, and question answering. GPT is a decoder trained to predict the next token left-to-right, so it is best for generating text like chat and writing.

What is transfer learning and fine-tuning for NLP?

Transfer learning means starting from a model already pretrained on huge amounts of text, then fine-tuning it on your smaller task-specific dataset. You add a small task head and train for a few epochs, which reaches high accuracy with far less data than training from scratch.

Which model should I use for summarization or translation?

Use an encoder-decoder model such as T5 or BART. They read the entire input with an encoder and then generate new text with a decoder, which is exactly the shape of summarization and translation tasks.

Do I need to build transformers from scratch?

No. For most tasks you call Hugging Face's pipeline() for instant NER, question answering, or summarization, and use the Trainer API (often with LoRA) to fine-tune. Building a transformer by hand is a learning exercise, not how production NLP is shipped.

Related lessons