10 Papers Every AI Engineer Must Know

The foundational research that shaped modern AI — from transformers to retrieval-augmented generation.

Attention Is All You Need (2017)

Vaswani et al. — the paper that launched a thousand models.

Read the paperVaswani et al., “Attention Is All You Need” (2017) — arXiv:1706.03762 ↗

📐
Self-Attention
Every token attends to every other token in parallel.
⚡
Parallelization
Unlike RNNs, all positions processed simultaneously.
🔑
Key Innovation
Q·KT / √dk — scaled dot-product attention.
🏗️
Architecture
Encoder-decoder with multi-head attention + feed-forward layers.
In Plain English

Simple: At a loud dinner party you do not hear every voice equally — you tune in to the two people whose words actually bear on what you are about to say. Every word in a sentence does that, to every other word, at the same moment.

Technical: Scaled dot-product attention scores every query against every key, softmaxes those scores into weights, and returns the weighted sum of the values. It is an all-pairs operation with no recurrence, which is why the whole sequence resolves in one parallel step instead of one timestep at a time (Vaswani et al., 2017).

GPT — Generative Pre-trained Transformers (2018–2024)

The decoder-only architecture that scaled to change the world.

Read the papersGPT-1: “Improving Language Understanding by Generative Pre-Training” (2018) — OpenAI PDF ↗GPT-3: “Language Models are Few-Shot Learners” (2020) — arXiv:2005.14165 ↗GPT-4 Technical Report (2023) — arXiv:2303.08774 ↗

2018
GPT-1
Decoder-only, unsupervised pre-training + supervised fine-tuning.
2019
GPT-2
“Too dangerous to release” — 1.5B params, zero-shot task transfer.
2020
GPT-3
175B params, in-context learning, few-shot prompting.
2023
GPT-4
Multimodal, extended reasoning, tool use.
In Plain English

Simple: Nobody redesigned the engine between 2018 and 2023. They kept the same engine and built it bigger, on more fuel — and past a certain size it started doing tricks nobody had taught it, like following an instruction it had only seen described.

Technical: The GPT line holds the decoder-only, next-token-prediction objective fixed across four generations and scales parameters, data and compute. GPT-3 (2020) is the paper that named the payoff: in-context learning, where few-shot examples in the prompt substitute for gradient updates — a capability the training objective never optimised for.

BERT — Bidirectional Understanding (2018)

BERT’s job is to understand text, not continue it. It was trained by a sophisticated fill-in-the-blank exercise: hide some words, guess them back. The [MASK] sat on the [MASK] — and to answer cat and mat, it has to read the whole sentence, including the part that comes after the gap. That is the difference from GPT, which only ever reads leftwards.

Read the paperDevlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding” (2018) — arXiv:1810.04805 ↗

🔄
It reads both ways
Every word is understood using the words before it and the words after it, at the same time.
🎭
Guess the hidden word masked LM
Blank out about one word in seven, then make the model put them back. Cheap to set up on any text you have, and no human has to label anything.
📝
Do these two sentences belong together? next-sentence prediction
A second exercise: show it two sentences and ask whether the second really followed the first. Later work questioned whether this one earned its keep.
🏆
Why it still matters
Search, classification and ranking still lean on this family. RoBERTa, ALBERT and DistilBERT are all variations on it.
In Plain English

Simple: A writer composing a sentence can only use the words already on the page. A proofreader reading the finished page can use both sides of any gap — which is why proofreading catches things writing cannot. GPT writes; BERT proofreads.

Technical: BERT is an encoder-only transformer with no causal mask, so every position attends to the full sequence in both directions. That makes autoregressive generation impossible and representation quality excellent, which is why the family survives in classification, ranking and retrieval rather than in chat.

LoRA — Low-Rank Adaptation (2021)

Teaching a big model a new job used to mean rewriting the whole textbook. LoRA is a sticky note clipped to the page: the book is left exactly as it was, and the note carries the correction. For a 7-billion-parameter model the note is around 0.1% of the book’s size — which is what turns fine-tuning from a data-centre job into something you can run on one machine.

Read the paperHu et al., “LoRA: Low-Rank Adaptation of Large Language Models” (2021) — arXiv:2106.09685 ↗

🧮
Retrain the Book, or Clip On a Note
Watch what stops being trainable — then count what is left
trainable frozen grid = one weight matrix W  ·  strips = A (d×r) and B (r×d)
r = 8
7,000,000,000
trainable parameters
Every weight in the matrix is being updated. This is what fine-tuning meant before 2021.
Log scale — each equal step is a ten-fold change, so a share this small stays visible.
Illustrative, for a 7-billion-parameter model, using the figures in this slide’s own deep dives. LoRA trains 2×d×r parameters per adapted matrix, so the count is linear in r — that part is exact (Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, 2021). Strip thickness is not to scale: at r=8 the honest share is under 0.1% and would draw as nothing. The counter is the real number.
📉
How small the note is
A fraction of a percent of the model’s parameters are trained. Orders of magnitude fewer than full fine-tuning, so it fits in far less memory.
🔀
One book, many notes
Keep a note per task and swap them in and out. One copy of the model on disk serves every job.
In Plain English

Simple: Sticking the note on is cheap. The reason it works is separate and less obvious: the correction a new job needs is usually simple — a nudge in a few consistent directions, not a rewrite of every page. A short note can hold a simple correction. It could not hold a complicated one.

Technical: LoRA freezes W and learns ΔW = A·B with A of shape d×r and B of shape r×d, so ΔW is constrained to rank at most r. The bet is that the weight update fine-tuning wants has low intrinsic rank. Parameters trained per adapted matrix are 2·d·r, linear in r — and at inference B·A merges back into W, so latency is unchanged (Hu et al., 2021).

PEFT — Parameter Efficient Fine-Tuning

A family of techniques that make large model adaptation practical.

Read the papersAdapters — Houlsby et al., “Parameter-Efficient Transfer Learning for NLP” (2019) — arXiv:1902.00751 ↗Prefix tuning — Li & Liang, “Prefix-Tuning: Optimizing Continuous Prompts for Generation” (2021) — arXiv:2101.00190 ↗

🔧
LoRA
Low-rank weight matrices
📌
Prefix Tuning
Learnable prompt prefixes
🎯
Adapter Layers
Small bottleneck modules
❄️
Frozen Base
Original model stays unchanged
🔄
Task Switching
Hot-swap task-specific weights
📊
Result
Full fine-tuning quality at fraction of cost
In Plain English

Simple: There is more than one place to clip the note. You can write in the margin next to the text (a small module inserted between layers), or you can staple a covering memo to the front so everything after it is read differently (a learned preamble). LoRA is one of these options, not the whole family.

Technical: PEFT is the umbrella for methods that adapt a frozen backbone by training a small set of new parameters. Adapters insert bottleneck MLP modules inside each block; prefix tuning prepends learned key/value vectors to every attention layer; LoRA reparameterises the weight update itself. They differ in where the trainable parameters sit and in whether they add inference cost — adapters and prefixes do, merged LoRA does not.

ViT — Vision Transformer (2020)

Read a picture like a page. Chop the photo into a grid of small square tiles, line the tiles up in reading order, and hand them to a transformer as if each tile were a word. A 224×224 photo cut into 16×16 tiles becomes a 196-“word” sentence — and from there the machinery is the same machinery that reads text.

Read the paperDosovitskiy et al., “An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale” (2020) — arXiv:2010.11929 ↗

🖼️
Tiles instead of words patch embedding
Cut the image into 16×16 squares, flatten each one, and feed the sequence in. Position markers tell the model where each tile sat.
🔄
Nothing about it is vision-specific
The same transformer encoder used for text, unmodified. No convolutions anywhere — which was the surprising part.
📊
It needs a lot of pictures JFT-300M
On a small dataset it loses to a CNN, because a CNN is built already knowing that a cat is a cat wherever it sits in the frame. Given a very large dataset, ViT learns that for itself and pulls ahead.
🌊
One architecture for everything
Once pictures and text run on the same design, a single model can take both. This is the step that made today’s multimodal models possible.
In Plain English

Simple: A child raised on a handful of photos still knows a cat is a cat upside down, or in the corner of the frame — that assumption comes built in. ViT is not born with it. Show it few pictures and it flounders; show it enormously many and it works the rule out itself, then applies it better than the built-in version.

Technical: A CNN hard-codes translation equivariance and locality as architectural priors. ViT removes them: patches are embedded, position is learned, and attention is global from layer one. With less data that missing prior costs it accuracy against a CNN; pre-trained on a very large corpus it overtakes, which is the scaling result the paper reports (Dosovitskiy et al., 2020).

VAE & GANs — Generative Pioneers

Two ways to make a machine produce a face that never existed. The sketcher learns the shape of “face” as a smooth space, so you can slide from one face to another and every point on the way is also a face. The forger never learns that space — it just gets very good at fooling an inspector. Sharper results, but there is no in-between to slide through.

Read the papersVAE — Kingma & Welling, “Auto-Encoding Variational Bayes” (2013) — arXiv:1312.6114 ↗GANs — Goodfellow et al., “Generative Adversarial Networks” (2014) — arXiv:1406.2661 ↗

 
VAE (2013) — the sketcher
GANs (2014) — the forger
How it learns
Squash the input down to a few numbers, then rebuild it from them latent distribution
Two networks compete — one makes fakes, one tries to spot them adversarial training
What you get
A smooth space you can move through — blend two faces, walk between styles
Individual samples that look strikingly real
Weak spot
Outputs come out slightly soft, a side effect of optimising for faithful reconstruction
Training is unstable, and it can collapse to producing the same few outputs
Deep dive
Diffusion models — the next slide — eventually beat both, and they did it by borrowing from each: the smooth, navigable space the sketcher built, with output sharp enough to satisfy the forger’s inspector.
In Plain English

Simple: The sketcher is graded on how closely the redrawing matches the original, so when it is unsure it hedges — and a hedged line is a soft line. The forger is graded only on whether the inspector is fooled, so hedging earns nothing. Each one’s weakness is the direct cost of what it is being graded on.

Technical: A VAE maximises a reconstruction likelihood plus a KL term pulling the latent posterior towards a prior; the pixel-wise reconstruction loss is what averages away high-frequency detail, and the KL term is what buys the continuous latent space you can interpolate through. A GAN optimises a minimax objective against a discriminator with no reconstruction term at all — hence sharpness, and hence mode collapse and unstable training as the standing failure modes.

Diffusion Models (2020–2022)

Un-blurring. Adding fog to a photograph is free — you need no skill and no training to make a picture worse. So the model is never taught to draw. It is taught only to remove one step of fog. Repeat that a few dozen times, starting from nothing but fog, and a picture comes out. Drag the slider to walk the fog in, then read it right to left to see what the model actually does.

Read the papersHo et al., “Denoising Diffusion Probabilistic Models” (2020) — arXiv:2006.11239 ↗Latent diffusion — Rombach et al., “High-Resolution Image Synthesis with Latent Diffusion Models” (2022) — arXiv:2112.10752 ↗

🌫️
Add the Fog, Then Take It Away
Six steps from photograph to pure noise — and back
step 0
Six steps here for legibility; the original formulation used on the order of 1,000 (Ho et al., Denoising Diffusion Probabilistic Models, 2020); later samplers cut generation to roughly 20–50 steps.
🎨
Do it small, then enlarge latent diffusion
Fifty passes over a full-resolution image is ruinous. Shrink it first, do all the work on the small version, expand at the end — the step that put image generation on ordinary hardware.
🏆
What it unlocked
Every text-to-image tool you have heard of runs on this idea. Steering it with a text prompt is a matter of telling the denoiser what it is meant to be uncovering.
In Plain English

Simple: Un-blurring in one heroic go is a job for an artist. Un-blurring by a hair, over and over, is a job for a technician — and a technician you can actually train, because you can manufacture unlimited practice material for free just by blurring photographs you already have.

Technical: The forward process is a fixed Gaussian noising schedule with no learned parameters, so training pairs are generated for nothing and any timestep can be sampled directly. The network learns only to predict the noise added at a given step; sampling subtracts that prediction iteratively. Latent diffusion runs the loop in a compressed autoencoder space rather than pixel space, which is the change that brought the cost down to consumer hardware (Ho et al., 2020; Rombach et al., 2022).

RAG — Retrieval-Augmented Generation (2020)

Lewis et al. — grounding LLMs in real, retrievable knowledge.

Read the paperLewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks” (2020) — arXiv:2005.11401 ↗

The Problem
Knowledge Cutoffs & Hallucination
LLMs have knowledge cutoffs and hallucinate facts they don’t know.
The Solution
Retrieve, Then Generate
Retrieve relevant documents, inject into context, then generate.
Architecture
The RAG Pipeline
Query → Embed → Vector Search → Top-K chunks → LLM + Context → Answer.
Impact
Enterprise AI
Knowledge bases, chatbots grounded in real data, enterprise search.
In Plain English

Simple: An open-book exam beats a closed-book one, and not because the student got cleverer. The student can now cite the page. Closed-book, a confident student who has forgotten the fact will simply invent one that sounds right — and sound exactly as confident doing it.

Technical: RAG splits knowledge from parameters: a retriever fetches passages at query time and the generator conditions on them, so updating the corpus updates the answers with no retraining. It narrows the hallucination surface rather than closing it — a model can still misread, over-generalise from, or ignore a retrieved passage, and a wrong retrieval is a wrong answer with a citation attached (Lewis et al., 2020).

Knowledge Check

Eight questions on this module. Answer to see why — the explanation appears whether you were right or wrong.

Question 1 of 0
Score 0/0

The Complete Picture

You now know the 10 papers that define modern AI. From attention mechanisms to retrieval-augmented generation — this is the foundation every AI engineer needs.

TransformersGPTBERTLoRAPEFTViTVAEGANsDiffusionRAG

Every paper here is about the shape of a model. Composing many model calls into a working system is a separate layer with its own rules — and no landmark paper of its own yet.

Next: Module 8 — The Future of GenAI. You have seen where these systems came from; the last module asks where they are going, and what that means for you.