10 Papers Every AI Engineer Must Know
The foundational research that shaped modern AI — from transformers to retrieval-augmented generation.
Attention Is All You Need (2017)
Vaswani et al. — the paper that launched a thousand models.
Read the paperVaswani et al., “Attention Is All You Need” (2017) — arXiv:1706.03762 ↗
Simple: At a loud dinner party you do not hear every voice equally — you tune in to the two people whose words actually bear on what you are about to say. Every word in a sentence does that, to every other word, at the same moment.
Technical: Scaled dot-product attention scores every query against every key, softmaxes those scores into weights, and returns the weighted sum of the values. It is an all-pairs operation with no recurrence, which is why the whole sequence resolves in one parallel step instead of one timestep at a time (Vaswani et al., 2017).
GPT — Generative Pre-trained Transformers (2018–2024)
The decoder-only architecture that scaled to change the world.
Read the papersGPT-1: “Improving Language Understanding by Generative Pre-Training” (2018) — OpenAI PDF ↗GPT-3: “Language Models are Few-Shot Learners” (2020) — arXiv:2005.14165 ↗GPT-4 Technical Report (2023) — arXiv:2303.08774 ↗
Simple: Nobody redesigned the engine between 2018 and 2023. They kept the same engine and built it bigger, on more fuel — and past a certain size it started doing tricks nobody had taught it, like following an instruction it had only seen described.
Technical: The GPT line holds the decoder-only, next-token-prediction objective fixed across four generations and scales parameters, data and compute. GPT-3 (2020) is the paper that named the payoff: in-context learning, where few-shot examples in the prompt substitute for gradient updates — a capability the training objective never optimised for.
BERT — Bidirectional Understanding (2018)
BERT’s job is to understand text, not continue it. It was trained by a sophisticated fill-in-the-blank exercise: hide some words, guess them back. The [MASK] sat on the [MASK] — and to answer cat and mat, it has to read the whole sentence, including the part that comes after the gap. That is the difference from GPT, which only ever reads leftwards.
Read the paperDevlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding” (2018) — arXiv:1810.04805 ↗
Simple: A writer composing a sentence can only use the words already on the page. A proofreader reading the finished page can use both sides of any gap — which is why proofreading catches things writing cannot. GPT writes; BERT proofreads.
Technical: BERT is an encoder-only transformer with no causal mask, so every position attends to the full sequence in both directions. That makes autoregressive generation impossible and representation quality excellent, which is why the family survives in classification, ranking and retrieval rather than in chat.
LoRA — Low-Rank Adaptation (2021)
Teaching a big model a new job used to mean rewriting the whole textbook. LoRA is a sticky note clipped to the page: the book is left exactly as it was, and the note carries the correction. For a 7-billion-parameter model the note is around 0.1% of the book’s size — which is what turns fine-tuning from a data-centre job into something you can run on one machine.
Read the paperHu et al., “LoRA: Low-Rank Adaptation of Large Language Models” (2021) — arXiv:2106.09685 ↗
Simple: Sticking the note on is cheap. The reason it works is separate and less obvious: the correction a new job needs is usually simple — a nudge in a few consistent directions, not a rewrite of every page. A short note can hold a simple correction. It could not hold a complicated one.
Technical: LoRA freezes W and learns ΔW = A·B with A of shape d×r and B of shape r×d, so ΔW is constrained to rank at most r. The bet is that the weight update fine-tuning wants has low intrinsic rank. Parameters trained per adapted matrix are 2·d·r, linear in r — and at inference B·A merges back into W, so latency is unchanged (Hu et al., 2021).
PEFT — Parameter Efficient Fine-Tuning
A family of techniques that make large model adaptation practical.
Read the papersAdapters — Houlsby et al., “Parameter-Efficient Transfer Learning for NLP” (2019) — arXiv:1902.00751 ↗Prefix tuning — Li & Liang, “Prefix-Tuning: Optimizing Continuous Prompts for Generation” (2021) — arXiv:2101.00190 ↗
Simple: There is more than one place to clip the note. You can write in the margin next to the text (a small module inserted between layers), or you can staple a covering memo to the front so everything after it is read differently (a learned preamble). LoRA is one of these options, not the whole family.
Technical: PEFT is the umbrella for methods that adapt a frozen backbone by training a small set of new parameters. Adapters insert bottleneck MLP modules inside each block; prefix tuning prepends learned key/value vectors to every attention layer; LoRA reparameterises the weight update itself. They differ in where the trainable parameters sit and in whether they add inference cost — adapters and prefixes do, merged LoRA does not.
ViT — Vision Transformer (2020)
Read a picture like a page. Chop the photo into a grid of small square tiles, line the tiles up in reading order, and hand them to a transformer as if each tile were a word. A 224×224 photo cut into 16×16 tiles becomes a 196-“word” sentence — and from there the machinery is the same machinery that reads text.
Read the paperDosovitskiy et al., “An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale” (2020) — arXiv:2010.11929 ↗
Simple: A child raised on a handful of photos still knows a cat is a cat upside down, or in the corner of the frame — that assumption comes built in. ViT is not born with it. Show it few pictures and it flounders; show it enormously many and it works the rule out itself, then applies it better than the built-in version.
Technical: A CNN hard-codes translation equivariance and locality as architectural priors. ViT removes them: patches are embedded, position is learned, and attention is global from layer one. With less data that missing prior costs it accuracy against a CNN; pre-trained on a very large corpus it overtakes, which is the scaling result the paper reports (Dosovitskiy et al., 2020).
VAE & GANs — Generative Pioneers
Two ways to make a machine produce a face that never existed. The sketcher learns the shape of “face” as a smooth space, so you can slide from one face to another and every point on the way is also a face. The forger never learns that space — it just gets very good at fooling an inspector. Sharper results, but there is no in-between to slide through.
Read the papersVAE — Kingma & Welling, “Auto-Encoding Variational Bayes” (2013) — arXiv:1312.6114 ↗GANs — Goodfellow et al., “Generative Adversarial Networks” (2014) — arXiv:1406.2661 ↗
Simple: The sketcher is graded on how closely the redrawing matches the original, so when it is unsure it hedges — and a hedged line is a soft line. The forger is graded only on whether the inspector is fooled, so hedging earns nothing. Each one’s weakness is the direct cost of what it is being graded on.
Technical: A VAE maximises a reconstruction likelihood plus a KL term pulling the latent posterior towards a prior; the pixel-wise reconstruction loss is what averages away high-frequency detail, and the KL term is what buys the continuous latent space you can interpolate through. A GAN optimises a minimax objective against a discriminator with no reconstruction term at all — hence sharpness, and hence mode collapse and unstable training as the standing failure modes.
Diffusion Models (2020–2022)
Un-blurring. Adding fog to a photograph is free — you need no skill and no training to make a picture worse. So the model is never taught to draw. It is taught only to remove one step of fog. Repeat that a few dozen times, starting from nothing but fog, and a picture comes out. Drag the slider to walk the fog in, then read it right to left to see what the model actually does.
Read the papersHo et al., “Denoising Diffusion Probabilistic Models” (2020) — arXiv:2006.11239 ↗Latent diffusion — Rombach et al., “High-Resolution Image Synthesis with Latent Diffusion Models” (2022) — arXiv:2112.10752 ↗
Simple: Un-blurring in one heroic go is a job for an artist. Un-blurring by a hair, over and over, is a job for a technician — and a technician you can actually train, because you can manufacture unlimited practice material for free just by blurring photographs you already have.
Technical: The forward process is a fixed Gaussian noising schedule with no learned parameters, so training pairs are generated for nothing and any timestep can be sampled directly. The network learns only to predict the noise added at a given step; sampling subtracts that prediction iteratively. Latent diffusion runs the loop in a compressed autoencoder space rather than pixel space, which is the change that brought the cost down to consumer hardware (Ho et al., 2020; Rombach et al., 2022).
RAG — Retrieval-Augmented Generation (2020)
Lewis et al. — grounding LLMs in real, retrievable knowledge.
Read the paperLewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks” (2020) — arXiv:2005.11401 ↗
Simple: An open-book exam beats a closed-book one, and not because the student got cleverer. The student can now cite the page. Closed-book, a confident student who has forgotten the fact will simply invent one that sounds right — and sound exactly as confident doing it.
Technical: RAG splits knowledge from parameters: a retriever fetches passages at query time and the generator conditions on them, so updating the corpus updates the answers with no retraining. It narrows the hallucination surface rather than closing it — a model can still misread, over-generalise from, or ignore a retrieved passage, and a wrong retrieval is a wrong answer with a citation attached (Lewis et al., 2020).
Knowledge Check
Eight questions on this module. Answer to see why — the explanation appears whether you were right or wrong.
The Complete Picture
You now know the 10 papers that define modern AI. From attention mechanisms to retrieval-augmented generation — this is the foundation every AI engineer needs.
Every paper here is about the shape of a model. Composing many model calls into a working system is a separate layer with its own rules — and no landmark paper of its own yet.
Next: Module 8 — The Future of GenAI. You have seen where these systems came from; the last module asks where they are going, and what that means for you.