The Future of GenAI

Current limitations, the evolving landscape, and your next steps in the AI revolution.

What AI Cannot Do (Yet)

Understanding limitations is just as important as understanding capabilities.

🧮
Deep Reasoning Gaps
Multi-step math and formal logic still trip up even the best models.
🌍
No World Model
LLMs don’t truly “understand” physics, time, or spatial relationships.
💡
Genuine Novelty
AI remixes patterns from training data. Truly novel breakthroughs remain human.
🏔️
Very Long Horizon
Tasks spanning days or weeks with hundreds of decisions still need human oversight.
In Plain English

Simple: The trouble is not that it fails. Everything fails sometimes. The trouble is that it fails in exactly the same fluent, confident voice it uses when it is right — so the failure does not announce itself the way a person hesitating would.

Technical: These four are limitations of the current paradigm, not permanent facts about machines, and the honest statement is that we cannot say which will move next. What makes them costly today is the missing calibration signal: output fluency is uncorrelated with correctness, so the usual human heuristic for detecting uncertainty does not apply. Design for that, rather than for the failure rate.

Using Limitations to Your Advantage

Every limitation is a design constraint you can work with, not against.

Strategy 1
Always Verify
Build verification into every workflow. Tests, linters, and CI/CD catch AI mistakes automatically — and keep the thing that checks the work separate from the thing that produced it.
Strategy 2
Scope Tasks Tightly
Break large tasks into small, verifiable chunks. Each chunk gets its own agentic loop — one job, one clear output, one obvious way to fail.
Strategy 3
Ground in Sources
Use RAG and tool access to give the model real data. Never let it guess when it can look up.
Strategy 4
Human-AI Partnership
Humans set direction, make judgment calls, and verify. AI handles execution at scale.
In Plain English

Simple: You would not hand a new colleague a six-month project on day one and check back at the end. You would give them something you can look at on Friday. None of these four strategies makes the model better — they make its mistakes cheap and early instead of expensive and late.

Technical: All four shorten the distance between an error and its detection. Tight scoping caps how far a wrong step can propagate; verification makes correctness machine-checkable rather than eyeballed; grounding removes the class of errors caused by guessing at retrievable facts; and human review is the fallback for the residue none of the other three can catch. The strategies compose, and the weakest of them sets your actual error floor.

Bias, Fairness & Harm

A model learns the world from its training data — including the world’s skews. Bias isn’t a bug someone forgot to fix; it’s the default, and keeping it out takes deliberate work.

🩻
Where Bias Enters
Training data, human labelling, and feedback loops each add their own slant.
⚖️
Fairness Conflicts
There are several definitions of “fair” — and you usually can’t satisfy them all at once.
🎭
Two Kinds of Harm
Representational harm shapes how groups are portrayed; allocative harm decides who gets what.
🛠️
What Actually Helps
Concrete mitigations: balanced data, measurement, and a human in the loop.
In Plain English

Simple: A mirror is not being unkind when it shows you a crooked collar. Train a system on records of how decisions were made in the past and it will reproduce how they were made — including the parts nobody would defend out loud. Nothing has to go wrong for that to happen; it is the system working.

Technical: Bias enters at every stage — corpus composition, annotation, objective choice, and the feedback loop once deployed — so there is no single place to fix it. The harder point is that “fair” is not one target: several formal fairness criteria are provably incompatible except in degenerate cases, so you are choosing which definition to satisfy, not whether to be fair. That choice is a stated, reviewable decision or it is an unstated one.

Why Benchmarks Mislead

“State of the art on MMLU” sounds decisive. It usually isn’t. A benchmark measures performance on one frozen test — and there are four good reasons that number won’t match what you see in production.

🧫
Contamination
If test questions leaked into training data, the model is recalling — not reasoning.
📈
Teaching to the Test
Optimising for a benchmark lifts the score without lifting the skill behind it.
🌉
The Benchmark Gap
Clean multiple-choice questions look nothing like your messy, real-world inputs.
🏆
Reading a “SOTA” Claim
What to check before “SOTA on X” means anything for your use.
In Plain English

Simple: A driving test tells you someone can pass a driving test. It is genuine evidence — and it is not the same claim as “safe on an unfamiliar motorway in the rain”. The score is real; what people do with it is to quietly upgrade it into a promise it never made.

Technical: A benchmark is a measurement on one frozen, public distribution, and every property that makes it usable also makes it leaky: being public invites contamination, being fixed invites optimisation against it, and being clean removes the ambiguity that dominates production inputs. None of that makes the number meaningless — it makes it a measurement whose scope you have to state. The transferable habit is to ask what was measured, on what, when, and by whom, and to treat an unlabelled number as unmeasured.

The Shape of AI Governance

Laws will change; the shape of good governance is steadier. Wherever you look, four pillars keep reappearing — and they map cleanly onto how you should run your own AI features.

Pillar 1
Risk Tiers
Match the rules to the stakes. A spam filter and a loan-approval model don’t deserve the same scrutiny.
Pillar 2
Transparency & Disclosure
People should know when they’re dealing with AI — and what it was trained on.
Pillar 3
Human Oversight
A person can review, override, and be appealed to on high-stakes decisions.
Pillar 4
Accountability
When the system causes harm, a named someone is answerable — not “the algorithm”.
Worked example — the EU AI Act: it sorts systems into tiers by risk. A few uses are simply banned; “high-risk” ones (hiring, credit, medical) carry hard obligations; most everyday tools face only light disclosure rules. Same idea you already use in engineering: spend your safety budget where the blast radius is largest.
In Plain English

Simple: Building regulations do not tell you how to design a house. They tell you the stairs need a handrail, somebody has to sign off the wiring, and if the roof falls in there is a name on the certificate. Good AI governance is the same trick applied to systems that make decisions about people.

Technical: The four pillars are outcome-based rather than technique-based, which is why they have outlasted several waves of specific regulation: they constrain what a system must be able to demonstrate, not how it must be built. That is also what makes them portable into engineering practice — risk tiering is where you spend review effort, transparency is what you log, oversight is your override path, and accountability is a named owner. The statutes will keep changing; these four keep reappearing.

Key Trends Shaping the Future

Six directions the field is moving in, drawn from where research effort and product releases have actually been concentrating. Each card is a headline — open one for what it means in practice, and treat the specific capability figures inside as a snapshot rather than a fixed number.

👁️
Multimodal
Text, images, audio, video — all in one model
🤖
Agentic
AI that acts autonomously with tools
🧠
Reasoning
Extended thinking and chain-of-thought
⚡
Speed & Cost
Models getting faster and cheaper exponentially
📐
Longer Context
From 4K to 200K to 1M+ tokens
🌐
Open Ecosystem
MCP, open standards, interoperability
In Plain English

Simple: Six directions, and the one thing they share is that none of them is a finish line. A trend is a description of where effort has been going, which is a much weaker claim than a prediction — effort has gone in directions that turned out to be dead ends before.

Technical: Read these as areas of concentrated research and product investment, not as a roadmap. Two of them cut against each other in a way worth noticing: longer context and lower cost per task pull in opposite directions, since attention cost grows superlinearly in sequence length and the mitigations are approximations with their own trade-offs. Any specific capability number attached to these — window sizes, prices, latencies — is a snapshot, and this module has already told you what to do with an undated figure.

The Shape of the Work

One agent, one long chain, one context window. When that gets slow and expensive the instinct is to rewrite the prompt — but usually the prompt is fine and the shape is wrong, because sequence is not dependency. Ask it of every arrow: does the next step actually read the previous step’s output? Where the answer is no, the wait is imaginary and deleting it is free. Four shapes cover almost everything graph engineering

Chain
Every wait is a real wait
Fan out, join once
Wait only for the full set
Cheap path, deep path
Spend what it deserves
Checked loop
Repeat only on evidence
Neither cheap nor durable by default. A fleet per request burns far more than the chain it replaced — and any run long enough to be interrupted has to know where it can resume, or a crash at minute forty costs you forty minutes.
In Plain English

Simple: Four people cooking one dinner do not queue at the same chopping board. They split what is independent, meet where the dish has to come together, and someone tastes it before it leaves the kitchen — and the tasting is a separate job from the cooking, because whoever made it is the worst judge of it.

Technical: Two of these are earlier strategies from this module moved one layer out. “Always Verify” becomes a node whose only job is to stop weak work moving downstream; “Scope Tasks Tightly” becomes one bounded job per node returning a structured value, which matters more the longer the run, since prose gets summarised at every hop and every summary is lossy. On the name: the practice is documented — Anthropic has published on a lead agent coordinating parallel subagents with a separate citation stage, reporting better breadth-first results at materially higher token cost — but the label is a practitioner’s coinage rather than settled vocabulary. Learn the shapes; hold the name loosely.

Impact on Software Development

The honest version of this slide has no numbers on it. Productivity claims for AI-assisted development are quoted constantly and measured badly — the figures depend on the task, the language, the codebase, the tooling and the month, and the module you just read tells you what to do with a number that arrives without those. What can be described without a figure is the shape of the change, and that has been consistent.

The shift: AI doesn’t replace developers — it changes what “development” means. Less typing, more thinking. Less debugging, more designing. Less repetition, more creation.
✍️
The typing stops being the job
Boilerplate, scaffolding and the fourth near-identical CRUD endpoint are the work AI absorbs first. What is left is deciding what should exist — which was always the harder half.
🔍
Reviewing becomes a core skill
You read more code than you write, and much of it you did not write. Judging whether an answer is right — quickly, and without running it — is now a daily skill rather than an occasional one.
✅
Verification moves to the front
Tests, types, linters and CI stop being hygiene and start being the mechanism that makes delegation safe. This is Strategy 1 from earlier in this module, applied to your whole workflow — and the reason a checking step should be a separate step, not a second opinion from the same author.
🧱
Scoping is the new craft
Work arrives as tasks you hand off, so the skill is cutting a problem into pieces small enough to be checked — and noticing which of them never needed to wait for each other. Vague in, vague out — and a badly scoped task fails expensively.
In Plain English

Simple: When power tools arrived, carpentry did not stop being a craft — the sawing stopped being the hard part, and knowing what to build became the whole job. The four shifts on this slide are versions of that one sentence.

Technical: Notice what this slide deliberately does not have on it: a figure. Productivity claims for AI-assisted development vary with task, language, codebase, tooling and measurement method, and are frequently self-reported — so a single multiplier is not a quantity you can act on. What can be stated without a number is the redistribution of effort: generation becomes cheap, so review, verification and decomposition become the bottleneck and therefore the skill.

The Choices Ahead

Beyond “which model is smartest” sit three choices that shape cost, privacy, and control. None has a single right answer — they’re trade-offs you pick per use case.

🔓
Open vs Closed
Downloadable open weights vs. API-only frontier models — control vs. capability.
📱
On-Device vs Cloud
Local models keep data on your phone; the cloud brings raw power. Privacy vs. capability.
🐜
Smaller, Sharper Models
A tuned small model often beats a giant one on a narrow task — at a fraction of the cost.
🔭
The Research Frontier
World models, planning, and reasoning — where the next leaps are being chased.
In Plain English

Simple: Owning a car and calling a taxi are both correct answers — to different questions about how often you drive and who you want holding the keys. Anyone who tells you one of these three choices has a universal answer is selling one side of it.

Technical: Each is a trade against a different axis. Open weights buy inspectability, offline operation and no vendor dependency, at the cost of you owning the serving, safety and update work. On-device buys data locality and predictable latency, bounded by the memory and power budget of the device. And a small tuned model beating a large general one on a narrow task is a claim you can test cheaply for your own task — which makes it the one of the three worth measuring rather than debating.

Knowledge Check

Eight questions on this module. Answer to see why — the explanation appears whether you were right or wrong.

Question 1 of 0
Score 0/0

What You’ve Learned

Seven modules, one thread: where these systems came from, what they are, how to talk to them, how to let them act, how to feed them the right context, how to run them properly, and which papers made each step possible. Here it is in one place.

🔮
The AI Revolution
Neural networks, the perceptron, deep learning and attention — and the seven years from GPT-1 to now. How the field arrived here.
🧠
How LLMs Work
Tokens, embeddings, attention, temperature — the engine under the hood.
✍️
Prompt Engineering
System prompts, few-shot, XML tags, chain of thought — the art of communicating with AI.
🤖
Agentic AI
The loop (Gather/Act/Verify/Repeat), tools, MCP, skills, subagents — beyond chatting.
🔧
Context & Retrieval
The context window as a budget, model selection, RAG, chunking and agentic search — feeding it the right things.
🔧
Power User
Skills, MCP servers, subagents, hooks, permissions and orchestration — the full toolkit.
📄
The Landmark Papers
Attention, GPT and BERT, LoRA and PEFT, ViT, VAEs and GANs, diffusion, RAG — where all of it came from.

Your Next Steps

This week
Practice Daily
Use Claude Code for one real task every day. Start small, build confidence.
This month
Build Your CLAUDE.md
Create a CLAUDE.md for every project you work on. Document your stack, rules, and patterns.
This quarter
Build an MCP Server
Wrap one of your team’s internal tools as an MCP server. See the power firsthand.
Ongoing
Teach Others
The best way to master AI is to teach it. Share what you’ve learned with your team.

You’ve Finished GenAI 101

You started at “what is a token” and you finish knowing why a benchmark score is not a promise. That is the whole arc: not a list of tools, but enough of the machinery to judge a claim, choose an approach, and tell a good result from a plausible one.

LLMsPromptingAgentic LoopToolsMCPSkillsRAGSubagentsHooksOrchestrationLandmark Papers
🎯
What you can do now
Write a prompt that specifies the job rather than hinting at it. Set an agent loose on a task and build the verification that makes that safe. Decide whether a problem wants retrieval, fine-tuning or neither. Read a paper’s abstract and know roughly where it sits in the lineage of the ten in Module 7.
🛤️
Where to go next
The previous slide is the plan, and it is deliberately small: one real task a day this week, a CLAUDE.md for every project this month, one internal tool wrapped as an MCP server this quarter. Then teach it — explaining the agentic loop to a colleague is the fastest way to find the part you only half understand.
One thing worth keeping. Almost every specific in this course has a shelf life — model names, context sizes, prices, which technique is fashionable. The parts that will still be true are the ones you can reason from: the loop, the context budget, the difference between retrieving knowledge and changing behaviour, and the habit of asking a number where it came from. Re-check the specifics; keep the reasoning.