The Future of GenAI
Current limitations, the evolving landscape, and your next steps in the AI revolution.
What AI Cannot Do (Yet)
Understanding limitations is just as important as understanding capabilities.
Simple: The trouble is not that it fails. Everything fails sometimes. The trouble is that it fails in exactly the same fluent, confident voice it uses when it is right — so the failure does not announce itself the way a person hesitating would.
Technical: These four are limitations of the current paradigm, not permanent facts about machines, and the honest statement is that we cannot say which will move next. What makes them costly today is the missing calibration signal: output fluency is uncorrelated with correctness, so the usual human heuristic for detecting uncertainty does not apply. Design for that, rather than for the failure rate.
Using Limitations to Your Advantage
Every limitation is a design constraint you can work with, not against.
Simple: You would not hand a new colleague a six-month project on day one and check back at the end. You would give them something you can look at on Friday. None of these four strategies makes the model better — they make its mistakes cheap and early instead of expensive and late.
Technical: All four shorten the distance between an error and its detection. Tight scoping caps how far a wrong step can propagate; verification makes correctness machine-checkable rather than eyeballed; grounding removes the class of errors caused by guessing at retrievable facts; and human review is the fallback for the residue none of the other three can catch. The strategies compose, and the weakest of them sets your actual error floor.
Bias, Fairness & Harm
A model learns the world from its training data — including the world’s skews. Bias isn’t a bug someone forgot to fix; it’s the default, and keeping it out takes deliberate work.
Simple: A mirror is not being unkind when it shows you a crooked collar. Train a system on records of how decisions were made in the past and it will reproduce how they were made — including the parts nobody would defend out loud. Nothing has to go wrong for that to happen; it is the system working.
Technical: Bias enters at every stage — corpus composition, annotation, objective choice, and the feedback loop once deployed — so there is no single place to fix it. The harder point is that “fair” is not one target: several formal fairness criteria are provably incompatible except in degenerate cases, so you are choosing which definition to satisfy, not whether to be fair. That choice is a stated, reviewable decision or it is an unstated one.
Why Benchmarks Mislead
“State of the art on MMLU” sounds decisive. It usually isn’t. A benchmark measures performance on one frozen test — and there are four good reasons that number won’t match what you see in production.
Simple: A driving test tells you someone can pass a driving test. It is genuine evidence — and it is not the same claim as “safe on an unfamiliar motorway in the rain”. The score is real; what people do with it is to quietly upgrade it into a promise it never made.
Technical: A benchmark is a measurement on one frozen, public distribution, and every property that makes it usable also makes it leaky: being public invites contamination, being fixed invites optimisation against it, and being clean removes the ambiguity that dominates production inputs. None of that makes the number meaningless — it makes it a measurement whose scope you have to state. The transferable habit is to ask what was measured, on what, when, and by whom, and to treat an unlabelled number as unmeasured.
The Shape of AI Governance
Laws will change; the shape of good governance is steadier. Wherever you look, four pillars keep reappearing — and they map cleanly onto how you should run your own AI features.
Simple: Building regulations do not tell you how to design a house. They tell you the stairs need a handrail, somebody has to sign off the wiring, and if the roof falls in there is a name on the certificate. Good AI governance is the same trick applied to systems that make decisions about people.
Technical: The four pillars are outcome-based rather than technique-based, which is why they have outlasted several waves of specific regulation: they constrain what a system must be able to demonstrate, not how it must be built. That is also what makes them portable into engineering practice — risk tiering is where you spend review effort, transparency is what you log, oversight is your override path, and accountability is a named owner. The statutes will keep changing; these four keep reappearing.
Key Trends Shaping the Future
Six directions the field is moving in, drawn from where research effort and product releases have actually been concentrating. Each card is a headline — open one for what it means in practice, and treat the specific capability figures inside as a snapshot rather than a fixed number.
Simple: Six directions, and the one thing they share is that none of them is a finish line. A trend is a description of where effort has been going, which is a much weaker claim than a prediction — effort has gone in directions that turned out to be dead ends before.
Technical: Read these as areas of concentrated research and product investment, not as a roadmap. Two of them cut against each other in a way worth noticing: longer context and lower cost per task pull in opposite directions, since attention cost grows superlinearly in sequence length and the mitigations are approximations with their own trade-offs. Any specific capability number attached to these — window sizes, prices, latencies — is a snapshot, and this module has already told you what to do with an undated figure.
The Shape of the Work
One agent, one long chain, one context window. When that gets slow and expensive the instinct is to rewrite the prompt — but usually the prompt is fine and the shape is wrong, because sequence is not dependency. Ask it of every arrow: does the next step actually read the previous step’s output? Where the answer is no, the wait is imaginary and deleting it is free. Four shapes cover almost everything graph engineering
Simple: Four people cooking one dinner do not queue at the same chopping board. They split what is independent, meet where the dish has to come together, and someone tastes it before it leaves the kitchen — and the tasting is a separate job from the cooking, because whoever made it is the worst judge of it.
Technical: Two of these are earlier strategies from this module moved one layer out. “Always Verify” becomes a node whose only job is to stop weak work moving downstream; “Scope Tasks Tightly” becomes one bounded job per node returning a structured value, which matters more the longer the run, since prose gets summarised at every hop and every summary is lossy. On the name: the practice is documented — Anthropic has published on a lead agent coordinating parallel subagents with a separate citation stage, reporting better breadth-first results at materially higher token cost — but the label is a practitioner’s coinage rather than settled vocabulary. Learn the shapes; hold the name loosely.
Impact on Software Development
The honest version of this slide has no numbers on it. Productivity claims for AI-assisted development are quoted constantly and measured badly — the figures depend on the task, the language, the codebase, the tooling and the month, and the module you just read tells you what to do with a number that arrives without those. What can be described without a figure is the shape of the change, and that has been consistent.
Simple: When power tools arrived, carpentry did not stop being a craft — the sawing stopped being the hard part, and knowing what to build became the whole job. The four shifts on this slide are versions of that one sentence.
Technical: Notice what this slide deliberately does not have on it: a figure. Productivity claims for AI-assisted development vary with task, language, codebase, tooling and measurement method, and are frequently self-reported — so a single multiplier is not a quantity you can act on. What can be stated without a number is the redistribution of effort: generation becomes cheap, so review, verification and decomposition become the bottleneck and therefore the skill.
The Choices Ahead
Beyond “which model is smartest” sit three choices that shape cost, privacy, and control. None has a single right answer — they’re trade-offs you pick per use case.
Simple: Owning a car and calling a taxi are both correct answers — to different questions about how often you drive and who you want holding the keys. Anyone who tells you one of these three choices has a universal answer is selling one side of it.
Technical: Each is a trade against a different axis. Open weights buy inspectability, offline operation and no vendor dependency, at the cost of you owning the serving, safety and update work. On-device buys data locality and predictable latency, bounded by the memory and power budget of the device. And a small tuned model beating a large general one on a narrow task is a claim you can test cheaply for your own task — which makes it the one of the three worth measuring rather than debating.
Knowledge Check
Eight questions on this module. Answer to see why — the explanation appears whether you were right or wrong.
What You’ve Learned
Seven modules, one thread: where these systems came from, what they are, how to talk to them, how to let them act, how to feed them the right context, how to run them properly, and which papers made each step possible. Here it is in one place.
Your Next Steps
You’ve Finished GenAI 101
You started at “what is a token” and you finish knowing why a benchmark score is not a promise. That is the whole arc: not a list of tools, but enough of the machinery to judge a claim, choose an approach, and tell a good result from a plausible one.