GENAI 102
M08 · L04
When Code Can’t Score It

LLM-as-Judge & Human-in-the-Loop

Some outputs are right or wrong — code can check those. But when correctness is a judgment call (Is this summary faithful? Is this reply helpful?), code has nothing to compare against. So you reach for a different judge: an LLM, a human, or both.

01 / 08
GENAI 102
M08 · L04
The Automated Judge

Let a Model Grade the Output

Hand a model the output plus a rubric, and ask it to score and critique. It reads nuance the way a person would — at the speed and cost of software.

  • Give it the output — the answer you want assessed
  • Give it a rubric — what “good” means, in plain words
  • Get a score and a reason — nuance at scale, cheaply
  • But the judge can be wrong too — it is a model, not an oracle
02 / 08
GENAI 102
M08 · L04
Trusting the Judge

Making the Judge Reliable

A judge you can’t trust is worse than none. Three things earn that trust: a clear rubric, worked examples, and checking the judge against people.

How to trust its scores
Write a clear rubric · show examples of good and bad answers · then spot-check the judge against human labels, so you know its scores track reality
03 / 08
GENAI 102
M08 · L04
The Gold Standard

Keep a Human in the Loop

For quality and for high-stakes calls, people are still the gold standard. You don’t need them on everything — you need them where it counts.

  • Label a sample — hand-rate a slice to know the real quality
  • Resolve disputes — be the tiebreaker when the automated judge is unsure
  • Calibrate the automation — those labels are what you check the LLM-judge against
04 / 08
GENAI 102
M08 · L04
Picking Your Evaluator

Cheapest Eval That’s Trustworthy Enough

These three aren’t rivals — they’re a ladder. Use the cheapest one that’s reliable for the job.

  • Code — wherever the answer is checkable, it’s free and exact
  • LLM-judge — for nuance at scale, once you’ve verified it
  • Humans — for the stakes that truly demand them
05 / 08
GENAI 102
Build It
From Concept to Capstone

Build It

How to implement: pick one subjective quality in your tool, write a one-paragraph rubric for it, and have an LLM score a few of your outputs against it. Then read its scores yourself — do you agree? That check is the whole point.

  • Weekly AI Tasks tracker — an LLM-judge can rate “did the summary capture the week accurately?” while you spot-check it against a few weeks you know well.
  • Personal brand site — an LLM-judge rates tone and clarity; you, the human, make the final call on what’s true and worth including.
06 / 08
GENAI 102
Knowledge Check

Check what stuck

Three questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

07 / 08
GENAI 102
Summary
Recap

What you learned

When correctness is a judgment call, code can’t score it. An LLM-judge gives you nuance at scale — once you’ve made it reliable with a rubric, examples, and spot-checks against people. Humans stay the gold standard for quality and high stakes. Use the cheapest eval that’s trustworthy enough.

08 / 08