Code-Based Evals
The cheapest, most reliable eval is plain code. A deterministic check runs in milliseconds, costs nothing, and never disagrees with itself — the same input always gives the same verdict.
Facts, Not Opinions
If a rule can be written down, code can enforce it — instantly and without judgment.
- Exact match — does the output equal the expected answer?
- Valid format — is it valid JSON?
- Required fields present — nothing missing
- A number in range — is the score between 0 and 100?
- A rule satisfied — no forbidden words, length under a limit
Whenever “Correct” Is Checkable
Use a code-based eval any time correctness is objective. They are fast, free, and repeatable, so there is no reason to skip them.
The Limit
Code checks facts, not taste. Some questions have no rule to test against:
- Is this summary good?
- Is this tone right?
- Is the answer helpful and clear?
That is where LLM-as-judge and humans come in — the next lesson.
Build a Small Labeled Set
A handful of input → expected pairs is enough to start. It turns “seems fine” into a number you can track over time — so you know whether a change helped or hurt.
Build It
How to implement: write one deterministic check for your tool — for example, the output is valid JSON with the required fields — and run it on your examples.
- Weekly AI Tasks tracker — code-check that a parsed task has a valid date and a known status; reject and re-ask otherwise.
- Personal brand site — code-check that every project entry has a source link and no empty fields.
What you learned
Code-based evals are fast, free, and deterministic — use them wherever correctness is objectively checkable, and run them on every change. They can't judge nuance, so a small labeled set of input→expected pairs turns quality into a trackable number.