Evaluating LLMs Without Fooling Yourself

Every AI feature has a moment I call the demo glow. You try three prompts, the answers are great, and you're ready to ship. Then real users arrive with typos, edge cases, sarcasm and questions you never imagined, and the quality you felt turns out to be the quality of your three favourite examples.

Here's how to evaluate an LLM-powered feature honestly, without needing a research lab.
1. Build a test set before you tune anything
Collect 50 to 200 real or realistic inputs. Include:
- the common cases (what most users will ask)
- the hard cases (long, ambiguous, multi-part)
- the edge cases (empty input, wrong language, rude messages)
- the "should refuse or say I don't know" cases
Write down what a good answer looks like for each, at least roughly. This takes an afternoon and will save you months.
2. Decide what "good" means
"Is it good?" is not measurable. These are:
- Correct: does it contain the right facts or perform the right action?
- Grounded: is every claim supported by the provided documents?
- Format: is it valid JSON, under 100 words, in the right language?
- Safe: does it avoid leaking data or giving harmful advice?
Some of these you can check with plain code. Format is easy:
import json
def valid_json(output: str) -> bool:
try:
json.loads(output)
return True
except json.JSONDecodeError:
return False
pass_rate = sum(valid_json(run(case.input)) for case in cases) / len(cases)
3. Use a model as a judge, carefully
For fuzzy criteria like helpfulness, you can ask a second model to grade answers against a rubric. It's fast and surprisingly useful, but:
- give it a specific rubric, not "rate this 1-10"
- check its grades against your own on a sample
- watch for bias toward longer answers and toward its own writing style
4. Compare, don't just score
Absolute numbers are hard to interpret. Comparisons are easy: prompt A vs prompt B, model X vs model Y, on the same test set. Track the results in a simple table every time you change something.
5. Look at the failures
The score tells you how much is wrong. Reading the failures tells you why. Group them: missing information, wrong retrieval, formatting, hallucination. One category usually dominates, and that's what you fix next.
6. Keep evaluating after launch
Log real inputs (with privacy in mind), sample them regularly and add interesting failures to your test set. Your test set should grow like a museum of everything that ever went wrong.
The uncomfortable truth
Evaluation is boring compared to prompting. It's spreadsheets and reading outputs instead of the thrill of the perfect prompt. It's also the only thing that separates an AI product from an AI demo. The demo glow fades. A test set doesn't.