Evaluating LLMs Without Fooling Yourself

· 2 min read · Syed Omar Faruk Towaha
Evaluating LLMs Without Fooling Yourself

Every AI feature has a moment I call the demo glow. You try three prompts, the answers are great, and you're ready to ship. Then real users arrive with typos, edge cases, sarcasm and questions you never imagined, and the quality you felt turns out to be the quality of your three favourite examples.

Vibes vs measurement
Three examples tell you almost nothing. A few hundred start to tell the truth.

Here's how to evaluate an LLM-powered feature honestly, without needing a research lab.

1. Build a test set before you tune anything

Collect 50 to 200 real or realistic inputs. Include:

Write down what a good answer looks like for each, at least roughly. This takes an afternoon and will save you months.

2. Decide what "good" means

"Is it good?" is not measurable. These are:

Some of these you can check with plain code. Format is easy:

import json

def valid_json(output: str) -> bool:
    try:
        json.loads(output)
        return True
    except json.JSONDecodeError:
        return False

pass_rate = sum(valid_json(run(case.input)) for case in cases) / len(cases)

3. Use a model as a judge, carefully

For fuzzy criteria like helpfulness, you can ask a second model to grade answers against a rubric. It's fast and surprisingly useful, but:

4. Compare, don't just score

Absolute numbers are hard to interpret. Comparisons are easy: prompt A vs prompt B, model X vs model Y, on the same test set. Track the results in a simple table every time you change something.

5. Look at the failures

The score tells you how much is wrong. Reading the failures tells you why. Group them: missing information, wrong retrieval, formatting, hallucination. One category usually dominates, and that's what you fix next.

6. Keep evaluating after launch

Log real inputs (with privacy in mind), sample them regularly and add interesting failures to your test set. Your test set should grow like a museum of everything that ever went wrong.

The uncomfortable truth

Evaluation is boring compared to prompting. It's spreadsheets and reading outputs instead of the thrill of the perfect prompt. It's also the only thing that separates an AI product from an AI demo. The demo glow fades. A test set doesn't.

// related

// prefer the terminal?

Open the terminal blog and type read evaluating-llms-without-fooling-yourself.