Benchmarks are maps. They help us compare approaches, notice progress, and communicate results. But a map is not the street, and a benchmark score does not tell us how a system behaves when a user is impatient, a document is stale, or an API returns half the expected data.

Begin with failure stories

Before choosing a metric, write down the failures that would make the product untrustworthy. The answer may be fabricated citations, silent omissions, inconsistent formatting, or a confident response when the evidence is weak.

A two by two evaluation grid comparing answer quality and evidence quality
Answer quality and evidence quality should be measured separately.

Those stories become evaluation cases. Each case should have an input, an expected behavior, a reason that behavior matters, and enough metadata to group similar failures later.

Use a portfolio of signals

SignalWhat it tells us
Task successWhether the user reached the intended outcome
GroundednessWhether claims are supported by available evidence
AbstentionWhether the system knows when not to answer
LatencyWhether the experience remains usable

Keep evaluation close to development

An evaluation suite should be easy enough to run while making a change, not only before a release. The most useful loop is short: observe a failure, preserve it as a test, change the system, and compare the result.

The goal is not a perfect score. It is a clear picture of what the system can do today, where it is brittle, and whether the next change genuinely made it better.