Benchmarks are maps. They help us compare approaches, notice progress, and communicate results. But a map is not the street, and a benchmark score does not tell us how a system behaves when a user is impatient, a document is stale, or an API returns half the expected data.
Begin with failure stories
Before choosing a metric, write down the failures that would make the product untrustworthy. The answer may be fabricated citations, silent omissions, inconsistent formatting, or a confident response when the evidence is weak.
Those stories become evaluation cases. Each case should have an input, an expected behavior, a reason that behavior matters, and enough metadata to group similar failures later.
Use a portfolio of signals
| Signal | What it tells us |
|---|---|
| Task success | Whether the user reached the intended outcome |
| Groundedness | Whether claims are supported by available evidence |
| Abstention | Whether the system knows when not to answer |
| Latency | Whether the experience remains usable |
Keep evaluation close to development
An evaluation suite should be easy enough to run while making a change, not only before a release. The most useful loop is short: observe a failure, preserve it as a test, change the system, and compare the result.
The goal is not a perfect score. It is a clear picture of what the system can do today, where it is brittle, and whether the next change genuinely made it better.