Engineering
·
3 min read
Evaluation sets are product decisions, not test fixtures
Which questions go into the eval set decides what the system is allowed to be bad at. That choice belongs with the people who own the workflow.

By
Elena Kowalczyk
,
Principal Applied AI Lead
3 min read
·

Engineers tend to treat the evaluation set the way they treat unit tests: something we write so the build goes green. For a retrieval or agent system, that framing hides the most important decision in the project.
What is in the set is what you promised.
If the set has no questions about policy exceptions, the system can fail every exception and still report 94%. The people who own the workflow know which questions matter, and they are rarely in the room when the set is written.
We write it with them.
Two working sessions with the operators: one to collect real questions from the last quarter, one to rank them by the cost of a wrong answer. The ranking becomes the weighting. A wrong answer about a payment deadline counts for more than a wrong answer about office hours.
It changes after launch.
Every escalation from production is a candidate for the set. We review the candidates monthly with the system owner, and the set grows with the system instead of fossilising at the version we shipped.

Written by
Elena Kowalczyk
Principal Applied AI Lead
View profile →
02 · Keep reading
More from the team.

03 · Start
Have a use case
like this one?
Most of what’s in this article came out of a real engagement. Tell us about yours.






