Engineering

·

3 min read

Evaluation sets are product decisions, not test fixtures

Which questions go into the eval set decides what the system is allowed to be bad at. That choice belongs with the people who own the workflow.

Elena Kowalczyk

By

Elena Kowalczyk

,

Principal Applied AI Lead

3 min read

·

In this article

Filed under

Engineering

Published

Reading time

3 min read

Written by

Elena Kowalczyk

Engineers tend to treat the evaluation set the way they treat unit tests: something we write so the build goes green. For a retrieval or agent system, that framing hides the most important decision in the project.

What is in the set is what you promised.

If the set has no questions about policy exceptions, the system can fail every exception and still report 94%. The people who own the workflow know which questions matter, and they are rarely in the room when the set is written.

We write it with them.

Two working sessions with the operators: one to collect real questions from the last quarter, one to rank them by the cost of a wrong answer. The ranking becomes the weighting. A wrong answer about a payment deadline counts for more than a wrong answer about office hours.

It changes after launch.

Every escalation from production is a candidate for the set. We review the candidates monthly with the system owner, and the set grows with the system instead of fossilising at the version we shipped.

Elena Kowalczyk

Written by

Elena Kowalczyk

Principal Applied AI Lead

View profile →

03 · Start

Have a use case
like this one?

Most of what’s in this article came out of a real engagement. Tell us about yours.

Create a free website with Framer, the website builder loved by startups, designers and agencies.