Engineering
·
5 min read
What we look for in a RAG eval set
A 94% accuracy score means nothing without knowing what's in the question set and who wrote the answers. Here's the bar we hold retrieval systems to before launch.

By
Elena Kowalczyk
,
Principal Applied AI Lead
5 min read
·

Every RAG system we’ve evaluated passes its own demo. That’s a low bar — the questions in a demo are usually the ones the person building it already knows the system handles well. The real test is a gold-standard evaluation set, and most of the ones we see when we inherit an existing system are too small, too easy, or too disconnected from how people actually ask questions.
Here’s what we require before we call a retrieval system production-ready.
Real questions, not written-for-the-test questions. We pull from actual support tickets, actual claims notes, actual research requests — whatever your team already fields. A question set written by the engineering team tends to be phrased the way engineers phrase things, which is not how the eventual users phrase things. On the Cobalt Mutual engagement, we built our evaluation set from 340 historical claims questions pulled directly from adjuster notes, not from a workshop.
Enough volume to catch the tail, not just the average. A 20-question eval set will tell you the system handles common cases. It won’t tell you anything about the unusual policy wording, the contradictory internal bulletin, or the question that spans two source documents. We aim for at least a few hundred questions before we trust an aggregate accuracy number, and we always break that number down by question category rather than reporting one blended score.
A verified correct answer for every question, with its source. Not “an answer that sounds right” — an answer a subject-matter expert has signed off on, tied to the exact passage it should cite. This is the slow part of building an eval set, and it’s also the part that can’t be skipped, because it’s the only thing that lets you tell the difference between “the system is grounded and correct” and “the system is grounded and citing the wrong thing confidently.”
A defined threshold for “I don’t know.” A retrieval system that always produces an answer, even when its retrieval confidence is low, will look more impressive in a demo and be less trustworthy in production. We set an explicit confidence floor below which the system says it can’t find a grounded answer, and we evaluate that refusal behavior as part of the accuracy number, not as a separate footnote.
Re-running the eval set after every meaningful change. Not just at launch. A new document source, an updated chunking strategy, a policy revision — any of these can move accuracy without anyone noticing until a user complains. We treat the gold question set the way a good engineering team treats a regression test suite: something that runs continuously, not something built once for a launch review.
The number we publish on our own case studies — a 94% accuracy bar against a held-out evaluation set — only means something because of the process above. A 94% score against 20 easy questions and a 94% score against 340 real ones pulled from actual claims files are not the same claim, even though they look identical on a slide. If a vendor shows you an accuracy number, the first question worth asking isn’t “how high is it” — it’s “what’s the question set, and who wrote the answers.”
We’d rather ship a system with a lower number we trust than a higher one we don’t.

Written by
Elena Kowalczyk
Principal Applied AI Lead
View profile →
02 · Keep reading
More
from the team.

03 · Start
Have a use case
like this one?
Most of what’s in this article came out of a real engagement. Tell us about yours.





