Your RAG Eval Set Was Written by the Team That Built It. Your Users Ask Different Questions.


Most enterprise RAG pilots have an evaluation set. Usually it is a spreadsheet of fifty or a hundred questions, each with an expected answer and a source document. The pilot runs against it, the scores look strong, and the steering committee signs off on a wider rollout.
Then real staff start using it. Within a few weeks the support channel fills with "it couldn't find the policy" and "it gave me the old rate". The eval dashboard is still green. Nobody has changed the model. The retrieval is quietly failing on questions the eval set never asked.
The cause is rarely the vector database or the embedding model. It is who wrote the questions. In most pilots, the eval set was written by the same people who built the system, and they ask questions very differently from the people who will use it.
Builders know the corpus. They have read the documents, chunked them, debugged them and watched the retriever pull them back. When they sit down to write test questions, they do it with the documents open in another tab.
That shapes the questions in ways that are hard to notice from the inside:
None of this is carelessness. It is the curse of knowledge. Once you know where the answer lives, it is very hard to phrase a question the way a confused, busy person would.
Picture an HR policy assistant at a mid-sized Australian insurer. The platform team loads the leave, travel, expenses and code of conduct policies, then writes eighty test questions. A typical one: "What is the parental leave entitlement for primary carers?" The assistant retrieves the right section and answers correctly almost every time.
In production, a team leader types: "new dad in my team, how many wks paid + can he split it". That one message has an abbreviation, a casual synonym, two questions joined together, and an assumption about splitting leave that sits in a separate flexible work procedure. The retriever pulls the general leave policy, the answer covers half the question, and the citation looks fine. The team leader gives up and emails HR, which is exactly what the assistant was meant to reduce.
The eval set had no way of catching that, because nobody on the build team would ever phrase a question like that.
If you pull a week of real, de-identified queries from almost any internal assistant, the same patterns show up. They look nothing like the golden set.
Each of these hurts retrieval in a different way. Together they describe the gap between the questions you tested and the questions you get.
An eval score is only a measure of how well the system handles the questions in the eval set. If those questions come from a different distribution to production traffic, a high score tells you very little about production.
A few things make this worse:
We have written before about why RAG demos flatter retrieval and why offline evals can look green while people rewrite every answer. This is the upstream cause of both: the questions themselves were never representative.
The single most useful change is to stop inventing questions and start sampling them. Take a regular random sample of real queries from the logs, plus a targeted sample of queries that got a thumbs down, were abandoned, or were followed by a human escalation.
In Australia, treat those logs as likely to contain personal information under the Privacy Act 1988. Staff type names, customer details, health information and member numbers into internal assistants all the time. Before anyone labels a sample:
The OAIC and CSIRO's Data61 have published a De-identification Decision-Making Framework that is a sensible reference point. The goal is a test set that keeps the shape of real questions (the shorthand, the typos, the double asks) without keeping the people in them.
Deliberately add questions the corpus cannot answer: real ones from the logs, plus a few written on purpose. Mark the expected behaviour as "declines and points to the right channel". Then score it. A system that confidently answers an unanswerable question should fail that case just as clearly as one that gets a factual answer wrong.
This is also where outdated-policy questions belong. The expected result is not just the right number, it is the current number with the current document cited, ideally with a note that the rule has changed.
For every case, record two things. First, did the retriever return the document or section a person would need, in the top few results? Second, given what was retrieved, was the final answer correct, complete and faithful to the source?
Keeping these apart tells you where to spend effort. If retrieval misses on shorthand queries, the fix is query rewriting, synonym lists, metadata or better chunking, not a new prompt. If retrieval is fine but answers drift, the fix sits in generation.
Engineers should build the harness. Business owners should own the questions. The HR policy lead, the claims operations manager or the contact centre team leader knows what good looks like and which wrong answers carry real risk.
In practice that means a named owner per domain who reviews new samples, agrees the expected answers and signs off before a release. It does not need to take long. An hour a fortnight with a well prepared sample is often enough to keep the set honest.
Set a regular refresh (monthly is a reasonable starting point) where a fresh sample of real queries is labelled and added. Also trigger a refresh whenever the corpus changes in a meaningful way: a policy rewrite, a new document collection, a system rename, or a new team getting access. Retire cases that no longer reflect how people ask, but keep a stable core so you can still compare releases over time.
This pairs naturally with document ownership. If nobody owns what goes stale in the corpus, the eval set will drift right along with it.
Track citation correctness as its own metric. Does the cited document actually support the claim, and is it the current version? An answer that is right by luck while citing an archived policy is a problem waiting to happen, particularly in regulated settings where someone may later need to show why a decision was made.
Before rebuilding anything, compare your current eval set with a de-identified sample of real queries side by side. Look at a few simple things: average query length, how often abbreviations or internal jargon appear, how many questions have no answer in the corpus, and how many need more than one document.
If the two profiles look very different, you have your answer. Run the existing system against the real sample and score it the same way. The difference between that score and the dashboard score is the honest measure of how much your eval set was flattering you.
TruFyre works with Australian and APAC organisations to take RAG and GenAI systems from pilot to dependable production. That includes designing evaluation harnesses that separate retrieval from generation, setting up privacy-aware query sampling and labelling workflows, and helping business owners take real ownership of what "correct" means in their domain.
If your RAG pilot passed its tests but staff still do not trust it, the questions in your eval set are a good place to start. You can find us at trufyre.ai.
