Your RAG Eval Set Was Written by the Team That Built It. Your Users Ask Different Questions.

Release date:
Hero Vector
Abstract editorial graphic dusty rose espresso brown oat neat stack of question cards linked by a retrieval path to an index, scattered messy cloud of shapes beyond, no text
Vector ImageVector ImageVector Image
Blog detail
Vector ImageVector ImageVector Image

Most enterprise RAG pilots have an evaluation set. Usually it is a spreadsheet of fifty or a hundred questions, each with an expected answer and a source document. The pilot runs against it, the scores look strong, and the steering committee signs off on a wider rollout.

Then real staff start using it. Within a few weeks the support channel fills with "it couldn't find the policy" and "it gave me the old rate". The eval dashboard is still green. Nobody has changed the model. The retrieval is quietly failing on questions the eval set never asked.

The cause is rarely the vector database or the embedding model. It is who wrote the questions. In most pilots, the eval set was written by the same people who built the system, and they ask questions very differently from the people who will use it.

The eval set that marks its own homework

Builders know the corpus. They have read the documents, chunked them, debugged them and watched the retriever pull them back. When they sit down to write test questions, they do it with the documents open in another tab.

That shapes the questions in ways that are hard to notice from the inside:

  • They borrow the document's vocabulary. If the policy says "parental leave entitlement", the test question says "parental leave entitlement". Retrieval loves that. Lexical and semantic overlap is about as high as it will ever get.
  • They only ask questions that have answers. Nobody writes a test case for a document they know is not in the index. So the system is never tested on what it should do when the honest answer is "I don't know".
  • They ask one thing at a time. Clean, single-intent questions map neatly to a single chunk. Real questions often do not.
  • They ask about the current version. Builders test against the documents as they stand today. Staff ask about the version they remember.

None of this is carelessness. It is the curse of knowledge. Once you know where the answer lives, it is very hard to phrase a question the way a confused, busy person would.

A hypothetical scenario

Picture an HR policy assistant at a mid-sized Australian insurer. The platform team loads the leave, travel, expenses and code of conduct policies, then writes eighty test questions. A typical one: "What is the parental leave entitlement for primary carers?" The assistant retrieves the right section and answers correctly almost every time.

In production, a team leader types: "new dad in my team, how many wks paid + can he split it". That one message has an abbreviation, a casual synonym, two questions joined together, and an assumption about splitting leave that sits in a separate flexible work procedure. The retriever pulls the general leave policy, the answer covers half the question, and the citation looks fine. The team leader gives up and emails HR, which is exactly what the assistant was meant to reduce.

The eval set had no way of catching that, because nobody on the build team would ever phrase a question like that.

What real query logs actually look like

If you pull a week of real, de-identified queries from almost any internal assistant, the same patterns show up. They look nothing like the golden set.

  • Shorthand and abbreviations. "PL", "TOIL", "LSL", "wfh", "approx". Each team has its own.
  • Internal jargon and old names. Staff use the name of the system before the last migration, or the nickname for a form, or the name of a team that was restructured two years ago.
  • Typos and fragments. Questions typed on a phone between meetings. Half sentences. No punctuation.
  • Multi-part questions. Two or three asks in one message, often depending on each other.
  • Questions with no answer in the corpus. Someone asks about a policy that does not exist, or one that lives in a system the assistant cannot see. The right response is a clear refusal and a pointer to who can help.
  • Outdated-policy questions. "Is it still 50 cents a km?" The person is anchored on an old rule, and the retriever may happily find the archived document that agrees with them.
  • Questions that need two documents. An eligibility rule in one policy and a process in another. Retrieval returns one, and the answer sounds complete when it is not.

Each of these hurts retrieval in a different way. Together they describe the gap between the questions you tested and the questions you get.

Why the dashboard stays green

An eval score is only a measure of how well the system handles the questions in the eval set. If those questions come from a different distribution to production traffic, a high score tells you very little about production.

A few things make this worse:

  • Retrieval and answer quality get blended into one number. A fluent answer built from the wrong chunk can still score well with an LLM judge, especially if the judge is only checking whether the answer sounds relevant.
  • Refusals are never scored. If every test question has an answer, a system that always answers looks perfect. A system that sensibly declines looks like it failed.
  • The set never changes. The corpus is refreshed, policies are rewritten, a new business unit is onboarded, and the same eighty questions keep passing.

We have written before about why RAG demos flatter retrieval and why offline evals can look green while people rewrite every answer. This is the upstream cause of both: the questions themselves were never representative.

How to build an eval set that looks like your users

Sample real queries, carefully and regularly

The single most useful change is to stop inventing questions and start sampling them. Take a regular random sample of real queries from the logs, plus a targeted sample of queries that got a thumbs down, were abandoned, or were followed by a human escalation.

In Australia, treat those logs as likely to contain personal information under the Privacy Act 1988. Staff type names, customer details, health information and member numbers into internal assistants all the time. Before anyone labels a sample:

  • Check that using query logs for quality testing fits the purpose staff and customers were told about, and involve your privacy officer early.
  • Strip or replace names, identifiers, account numbers and free-text details that could re-identify someone. Automated redaction is a good first pass, not a final one.
  • Keep labelled samples in a controlled location with access limited to the people doing the labelling, and set a retention period.

The OAIC and CSIRO's Data61 have published a De-identification Decision-Making Framework that is a sensible reference point. The goal is a test set that keeps the shape of real questions (the shorthand, the typos, the double asks) without keeping the people in them.

Include unanswerable questions and score correct refusals

Deliberately add questions the corpus cannot answer: real ones from the logs, plus a few written on purpose. Mark the expected behaviour as "declines and points to the right channel". Then score it. A system that confidently answers an unanswerable question should fail that case just as clearly as one that gets a factual answer wrong.

This is also where outdated-policy questions belong. The expected result is not just the right number, it is the current number with the current document cited, ideally with a note that the rule has changed.

Separate retrieval metrics from answer metrics

For every case, record two things. First, did the retriever return the document or section a person would need, in the top few results? Second, given what was retrieved, was the final answer correct, complete and faithful to the source?

Keeping these apart tells you where to spend effort. If retrieval misses on shorthand queries, the fix is query rewriting, synonym lists, metadata or better chunking, not a new prompt. If retrieval is fine but answers drift, the fix sits in generation.

Give ownership of the set to the business, not the engineers

Engineers should build the harness. Business owners should own the questions. The HR policy lead, the claims operations manager or the contact centre team leader knows what good looks like and which wrong answers carry real risk.

In practice that means a named owner per domain who reviews new samples, agrees the expected answers and signs off before a release. It does not need to take long. An hour a fortnight with a well prepared sample is often enough to keep the set honest.

Refresh on a cadence and when the corpus changes

Set a regular refresh (monthly is a reasonable starting point) where a fresh sample of real queries is labelled and added. Also trigger a refresh whenever the corpus changes in a meaningful way: a policy rewrite, a new document collection, a system rename, or a new team getting access. Retire cases that no longer reflect how people ask, but keep a stable core so you can still compare releases over time.

This pairs naturally with document ownership. If nobody owns what goes stale in the corpus, the eval set will drift right along with it.

Check the citation, not just the answer

Track citation correctness as its own metric. Does the cited document actually support the claim, and is it the current version? An answer that is right by luck while citing an archived policy is a problem waiting to happen, particularly in regulated settings where someone may later need to show why a decision was made.

A quick way to see the gap

Before rebuilding anything, compare your current eval set with a de-identified sample of real queries side by side. Look at a few simple things: average query length, how often abbreviations or internal jargon appear, how many questions have no answer in the corpus, and how many need more than one document.

If the two profiles look very different, you have your answer. Run the existing system against the real sample and score it the same way. The difference between that score and the dashboard score is the honest measure of how much your eval set was flattering you.

A practical checklist

  • A regular, de-identified sample of real user queries is labelled and added to the eval set.
  • Your privacy officer has reviewed how query logs are collected, redacted, stored and retained.
  • The set includes shorthand, typos, multi-part questions and questions that need two documents.
  • Unanswerable and outdated-policy questions are included, and correct refusals are scored as passes.
  • Retrieval metrics and answer metrics are reported separately.
  • Citation correctness, including document version, is tracked as its own metric.
  • Each domain has a named business owner who signs off on expected answers.
  • The set is refreshed on a fixed cadence and whenever the corpus changes meaningfully.
  • Release decisions use scores on real query samples, not just builder-written questions.

How TruFyre helps

TruFyre works with Australian and APAC organisations to take RAG and GenAI systems from pilot to dependable production. That includes designing evaluation harnesses that separate retrieval from generation, setting up privacy-aware query sampling and labelling workflows, and helping business owners take real ownership of what "correct" means in their domain.

If your RAG pilot passed its tests but staff still do not trust it, the questions in your eval set are a good place to start. You can find us at trufyre.ai.

BG Image
Vector ImageVector ImageVector Image
We’re here to help
Vector ImageVector ImageVector Image

Ready to put AI to work in your business?

Talk to an AI expert about your goals.
Arrow Icon
Smart process automation
Arrow Icon
Direct access to our team. No bots.
Arrow Icon
We ask smart questions fast.

Book a Discovery Call

Your form has been submitted successfully. Thank you!
Please double-check your information and try again. If the issue continues, email us at info@trufyre.ai