Your Offline Evals Look Green. Production Still Rewrites Every Answer.

Release date:
September 13, 2026
Hero Vector
Vector ImageVector ImageVector Image
Blog detail
Vector ImageVector ImageVector Image

Your offline eval suite is green. Latency is fine. The golden set scores look better than last sprint. Support still spends half the day rewriting answers before they reach a customer.

That mismatch is not a mystery. It is what happens when you treat a static scorecard as proof the system works in the wild. Offline evals are necessary. They are not the same thing as production fitness.

Green scores can hide the wrong job

Most teams build offline evals around questions they already know how to ask. A curated set of prompts. A handful of reference answers. A judge model or a rubric that rewards similarity to yesterday's "good" output.

That is useful for catching regressions. It is weak at catching the work that actually burns hours: ambiguous requests, incomplete context, policy edge cases, and the awkward follow-ups that never made it into the golden set.

If production users keep rewriting answers, they are telling you the eval suite is measuring a different job from the one they are doing.

What rewriting usually means

When operators rewrite, they rarely fix spelling. They fix trust.

  • They soften a claim that sounded too certain.
  • They add the caveat the model skipped.
  • They swap in the local policy version nobody indexed.
  • They remove a citation that points at a stale page.
  • They restructure the answer so a human can send it without embarrassment.

Those edits are a product signal. If you only log model scores and not rewrite effort, you will keep celebrating green bars while the real cost sits in Slack threads and support macros.

A pattern we keep seeing

An Australian insurer built a strong RAG assistant for claims handlers. Offline faithfulness scores sat above the team's threshold for weeks. In the contact centre, handlers still pasted the draft into a second document and rewrote the customer-facing paragraph almost every time.

The offline set rewarded answer completeness against the knowledge base. The handlers cared about tone, eligibility exceptions, and which product version applied to that policy year. The eval never asked those questions, so the model never learned they mattered.

Once the team logged rewrite reasons for two weeks, the top three tags were "wrong product year", "too confident", and "missing eligibility caveat". None of those showed up in the golden set. The next eval release changed overnight.

Make production the second judge

Keep the offline suite. Add a production loop that treats human edits as first-class evidence.

  1. Capture rewrite diffs where work happens. If people rewrite in the tool, store before and after. If they rewrite outside it, you have a workflow problem, not just an eval problem.
  2. Tag the reason, not just the sentiment. "Bad answer" is useless. "Wrong jurisdiction", "missing disclaimer", and "hallucinated step" are actionable.
  3. Sample live traffic into the eval set every week. Promote hard production cases into the offline suite so green scores mean something closer to reality.
  4. Score usefulness, not only similarity. Ask whether a trained human would send the answer with light edits, heavy edits, or a full rewrite.
  5. Separate retrieval failures from generation failures. Rewriting a confident wrong answer is different from rewriting a vague right answer. Your fixes will differ.

What good looks like

Healthy teams still rewrite. They just rewrite less, and they know why.

A practical target is not zero edits. It is a shrinking share of heavy rewrites, a rising share of light polish, and a clear backlog of eval cases that map to the reasons people still intervene. When leadership asks "are the evals good?", the honest reply includes both the offline score and the rewrite rate by reason code.

If those two numbers disagree, believe production. Then update the suite until they start telling the same story.

Start this week

Pick one production surface where people already rewrite model output. Log fifty consecutive edits with a short reason tag. Do not wait for a perfect taxonomy. Review the tags with the people who did the rewriting. Promote the top five failure modes into your next offline release.

You will probably find that your green suite was never wrong about the tests it ran. It was incomplete about the work your users actually do. Fix that gap and the rewrites start to fall for a reason you can defend.

BG Image
Vector ImageVector ImageVector Image
We’re here to help
Vector ImageVector ImageVector Image

Ready to put AI to work in your business?

Talk to an AI expert about your goals.
Arrow Icon
Smart process automation
Arrow Icon
Direct access to our team. No bots.
Arrow Icon
We ask smart questions fast.

Book a Discovery Call

Your form has been submitted successfully. Thank you!
Please double-check your information and try again. If the issue continues, email us at info@trufyre.ai