You Built Human Oversight. Nobody Owns the Escalation Queue.




The governance deck says every high-risk answer goes through a human. The architecture diagram even has a tidy box labelled "reviewer". Then the agent starts escalating twenty times an hour, the queue has no named owner, and the same three people quietly stop opening it. Oversight did not fail in policy. It failed in operations.
That pattern shows up across Australian enterprise AI programmes. Human-in-the-loop sounds responsible. Without ownership, capacity, SLAs, and a path for unresolved cases, it is theatre. The model keeps moving. The review queue does not.
Most "human oversight" designs stop at the diagram. Route low-confidence outputs to a person. Log the decision. Resume the workflow. That is a control sketch, not a service.
A working escalation path needs the boring details: which role owns the queue on Tuesday, what "done" means, how long a case can sit, what happens after hours, and who can override when the reviewer disagrees with the model. If those answers live only in a slide, production will invent its own answers under load.
Invented answers look like auto-approve after ten minutes, Slack pings that nobody acknowledges, or a shared mailbox that becomes a graveyard. The audit trail still says "human reviewed". Reality says the human never had a chance.
Pilots hide the problem. Ten escalations a day feel manageable. A contact-centre assistant, a claims triage bot, or a procurement agent can push hundreds once traffic is real. Reviewer capacity does not scale with token spend.
When volume jumps, teams usually tighten the model threshold so fewer items escalate. That can be sensible. It can also silently move risk back into automation without updating the risk register. The oversight story stays on the website. The actual control threshold drifted in a config file.
Measure both sides. Track escalation rate, median time to first review, age of the oldest open case, and how often items expire into default approve or default deny. If you only watch model accuracy, you will miss the queue dying.
Name a single accountable owner for the escalation service, not a committee. That owner does not need to clear every ticket. They need budget for staffing, clear playbooks, and authority to pause the agent when the queue breaches its SLA.
Pair them with the process owner on the floor. An AI platform team can run the tooling. A claims, credit, or HR operations lead owns whether a delayed human decision is acceptable. Split those roles and you get tickets with no one who can say stop.
Write the handoff. What context does the reviewer see. Can they see the retrieved documents, the tool calls, and the confidence reason. Can they edit the draft or only approve or reject. Ambiguous UI creates rubber-stamping, which is worse than no review because it creates false assurance.
Some escalations will not resolve cleanly. The customer hangs up. The document is missing. The policy conflicts. Your design needs an explicit terminal state: hold, reject with reason, escalate to a senior desk, or schedule a human callback. Leaving items open forever trains staff to ignore the queue.
Also design for reviewer disagreement. If three reviewers reverse the model often on the same intent, that is not a staffing problem alone. It is a signal to retrain, rewrite policy, or narrow the agent's scope. Feed those reversals back into evaluation sets. Oversight without learning is unpaid labour.
Before you call human oversight live, prove four things in a rehearsal week, not a workshop:
If any of those are missing, you do not have human oversight. You have a hope that someone will notice.
Responsible AI controls fail in the same place most operating models fail: ownership under load. Build the escalation queue as a real service with capacity, SLAs, and a kill switch. Put the human in the loop only where a human can actually finish the job. Anything else is a diagram that looks safe until Tuesday morning.
