Foundations
Human in the loop is the job
Tobiloba Odejinmi · 21 Apr 2026 · 6 min · 1,280 words

Direct answer
Yes. Human in the loop is the job, not a safety sticker you add after the demo. An AI employee does the first pass, and a named person takes gated cases with working state rather than a transcript dump. If you cannot staff that person, you have a model with permissions, not a loop.
- The reviewer is a role with a rota, not a checkbox in a vendor UI.
- Handoff is working state: facts, policy, next action.
- Aim for a modest escalation rate on a scoped queue, around 10 to 15 percent.
- Hard gates beat confidence scores on refunds, legal, and medical.
What does human in the loop actually mean?
It means the design assumes a person will take a slice of the work, forever. Not until the model 'gets there'. Models get better at fluency. Fluency is not the same as authority.
I treat the person as the product. The AI employee is how you stop that person from drowning in the easy eighty percent. The job you are hiring for is the twenty, or the fifteen, or whatever your gate produces.
If that sounds like you still need staff, good. You do. IBM's 2026 picture of 'Frontier Firms' is human-led and agent-operated, and only a small share of leaders say they run that way. Nine percent. The rest are still arguing about tools. The loop is the argument that matters.
Why is a transcript dump not a handoff?
Because a transcript is a story of confusion. The customer changed the order. The bot asked twice. The policy snippet was old. A person should not replay that novel to find the address.
Working state is a form. Case id. What was verified. What is still unknown. Which rule matched. What the system did. What it must not do. What you need from the person in the next ten minutes.
I have seen this in support and in clinic-adjacent ops. The teams that stay calm at night are the teams whose handoff looks like a chart, not a novel. You can teach a chart. You cannot staff a novel.
What is a hard gate?
A hard gate is a stop that does not care how sure the model feels. Refunds over a number. Any medical instruction. Any legal claim. Any change to credit terms. Any outbound that names a diagnosis, a debt, or a firing.
Soft gates are 'if confidence is low'. Soft gates fail the week you are behind, because someone lowers the threshold. Hard gates fail in the open. You can argue about the list. You should not argue about whether the list exists.
Write the list with the person who already gets the angry call. They know the cases. Product people know the demo. Those are different educations.
- Money out: person required.
- Health instruction: person required.
- Legal or HR letter: person required.
- Everything else: policy plus a miss budget you review weekly.
What escalation rate should you aim for?
You want enough escalations that the gate is real, and few enough that the first pass is doing work. Around 10 to 15 percent is a sane place to start on a single process that you actually understand.
If you are at 40 percent, the policy is thin or the input is wilder than you admitted. Fix the form. Fix the SOP. Do not train the reviewer to rubber-stamp so the dashboard looks better.
If you are at 2 percent, go read the closed cases. Either you picked a toy, or the system is closing things it should not. I would rather find that in a review than in a complaint from a provider.
Who is on the rota?
Name two people. Write the hours. Write what happens on Sunday. If the answer is 'we will figure it out', you figured out that nobody is on the rota.
The rota should already understand the process. Training a stranger on exceptions while the model is live is how you get two sources of error. Train them on the SOP first, with the model off, on a sample of old cases.
At 10mg, the people who keep clinics online are not the people who write slogans. Same here. Put the loop on the people who already stay late. Then use the first pass to give them their evenings back. That is the sale. Not 'the AI is in charge'.
What does a good miss look like?
A good miss is caught, explained, and turned into a rule or a new gate. It is not a blame meeting. It is a change to the page the worker reads.
A bad miss is silent. The customer ate it. Nobody can replay it. The team says 'the model hallucinated' as if that were a weather event. It is not weather. It is a missing check.
Review misses once a week for the first month. Short meeting. Ten cases. One change to the SOP. If you cannot spare that hour, you cannot spare the hire. The loop is the job.
Write one miss in the SOP as an example, not as shame. 'On 12 March we closed X because the policy did not mention Y. The stop list now includes Y.' New reviewers should see that page in week one. A loop that only lives in a meeting will die when the meeting gets cancelled.
What do you give the person on the loop?
A screen that shows the file, not the novel. Case id. Fields. Policy line. What the worker did. The button that agrees, the button that rejects, the box for a reason. If they have to open six tabs, you did not build a loop. You built a scavenger hunt.
A way to pause the worker. One control. One person who may use it. When a miss pattern appears, you pause, you fix the SOP, you resume. If pause requires a ticket to a vendor, you will not pause. You will hope. Hope is not a control.
Time. The loop is not extra work you add to a full day and then act shocked when people rubber-stamp. If the first pass saves three hours, some of those hours belong to the rejects. Steal all of them back to 'more tickets' and the quality will fall in week two.
I have kept clinics online. The people who do that job well get a runbook they can read at 2am. Give your reviewer the same respect. A pretty dashboard that does not say what to do next is decoration. Decoration does not close a case.
Questions people ask
What does human in the loop mean for an AI employee?
A person is in the path for cases the system must not close. They receive a file they can act on. They are not there to 'watch the AI' as a hobby.
Is human in the loop just for compliance?
No. It is how you keep quality when the input is messy. Compliance is one reason. Customers and sleep are others.
What is a good escalation rate?
On a well-scoped process, something like 10 to 15 percent is a working target from 2026 handoff research and from queues I have watched. Much higher and you built a router. Much lower and you should check the gate.
Why is a chat transcript a bad handoff?
Because the person has to reconstruct the case. That wastes the time you thought you saved. Send the verified facts, the rule that fired, and the action required.
Who should be the human in the loop?
Someone who already owns the exception today. Not an intern who has never seen the weird case. Not 'the Slack channel'. A name, a backup, and hours.
Written by
Tobiloba Odejinmi
Head of Engineering at 10mg Health. I have run engineering at Zeeh Africa and sold Insurpass and Shopl. I still write the code. If you have one process that still runs on people copying things, we can look at it in thirty minutes.


