People

From pilot to production AI employees

Tobiloba Odejinmi · 23 Jul 2026 · 6 min · 922 words

A notebook with a seven-day grid sketched in pencil

Direct answer

Production is not a bigger pilot. It is a named owner, a stop condition, evaluation on your real cases, monitoring a stranger would shout about, and a handover someone else can run. If those are missing, you still have a demo, even if it has a logo and a Slack channel. IBM’s frontier share is about 9 percent for a reason.

  • A pilot without an owner will not grow one after launch.
  • Evaluation is your pile, not a public leaderboard.
  • Production means pause, logs, and a backup person.
  • The first real Monday is the test, not the demo recording.

Why most pilots stay pilots

IBM still puts frontier firms at about 9 percent. I do not treat that as a moral ranking. I treat it as a description of who put a name on the work. The other 91 percent are not stupid. They ran a pilot, got a nice paragraph, and then nobody wanted the queue.

A pilot is allowed to be ugly. Production is not allowed to be ownerless. The stall happens when leadership keeps asking for another use case instead of staffing the first one. That feels like progress in a slide. It is how you get six demos and zero Mondays. I would rather you keep one job alive through a bad week than open a second pilot so the roadmap looks busy.

What production actually means

Real tickets. Real CVs. Real follow-ups. A stop button the owner can use. Logs you can read without me. Escalations you can list. A backup person. That is the bar. If any of those are “we’ll add it after,” you are not in production. You are in a longer pilot with customers inside it.

Anthropic’s 2026 agents report is useful as a cold shower. The agents that hold up are bounded and evaluated. They are not colleagues you hired in a keynote. If your production plan is “let it loose and watch,” you are not deploying an employee. You are deploying a hope.

The first real Monday

I have sat in reviews where the uptime slide was green and a clinic still could not open the path they needed. Monday morning is the test. The same is true here. The queue that arrives after a weekend will not match the five golden examples from the demo.

Sit with the owner that morning. Take the escalations. Write the two rules you did not have on Friday. If the owner cannot do that job and also do their old job, you understaffed production. The model did not fail. The calendar did.

Evaluation that survives contact

Build a set from last month’s real work, including the ugly ones. Expected fields. Expected escalations. Run it before you go live and whenever you change the prompt, the model, or the sources. A pass on a vendor’s sample is not your pass.

When evaluation is weak, teams argue taste. When evaluation is a shape, they argue a field. I know which argument I want at 5pm. Structured output is what makes the set cheap to score. Essays are what make it a book club.

Monitoring after the glow

Watch whether a case can start, whether a decision comes back, whether escalations age, whether required fields start arriving empty. Those are the shouts. Token charts are optional. I will take a boring counter over a pretty dashboard that nobody can act on.

Write the stop condition next to the monitor. If schema failures jump, pause. If a tool goes dark, pause. If the owner is out and the backup is also out, pause. Production includes the right to stop. Pilots pretend stopping is drama. It is hygiene.

The handover you skip at your own cost

I learned what buyers open first: cost, uptime, and whether one person can explain the database. An AI employee that only I can run fails that meeting, and it fails the 2am version of that meeting when I am not in the room.

Handover is docs a tired person can follow, the owner, the backup, the monitors, and a call. If your pilot never had a handover, it was never going to be production. It was a visit. Visits do not take the pile off a team. Production starts the week the backup owner can pause it without asking me where the switch is.

Questions people ask

When is an AI employee in production?

When it runs on real work, an owner can pause it, exceptions have a queue, and someone who did not build it can explain a miss. A shared sandbox link is not production.

Why do so many pilots stall?

Nobody wanted the exceptions. The tools had no door. The process was still folklore. Or the pilot was a tour of a model, and the company never chose a job. BCG keeps pointing at the operating model. That is the stall.

Do we need a new team to go live?

You need an owner and a reviewer path. You do not need an “AI department” first. Departments are how pilots become programs with no pile.

What should we measure in the first month?

Review hours, miss rate on a sample, aging escalations, and whether the owner still believes the job description. Do not lead with token spend. Spend will lie if the work is wrong.

Can we stay in pilot longer?

Yes, if you are still writing the process. No, if you are using “pilot” to avoid naming an owner while real customers already see the output. That is production without the honesty.

Written by

Tobiloba Odejinmi

Head of Engineering at 10mg Health. I have run engineering at Zeeh Africa and sold Insurpass and Shopl. I still write the code. If you have one process that still runs on people copying things, we can look at it in thirty minutes.