People

Write a job description for the AI employee

Tobiloba Odejinmi · 11 Feb 2026 · 6 min · 921 words

A headset and notepad on a dark support desk

Direct answer

Write the AI employee a job description the way you would for a junior hire. What it does every day, which tools it may touch, what it must never do, who it escalates to, and how you will measure the work. If that page is vague, the workflow will be vague. A model cannot invent a role you refused to specify.

  • A one-page JD beats a folder of prompts.
  • Write the “never do” list before the happy path.
  • Name the tools, the hours, and the person who takes overflow.
  • Success is a measurable output, not “it uses AI.”

A JD beats a prompt dump

I have been handed folders of prompts and told the role was obvious. It never is. Prompts encode taste. A job description encodes the work. When the two disagree, the prompt wins until a customer gets the wrong answer. Then everyone discovers there was no contract.

Write the JD first. Then we can talk about retrieval, tools, and which model sits underneath. Anthropic’s 2026 agents report is useful as a constraint here: the agents that hold up are bounded. A JD is how you write the bound in language a non-engineer will still understand in three months.

What it does every day

Be boring. “Takes inbound tickets that match these three intents and drafts a reply in the helpdesk.” “Reads new applications and returns a shortlist of ten with the reasons attached.” “Looks up the lead, writes the first note, flags the ones that need a call.”

If you cannot say the daily work in two sentences, you are describing a department. Departments do not go live in a week. One process does. I would rather ship a narrow JD and add a second job later than ship a “digital teammate” that does a bit of everything badly.

Tools it may touch

List the systems by name. Helpdesk, CRM, inbox, calendar, the internal doc store. Say whether it may read, draft, or send. Sending is a different job from drafting. Most teams want sending on day one. Most teams are not ready.

If a tool is missing an API, write that down too. I will not pretend a pile of CSVs is a stable workplace. Zeeh grew because other companies could plug in before lunch. Your AI employee needs the same kind of boring door. If the door is a person forwarding emails, that person is still the integration.

What it must never do

This list is more important than the happy path. No refunds above a number. No legal language that is not in the approved set. No ranking of candidates without a human next look. No deleting records. No inventing a policy when retrieval comes back empty. Empty should escalate.

Hiring makes this concrete. If the “job” includes screening people, NYC Local Law 144 and the EU AI Act are not footnotes. You need a reviewer, an audit trail, and a story you can tell a candidate. Put that in the JD. Do not discover it after the shortlist went out.

Hours, volume, and the person behind it

Say when it runs and what happens when the pile is bigger than the rules. Overnight drafts are fine. Overnight decisions about money are not. Write the overflow owner. Write the expected volume you measured, not the volume you hoped.

I ask for a week of real counts before we build. If nobody has counted the pile, we count it. A JD that says “handles support” with no volume is a wish. A JD that says “about 80 first-response tickets a day, three intents, escalate the rest” is a job. The overflow owner is part of that sentence. Without them, the volume number is a wish that lands on whoever happens to be online.

How you will know it is working

Pick measures the owner already believes. Time to first response. Review hours on the pile. Miss rate on a weekly sample. Number of escalations that sat more than a day. Do not pick “AI usage.” Usage is how you get rubber stamps. A green usage chart with a quiet miss rate is how teams talk themselves into staying live too long.

Put the measures on the same page as the “never do” list. If the only way to look green is to skip review, the JD is lying. I will not take a workflow live on a lying page. The model is not the hero. The loop is. The JD is how you describe the loop before anyone writes a line of glue.

Questions people ask

Why write a job description for software?

Because the failure mode is scope creep in prose. A JD forces you to say what “done” is. Prompts hide that argument inside a paragraph nobody owns.

How detailed should it be?

One page. Daily work, tools, forbidden actions, escalation path, and two or three measures. If you need a novel, the process is not ready. Split it.

Who writes it?

The person who owns the process today. I will help them cut it. I will not invent the job from a brainstorm. They already know the pile.

Is this the same as a prompt?

No. The JD is the contract. The prompt is one implementation detail. If you change models next quarter, the JD should still be true.

What goes in the “never do” list?

Anything that moves money, changes access, makes a legal promise, or ranks a person without a reviewer. Also: inventing a policy that is not in the source you gave it.

Written by

Tobiloba Odejinmi

Head of Engineering at 10mg Health. I have run engineering at Zeeh Africa and sold Insurpass and Shopl. I still write the code. If you have one process that still runs on people copying things, we can look at it in thirty minutes.