AI Employees: What “Digital Workers” Can and Cannot Do
Vendors now sell AI employees with names, job titles and start dates. Underneath is an AI agent: useful for bounded, checkable work, unable to be accountable for any of it. What they can do today, what to ask a vendor, and who answers when one gets it wrong.
7 min read
An “AI employee” is an AI agent sold with a job title: a language model with instructions, a set of connected tools and permission to act, packaged as a support rep, a sales development rep or a bookkeeper. Today these digital workers can do bounded work with a clear check, such as drafting replies, sorting tickets, researching accounts or making routine code changes, faster than a person and at any hour. What they cannot do is be accountable. They do not answer for a mistake, they cannot hold a license, and they do not reliably know when they are out of their depth. So the useful question is not whether to hire one, but which person owns its work, what it may do without asking, and how you will check it.
This post is about the claim and the supervision. For forty concrete jobs by team, see AI agent use cases; for six companies that have documented their own deployments, see real-world AI agent examples.
What vendors mean by AI workers and digital employees
Strip the marketing and most AI employee products share the same parts. There is a model, usually from one of the large labs. There is a system prompt that gives it a role and rules. There are connections to your tools: email, a CRM, a help desk, a code repository. There is a loop that lets it take several steps without a person typing each one. And there is a persona on top: a name, a face, sometimes a weekly report written in the first person.
None of that is a problem in itself. The persona becomes a problem when it implies things the software does not do: that it understands your business the way a person who has worked there for a year does, that it will notice when a situation is new, or that someone is responsible for it the way a manager is responsible for staff. The glossary entry for an AI agent covers the plain definition; the rest of this post is about the gap between that definition and the job title.
What AI employees can do today
The most honest public test of an AI employee so far is one a model maker ran on itself. In Project Vend (opens in a new tab), published June 27, 2025, Anthropic gave Claude a small store in its office and the task of running it at a profit, with web search, email, Slack and control of prices. It found suppliers well and resisted attempts to make it misbehave. It also gave in to requests for discounts, for a time told customers to pay into an account it had hallucinated, and claimed it would deliver products in person in a blazer and tie. Anthropic’s conclusion: “we would not hire Claudius.”
The second phase (opens in a new tab), published December 18, 2025, gave the agent running the store newer models, a CRM and more structure, and the store did better: it sourced items reliably and priced with a margin. The same report says the models “still needed a great deal of human support” and that “the gap between ‘capable’ and ‘completely robust’ remains wide.” That is a fair summary of the whole category. Digital workers do well at:
- Work with a mechanical check: tests that pass, a reconciliation that balances, a link that loads.
- Gathering and sorting: research briefs, ticket routing, grouping feedback by theme.
- First drafts a person will read: replies, summaries, release notes, job postings.
- High-volume routine steps inside a process someone else designed and owns.
What they cannot do
- Answer for a result. An agent cannot be disciplined, sued, licensed or promoted, so the accountability for its work lands on whoever deployed it.
- Hold the line under pressure on their own. Vend’s agent gave away discounts because customers asked nicely; a support agent can be argued into a refund the same way unless a rule and a person stand behind it.
- Know what they do not know. They will produce a fluent answer where a person would say “I need to check,” unless you tell them to stop and ask.
- Notice that the job has changed. A new policy, a new product or a new customer segment reaches the agent only when someone updates its instructions.
- Carry context between jobs by default. Whatever it learned yesterday is gone unless it was written somewhere it reads.
Reading a vendor claim: what the FTC has said
US regulators have already acted against AI claims that promised more than the product did. On September 25, 2024, the FTC announced Operation AI Comply (opens in a new tab), a set of cases against companies using AI hype to mislead customers, with the line that “there is no AI exemption from the laws on the books.”
The case closest to the AI employee pitch is DoNotPay (opens in a new tab), which called itself “the world’s first robot lawyer.” The FTC said the company had not tested whether its chatbot’s output matched a human lawyer’s and had not hired any attorneys itself. The order, finalized January 17, 2025, bars it from claiming its service substitutes for a professional service without evidence. Before you buy a digital worker, ask the questions that case turned on:
- What exactly does it do, end to end, and which steps does it take without a person approving them?
- What is the evidence that it does the job as well as the person it is said to replace, and was that tested on work like yours?
- What does it do when it is unsure: guess, ask, or stop?
- Where is the record of every action it took, and can you read it without the vendor?
- How do you pause it, narrow it or switch it off, and how fast does that take effect?
A vendor with good answers will give them plainly. Treat any figure in a sales deck as the vendor’s own claim until you have seen it hold on your own work.
Supervision: who answers for an AI employee
NIST’s AI Risk Management Framework (opens in a new tab) is intended for voluntary use, and its answer to accountability is that it stays with people. The framework’s core (opens in a new tab) asks that “accountability structures are in place so that the appropriate teams and individuals are empowered, responsible, and trained,” and that “processes for human oversight are defined, assessed, and documented.” For a small team, that comes down to four things:
- A named owner for each AI worker: one person who reads its output, changes its instructions and switches it off.
- A written job description that says what it may do alone, what it must ask about, and what it never does.
- A review rhythm: every output at first, then a sample, widening only after a run of good work, with that decision written down.
- A record of what it did, under its own name, kept where the owner can read it.
Where in a workflow a person should step in is the subject of human in the loop for AI agents. A one-page team policy for who may connect which agents is in AI agent governance for small teams.
A job description for an AI worker
Write it before you switch anything on. It doubles as the core of the agent’s instructions; turning it into a full system prompt is covered in system prompt examples.
Role: Support triage for our online store (US customers)
Owner: Dana Ruiz, support lead
Does alone: tags new tickets by topic and urgency; drafts replies
from the help center; files a bug for repeated complaints
Asks first: any refund, credit or exception to the 30-day return policy;
any reply to a complaint that mentions a lawyer or the BBB
Never: sends a reply without a person; changes an order; promises
a ship date
When unsure: leaves the ticket untagged with a note saying why
Checked by: Dana reads every draft for two weeks, then a daily sample
Record: every action logged under the agent's name
Switch off: Dana revokes its access; nothing else depends on itWhere a task board fits
An AI employee’s work needs the same place a person’s does: a list of what it was asked to do, what it did, and who checked it. On a fenbs board an AI assistant connects over MCP as the person who connected it, narrowed by the scopes they approved, and everything it does is recorded in History as, for example, “Claude via Dana.” Each job is a task with a note that says the problem and a plan that says how it will be done.
Only a person can approve a task for an assistant, by pressing “Let AI do this” with an optional limits line. An assistant then takes it with fenbs_next_approved_task, records how it was tested, and the card reads “AI done · check it” until a person presses “I’ve checked it.” Standing limits, such as “never send a customer reply without a person,” go on the Decisions and rules page, and every connected assistant reads the rules first. fenbs does not run or schedule agents, and it has no due dates or settable assignee; it is the record and the queue, not the worker.
Related
Ideas by team: AI agent use cases. Documented deployments: real-world AI agent examples. Pre-approving work for an assistant: AI agent approval workflow. Connecting one to a board: the MCP docs.