Real-World AI Agent Examples From Companies

Six AI agents that companies have documented themselves, at Stripe, Airbnb, Google, Uber, Klarna and Anthropic: what each one does, how people supervise it, and the pattern worth copying.

7 min read

The clearest AI agent examples are the ones companies have written up themselves: Stripe’s coding agents that end at a pull request, Airbnb’s and Google’s code migrations, Uber’s on-call assistant in Slack, Klarna’s customer service assistant and Anthropic’s sales agent. They differ in job, but they share a shape. Each has a narrow task, its own tools, an automatic check such as tests or user feedback, and a person at the end who reviews the work or takes over when the agent cannot finish. Below are the six, each with what the agent does, how it is supervised, and what to copy.

Every example here comes from the company’s own engineering blog, research paper or announcement. Any figure quoted is the company’s own and has not been checked independently. For ideas by team rather than documented deployments, see AI agent use cases; for writing one of these jobs as a card an assistant can finish, see AI agent task examples.

1. Stripe: coding agents that stop at a pull request

Stripe’s engineering blog describes Minions (opens in a new tab), its in-house coding agents, in a post from February 9, 2026. An engineer starts one from a Slack message or another internal tool, and it works “fully unattended” in an isolated developer environment that Stripe says is cut off from production resources and the internet.

  • What it does: writes the change from start to finish and opens a pull request that passes CI, with at most two rounds of CI allowed.
  • How it is supervised: people review the code before it merges, through the same review a human engineer’s change gets. Stripe reports that over a thousand merged pull requests each week are produced entirely by minions.
  • What to copy: the agent’s finish line is “ready for review,” never “shipped.”

2. Airbnb: a test migration with gates at every step

Airbnb’s engineering team wrote up how it moved nearly 3,500 React test files from Enzyme to React Testing Library using a large language model pipeline (opens in a new tab). Each file passed through steps, refactor, then test fixes, then lint and type checks, and moved on only when the step’s check passed.

  • What it does: rewrites a file, runs the checks, and retries with the errors fed back into the prompt when a check fails.
  • How it is supervised: engineers worked in “sample, tune, sweep” loops, trying prompt changes on a few failing files before running them on all of them, and finished the last files by hand. Airbnb puts the whole migration at six weeks against an original estimate of a year and a half.
  • What to copy: every step has a mechanical check, and the files the agent cannot finish go to a person with the agent’s attempt as a starting point.

3. Google: code migrations reviewed like any other change

Google engineers published an experience report on internal code migrations (opens in a new tab) in January 2025, covering changes such as moving identifiers from 32-bit to 64-bit integers. The authors are careful to say it is not a research study, only an account of what they did.

  • What it does: finds the places that need changing and proposes the edits, inside a workflow that also runs validation steps.
  • How it is supervised: the report says humans review the generated code “the same way as any other code,” with changes split and sent to the owners of each part of the codebase. It reports that 80 percent of the code modifications in landed changes were fully AI-authored, and that review and rollout were still largely human-driven.
  • What to copy: no separate, lighter review lane for agent work. The code owners who would review a person’s change review the agent’s.

4. Uber: an on-call assistant that hands off to people

Uber’s engineering blog describes Genie (opens in a new tab), an assistant that answers questions in the Slack channels where internal users ask on-call engineers for help. It answers from Uber’s internal wiki, its internal Stack Overflow and engineering documents, using retrieval rather than a fine-tuned model.

  • What it does: replies in the support channel with an answer drawn from the documentation, so the on-call engineer is not interrupted for questions the docs already answer.
  • How it is supervised: users rate each answer as Resolved, Helpful, Not Helpful or Not Relevant, the ratings are tracked as metrics, and anyone can escalate to the human on call. Uber’s post, from October 2024, reports a 48.9 percent helpfulness rate by its own measure.
  • What to copy: a feedback button on every answer and a one-step route to a person.

5. Klarna: a customer service assistant with live agents behind it

Klarna announced its AI customer service assistant (opens in a new tab) on February 27, 2024. It handles refunds, returns, payment issues, cancellations, disputes and invoice errors in chat, around the clock and in many languages.

  • What it does: resolves routine customer service errands in the chat window.
  • How it is supervised: Klarna says customers “can still choose to interact with live agents if they’d prefer.” Its first-month figures, including handling two-thirds of customer service chats, are Klarna’s own announcement.
  • What to copy: the customer always has a way to reach a person, and the agent works inside policies a person wrote.

6. Anthropic: a sales agent that hands larger deals to reps

On September 30, 2026, Anthropic described the buying agent its own sales team built (opens in a new tab) on its Claude Managed Agents product. It answers prospects’ questions about plans, security and data, recommends a plan, and can take a buyer through to a completed purchase. This is a vendor writing about its own product, so read it as such.

  • What it does: answers questions and handles straightforward purchases; for larger or more complex deals it hands the prospect to a sales rep with the full conversation attached.
  • How it is supervised: sales and content leads reviewed the system prompt before customers saw it, and changes were tested as separate versions before they went live.
  • What to copy: the handoff carries the context, so the person who takes over does not start from nothing.

One more is worth knowing because it shows what happens without that structure: Anthropic’s experiment letting an agent run a small office store, covered in AI employees.

What the six have in common

  • A narrow job. Not “engineering” or “support,” but one migration, one channel, one kind of errand.
  • An automatic check before a person sees it: CI, lint and type checks, user ratings.
  • A person at the finish line or one step away: code review, an on-call engineer, a live agent, a sales rep.
  • A defined handoff for the cases the agent cannot finish, with its work attached.
  • Instructions that people own and review, in at least one case before any customer saw them.

None of them is an agent left alone with a broad goal. The parts that look most autonomous, such as Stripe’s unattended runs, are the parts with the hardest boundaries around them: an isolated environment and a review gate at the end. Where people should step in is the subject of human in the loop for AI agents.

How to read an AI agent case study

  1. Who wrote it? An engineering blog about an internal tool, a research report and a vendor’s customer story have different reasons to be written.
  2. What exactly is measured, and against what baseline? “Helpfulness” and “resolution” mean whatever the author defined them to mean.
  3. Where is the person? If the write-up does not say who reviews the work, ask.
  4. What did it not do? The best write-ups, such as Airbnb’s, say what was finished by hand.

Copying the pattern on a board

A small team can copy the structure without building anything. Here is a fenbs board for a team trying the Stripe and Uber pattern on its own repository and help channel, with the four fixed lanes:

Example board
To Do
  FET-041  Agent drafts answers to help-channel questions from the docs
  BUG-044  Flaky test in checkout.spec (agent may fix; person reviews PR)
Next Up
  ENH-039  Upgrade date library (approved for AI; limits: no API change)
In Progress
  BUG-037  Retry double charge (Claude via Sam: reproducing on dev)
Completed
  ENH-035  Lint fixes across src/auth (AI done, checked by Priya)

Each card has a note with the problem and a plan with how it will be done. A person presses “Let AI do this” on the ones an assistant may take, with a limits line; the assistant takes them with fenbs_next_approved_task, records its test status and notes, and the card reads “AI done · check it” until a person presses “I’ve checked it.” History records every change under the name of the person or assistant that made it. fenbs does not run the agents, connect to CI or set due dates; it holds the queue, the limits and the record.

Related

What vendors mean by digital workers: AI employees. Writing the agent’s standing instructions: system prompt examples. Turning a pattern into cards: AI agent task board. Connecting an assistant: Claude integration.

Questions people ask.

What are some real examples of AI agents in companies?

Stripe uses in-house coding agents that open pull requests for human review. Airbnb and Google have used AI to migrate code at scale, with checks and code review. Uber runs an on-call assistant in Slack, Klarna a customer service assistant, and Anthropic a sales agent that hands larger deals to reps.

Do companies let AI agents work without human oversight?

Not in the documented examples. Each has a person at the end or one step away: code review before a merge, an on-call engineer to escalate to, live agents customers can choose, or a sales rep for larger deals.

Are the results in AI agent case studies reliable?

Treat them as the company’s own account. Check who wrote it, what was measured and against what baseline, and whether it says what the agent did not finish. Figures in vendor customer stories are rarely checked independently.

What makes these AI agent examples work?

A narrow job, an automatic check such as tests or user ratings, and a person at the end who reviews the work or takes over. Coding and support agents appear often in company write-ups because both jobs have a clear check: tests that pass, or a user who says the answer helped.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.