Runbook Template: Steps Anyone on Call Can Follow

A runbook is the page someone opens at 2 a.m. when an alert fires. The six sections that make one usable by whoever is on call, a template to copy, a filled example for a stuck job queue, which steps an AI agent may run, and how to keep the page true.

7 min read

A usable runbook template has six sections: the trigger that sends someone to the page, an impact check that says how bad it is, numbered steps that each show the expected output, a rollback for every step that changes something, an escalation line with names and a time limit, and a verify step that proves the problem is gone. Put an owner and a last-tested date at the top, and link the page from the alert itself. The template and a filled example for a stuck job queue are below, followed by the rules for letting an AI agent run parts of it and the habits that stop it going stale.

What a runbook is, and what it is not

AWS’s Well-Architected guidance defines a runbook (opens in a new tab) as “a documented process to achieve a specific outcome,” and says that at its simplest it is “a checklist to complete a task.” The same guidance asks for the outcome to be stated clearly, for special permissions and tools to be listed, and for guidance on error handling and escalation. That is most of a template already.

  • A runbook fixes a known problem: “the export queue is stuck, here is how to unstick it.” It is written for one situation and followed under pressure.
  • A playbook (opens in a new tab), in AWS’s terms, is for investigating: it guides someone through finding the cause, and once the cause is found it points to the runbook that fixes it.
  • An SOP is the agreed way to do routine work that happens every week whether or not anything is broken, such as onboarding or month-end close. It is covered in the SOP template.

The words overlap in practice. Google’s SRE book calls its alert-response pages playbooks, and its introduction (opens in a new tab) reports that recording best practices ahead of time in a playbook “produces roughly a 3x improvement in MTTR as compared to the strategy of ‘winging it.’” Whatever your team calls the page, the test is the same: can someone who has never seen this failure follow it and get the system back?

The six sections of a runbook

  1. Trigger: the alert name, symptom or customer report that sends someone here, and what is out of scope. If the page does not match what the reader sees, they need to know in ten seconds.
  2. Impact check: two or three read-only commands or dashboards that say how many users are affected and whether it is getting worse. This sets the pace for everything else.
  3. Steps with expected output: numbered, one action each, with the exact command and what a healthy result looks like. A step without expected output leaves the reader guessing whether it worked.
  4. Rollback: for every step that changes something, how to undo it. Write it next to the step, not at the bottom.
  5. Escalation: who to call, how, and after how long. “If the queue is not draining 15 minutes after step 5, page the platform lead.” A time limit stops someone struggling alone for an hour.
  6. Verify: how you know it is fixed, and how long to watch before you close the incident. Usually the same checks as the impact step, now showing healthy numbers.

Above the six sections goes a short header: the service, the owner, the date someone last ran the page end to end, and the access needed before step 1. Access is the part people forget. A runbook that needs a production role the on-call person does not have fails at step 1, at the worst time to find out.

A runbook template you can copy

Runbook template
# Runbook: [symptom, in the words of the alert]

Service: [name]            Owner: [role]
Last run end to end: [Month day, year] by [name]
Access needed: [roles, tools, VPN], request from [where]
Linked from: [alert names or dashboards]

## 1. Trigger
Use this when: [alert / symptom / report]
Not for: [look-alike problems, and where to go instead]

## 2. Impact check (read-only)
- [ ] [command or dashboard]  -> healthy: [value]  bad: [value]
- [ ] [who is affected, how many, getting worse?]

## 3. Steps
1. [READ]   [command]
   Expect: [output]
2. [CHANGE] [command]
   Expect: [output]
   Rollback: [command that undoes it]
3. [CHANGE] [command]
   Expect: [output]
   Rollback: [command]
Stop and escalate if any output does not match.

## 4. Escalation
- After [N] minutes without improvement: [role], via [channel]
- Customer-facing message needed? Tell [role].

## 5. Verify
- [ ] [same checks as step 2, now healthy]
- [ ] Watched for [N] minutes after the fix

## 6. Afterward
- File a task for anything in this page that was wrong or missing.
- Update "Last run end to end" above.

A filled example: a stuck job queue

An invented service that sends invoices through a background queue. The commands are placeholders for your own tools; the shape is what matters.

Runbook example
# Runbook: Invoice queue not draining

Service: billing-worker     Owner: Platform lead
Last run end to end: September 15, 2026 by Dana (drill)
Access needed: ops-readonly role; ops-admin for steps 3-4
Linked from: alert "InvoiceQueueDepthHigh"

## 1. Trigger
Use this when: queue depth above 500 for 10 minutes.
Not for: invoices sent but wrong (see "Invoice totals wrong").

## 2. Impact check (read-only)
- [ ] ops queue stats invoices     -> healthy: depth < 50, rate > 0
- [ ] Oldest job age                -> over 30 min means late invoices
- [ ] ops workers list billing      -> healthy: 3 workers, all "running"

## 3. Steps
1. [READ]   ops logs billing-worker --since 30m | grep ERROR
   Expect: nothing, or one repeated error naming a job id
2. [READ]   ops queue peek invoices --oldest
   Expect: a normal job; note its id
3. [CHANGE] ops queue move invoices <job-id> --to invoices-dead
   Only if step 1 shows the same job id failing over and over.
   Expect: depth starts falling within 2 minutes
   Rollback: ops queue move invoices-dead <job-id> --to invoices
4. [CHANGE] ops workers restart billing --one-at-a-time
   Expect: 3 workers "running" again within 5 minutes
   Rollback: none needed; restarts are safe
Stop and escalate if depth is still rising after step 4.

## 4. Escalation
- 15 minutes after step 4 with no drain: page the platform lead.
- Invoices more than 2 hours late: tell the support lead.

## 5. Verify
- [ ] Depth under 50 and falling; oldest job under 5 minutes
- [ ] Watched for 20 minutes; dead-letter job has a task filed

Notice what the example does not say: why the job failed. Finding that is investigation, and it belongs in a follow-up task, not in the middle of the fix. The runbook’s job is to get invoices moving again safely.

Writing steps anyone on call can follow

  • Write for the least experienced person on the rotation, at night, on a phone. Full commands, no “the usual flags”.
  • Mark each step READ or CHANGE. It tells a person where to slow down, and it tells an AI agent where to stop.
  • Say when to stop. Every change step needs a condition under which the reader stops and escalates rather than improvising.
  • Test it on someone else. AWS’s guidance is to validate a new runbook by having someone else on the team run it; a drill in a test environment counts.
  • Link it from the alert. Google’s SRE workbook chapter on on-call (opens in a new tab) says each alert should have a corresponding playbook entry, so the page is one click from the page that woke you up.

Runbooks an AI agent may run

An AI agent with shell access can follow a runbook, and the READ steps are where it earns its keep: gathering the impact numbers, the logs and the oldest job while a person is still opening a laptop. The CHANGE steps are different. Give the agent only read access by default, and make every change step a point where it stops, shows what it will run and waits for a person to say yes.

This is not a new idea from AI. AWS Systems Manager automation runbooks have an aws:approve step (opens in a new tab) that pauses the automation until designated approvers approve or reject it. Put the same pause in a runbook written for an agent: “Step 3 moves a job. Show the job id and the command, then stop until a person approves.” If an agent’s own actions become the incident, that is a different page: AI agent incident response. How approvals work in general is in the AI agent approval workflow.

Keeping runbooks current

The SRE workbook is blunt about decay: details “go out of date at the same rate as production environment changes.” Three habits keep up with it:

  • Update after every use. The person who followed the page fixes what was wrong before closing the incident, or files a task that says exactly which step failed.
  • Keep a last-run date and an owner at the top. A page nobody has run in a year is a guess; schedule a drill for it like any other piece of work.
  • Automate the deterministic ones. The workbook recommends automation when a playbook is a fixed list of commands run the same way every time. What remains in the runbook is judgment: when to stop, and who to call.

Where fenbs fits

fenbs does not store runbooks; keep them in your repository or docs, next to the code they operate. What fenbs holds is the work a runbook produces. Every gap found during an incident becomes a task: a bug with a ref such as BUG-212 when a step was wrong, an enhancement such as ENH-213 when a step was missing. The note says which runbook and which step, the plan says what will change, and the test status records whether someone has run the fixed page. Priority runs 1 to 10, 1 the most urgent, so a runbook that failed during a real incident can sit above routine work.

The line “an AI assistant may run READ steps, and a person approves every CHANGE step” is a rule, not a task. On the Decisions and rules page, a rule is a decision that holds from now on, with the person who made it recorded as the decider, and every connected AI assistant reads the rules first, before it touches the board. A runbook update with a clear plan can be pre-approved for an assistant with “Let AI do this” and a limits line such as “docs only”; only a person can give that approval.

Related

Routine procedures rather than fixes: SOP template. Errors an agent can read while following the page: Sentry MCP. What an assistant reads before it starts: AI context. When an agent is the cause: AI agent incident response.

Questions people ask.

What is the difference between a runbook and a playbook?

In AWS’s usage, a playbook guides an investigation to find the cause of an incident, and a runbook is the set of steps that fixes a known cause. Some teams, including Google SRE, use playbook for both. What matters is that each page says what it is for.

What should a runbook include?

The trigger, a read-only impact check, numbered steps with the exact command and expected output, a rollback for each change, an escalation line with a time limit, and a verify step. Add an owner, a last-run date and the access needed at the top.

Should an AI agent run runbooks?

It can run the read-only steps safely, which saves time early in an incident. Steps that change production should pause for a person to approve, the same way an approval step pauses an automation runbook.

How often should runbooks be updated?

After every use, by the person who followed them, and whenever the system they describe changes. A scheduled drill catches the pages nobody has needed recently.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.