On-Call Rotation for Small Teams: A Fair Setup
A fair on-call rotation for a team of three to eight: a schedule people can plan around, a written handoff, a runbook behind every alert, and alerts worth waking up for. What large SRE teams recommend, and how a small team adapts it.
6 min read
A fair on-call rotation for a small team has five parts. Shifts last a week, change hands at a fixed time midweek, and always have a named backup. Every shift ends with a short written handoff. Every alert that can wake someone links to a runbook. Only urgent, actionable problems page anyone; everything else waits for business hours. And the rules for swaps, time off after a bad night and pay are written down before the first shift, not argued over after it. The details, and what to do when your team is smaller than the textbooks assume, are below.
What large teams recommend, and why small teams cannot copy it
Google’s SRE book, in its chapter on being on-call (opens in a new tab), sets limits worth knowing. No more than 25% of an engineer’s time should go to on-call. With a primary and a secondary on week-long shifts, that means at least eight engineers for a team at a single site. It also caps the load at two incidents per 12-hour shift, because handling one properly, including the follow-up and postmortem, takes about six hours. Typical response times it cites are 5 minutes for user-facing, time-critical services and 30 minutes for less time-sensitive systems.
Most small teams have three to six people, so every one of them will be on call more than a quarter of the time. That is not a reason to give up on fairness; it is a reason to shrink what on-call has to cover. The levers a small team actually has:
- Page for less. The fewer alerts that can wake someone, the more often each person can reasonably be on call. Alert hygiene, below, is the biggest lever.
- Choose the response target honestly. If your customers can tolerate 30 minutes overnight, do not promise 5.
- Cover business hours and weekends differently. Overnight pages can be limited to “the service is down,” with everything else queued for morning.
- Borrow a backup. A second-line contact from a neighboring team or a contractor can hold the secondary slot a small team cannot fill.
Setting up an on-call schedule
Weekly shifts suit most small teams: long enough to build context, short enough that nobody is on call for a month straight. Hand over midweek, at a time when both people are working, so the incoming person starts with a normal workday rather than a weekend. Pair every primary with a secondary, who is paged if the primary does not acknowledge in time and who covers short gaps.
Handoff: Wednesday 10:00 a.m. Central, in the team channel Response: 15 min business hours, 30 min overnight and weekends Pages overnight: only "service down" and "data at risk" alerts Week of Primary Secondary Sep 30 Ana Ben Oct 7 Ben Chris Oct 14 Chris Dana Oct 21 Dana Ana Swaps: agree with the other person, then update the schedule. After a night with a page between midnight and 6 a.m.: start late the next day, no questions asked.
Notice that the secondary this week is the primary next week. They have already seen the week’s problems, which makes the handoff shorter. The schedule itself should live in whatever tool pages you, so a swap changes who is actually paged, not just a spreadsheet.
On-call handoffs
PagerDuty’s open incident response documentation (opens in a new tab) puts the rule simply: when your shift ends, let the next on-call know about issues that have not been resolved and other experiences of note. The same guide asks people to check their upcoming shifts and arrange swaps around travel and vacations. A handoff note needs only a few lines:
On-call handoff: Sep 30 -> Oct 7 (Ana -> Ben) Pages this week: 3 (2 overnight) Still open: - Export queue backs up after 2 a.m. batch; watching, see BUG-204 - Disk on db-2 at 81%, cleanup scheduled Thursday Changed this week: - Checkout flag on for all users Monday; kill switch still in place Noisy alerts to fix or delete: - "CPU high on worker-3" paged twice, both harmless (ENH-206) Heads-up for next week: - Payment provider maintenance Saturday 1-3 a.m. Central
The “noisy alerts” line is the one that improves the rotation over time. Every alert that woke someone for nothing should leave the shift as a task to fix, tune or delete it.
A runbook behind every alert
The person on call at 2 a.m. may never have seen this failure before. Each alert should link straight to a page that says how bad it is, what to run, what to expect and when to escalate. How to write one is covered in the runbook template; the on-call rule is only that no alert goes live without one. When a runbook step turns out to be wrong during a real page, the fix is part of the shift, not a someday job.
Alert hygiene
The same Google book, in its chapter on monitoring distributed systems (opens in a new tab), gives four rules for pages that hold up well at any team size:
- “Every time the pager goes off, I should be able to react with a sense of urgency.”
- “Every page should be actionable.”
- “Every page response should require intelligence. If a page merely merits a robotic response, it shouldn’t be a page.”
- “Pages should be about a novel problem or an event that hasn’t been seen before.”
Turn those into a weekly habit. At handoff, look at every page from the shift and give it one of four outcomes: it was real and the runbook worked; it was real and the runbook needs a fix; it should have been a ticket for the morning, not a page; or it should be deleted. A page with a robotic response is a candidate for automation. Alerts nobody acts on train people to ignore the pager, which is how a real one gets missed.
Making it fair
- Everyone who can change production is in the rotation, including leads. A rotation that exempts senior people tells everyone else what on-call is worth.
- Swaps are easy and never need a manager’s permission, only an updated schedule.
- A bad night buys a late start. Write the threshold down so nobody has to ask.
- Decide compensation before the first shift. Google’s SRE book says it offers time off in lieu or cash, capped at a proportion of salary. Small teams often choose time off.
- Check the law for your employees. The Department of Labor’s Fact Sheet #22 (opens in a new tab) says an employee required to stay on call on the employer’s premises is working, while one on call at home generally is not, unless added constraints on their freedom require that time to be paid. State rules can differ; ask an employment lawyer for anything beyond the basics.
After a real incident, write it up without blame and turn the action items into tasks; the incident postmortem template covers both. If an AI agent’s own actions caused the page, follow AI agent incident response instead.
Where fenbs fits in an on-call rotation
fenbs does not page anyone, hold a schedule or know who is on call; keep that in your paging tool. There is no assignee field to set either, so the board is not where you record whose week it is. What fenbs holds is everything the rotation produces: the bug behind a page, the noisy alert to fix, the runbook step that was wrong. Each one is a task with a ref, like BUG-204 or ENH-206 in the handoff note above, so the note can point at the work instead of describing it.
- Priority 1 is the most urgent on a 1 to 10 scale, so an open problem from last night ranks above planned work without a separate list.
- History records who moved each task and when, which makes “what did on-call touch this week?” a question the board can answer.
- Standing rules such as “every alert links a runbook” or “no overnight page without a service-down alert” belong on the Decisions and rules page. Each rule records the person who decided it, and every connected AI assistant reads the rules first.
- An AI assistant helping with the handoff can list the tasks filed since last Wednesday over MCP and draft the note for a person to check.
Related
The page behind every alert: runbook template. After the incident: incident postmortem template. Deciding what gets fixed first: bug severity vs priority. Letting an AI assistant read your errors: Sentry MCP.