Regression Testing Checklist: What to Re-Test Before Each Release
Regression testing checks that what worked last release still works after this one. How to choose what to re-test, what to automate and what to leave to a person, a checklist to copy, how every fixed bug becomes a permanent test, and how an AI agent can run the suite and file what fails.
7 min read
A regression testing checklist answers one question before each release: does everything that worked before still work? Build it from three sources: the areas this release changed and whatever depends on them, the flows that would hurt most if they broke, such as sign-in, payment and saving data, and every bug you have fixed before. Automate the checks that run the same way every time and leave a short list for a person: layout, new devices, and anything that needs judgment. Run the automated suite on every change, the full list before each release, and turn each newly fixed bug into a test so it cannot quietly return. The checklist to copy is below.
This page is about re-testing what already worked. The wider release check, covering browsers, accessibility, security and speed, is the QA checklist; writing a single test case well is covered in the test case template.
What regression testing is
The ISTQB glossary defines regression testing (opens in a new tab) as a type of change-related testing to detect whether defects have been introduced or uncovered in unchanged areas of the software. The key words are “unchanged areas”. Checking that the new feature works is testing the change; checking that the fix worked is a retest. Regression testing looks at everything around them: the report that shares a query with the new feature, the export that uses the date format someone just touched.
Smoke vs regression, in one line: ISTQB describes smoke testing (opens in a new tab) as a test type to gain sufficient confidence that a test object is ready for planned testing, so a smoke test asks “is this build worth testing?” in minutes, and regression testing is the planned testing that follows.
Choosing what to re-test
Re-running everything, every time, by hand, is not a plan; it is how regression testing gets skipped. Choose the scope from three lists and re-test their union.
- What changed, and what depends on it. Start from the tasks in the release and the files they touched, then follow the dependencies one step out: shared components, shared queries, shared configuration.
git diff --name-onlybetween the last release tag and this one is the honest list of what changed. - High-risk flows, every time. The paths where a failure costs money, data or trust: sign-up and sign-in, password reset, payment, saving and loading the main thing your product stores, permissions, and email or notifications. These are re-tested on every release whether or not anyone touched them.
- Past bugs. Every bug you have fixed is a place the code has already broken once. A bug that was reopened, or one in an area that keeps changing, is the most likely to come back.
Tier the result. A fast set runs on every change, typically the high-risk flows and recent bug tests; the full set runs before each release; and a slow set, such as long data imports or every supported device, runs nightly or weekly. Write the tiers down so nobody has to decide under deadline pressure what to leave out.
Automated vs manual
- Automate what repeats exactly: the same steps, the same data, a clear pass or fail. Sign-in, checkout in test mode, API responses, calculations, permissions. These are regression tests’ natural home, because they run on every change without anyone remembering to.
- Keep manual what needs eyes: visual layout, a new phone or browser, copy that reads wrong, flows that changed so much the old test no longer describes them.
- Watch for flaky tests. A test that fails one run in ten teaches everyone to ignore red. Fix it or quarantine it with a task, never let it stay in the release gate.
- Retire tests for features that are gone. A suite that grows forever gets slower until people stop running it.
The regression testing checklist
Release: [version] Build: [id] Environment: [staging / prod-like] Compared with: [last release tag] Run by: [name] Date: [Month day, year] Result per line: PASS / FAIL (bug ref) / SKIPPED (why) SCOPE [ ] List of tasks in this release attached [ ] Changed files and their direct dependents listed [ ] Areas to re-test agreed from: changes, high-risk flows, past bugs BUILD READY (smoke) [ ] App starts; health check passes; main page loads [ ] Sign in works with a test account AUTOMATED SUITE [ ] Full regression suite run on this build, not an earlier one [ ] Every failure has a bug ref or a known-flaky task ref [ ] Tests for bugs fixed in this release added and passing [ ] Suite run time noted: [minutes] (rising? file a task) HIGH-RISK FLOWS (every release) [ ] Sign up, sign in, sign out, password reset [ ] Core action of the product, start to finish [ ] Payment in test mode: success, decline, refund [ ] Data saved, reloaded and exported correctly [ ] Roles: a restricted user cannot see or change others' data [ ] Emails and notifications arrive; their links work CHANGED AREAS AND NEIGHBORS [ ] [area]: main path and one edge case [ ] [neighbor sharing a component or query]: still correct PAST BUGS [ ] Reopened bugs from the last three releases re-tested [ ] Bugs in changed areas re-tested MANUAL [ ] Layout on the smallest supported screen [ ] Anything the automated suite cannot see (list it) SIGN-OFF [ ] Every FAIL filed as a bug, with blocks / does not block decided [ ] Blocking bugs fixed and re-tested, or release moved
Treat the checklist as a living document with one owner. After each release, add a line for anything that broke in production that the list did not catch, and remove lines for features that no longer exist. Keep the filled copies: a run from three releases ago that shows which lines were skipped, and why, is often the fastest way to explain how a bug slipped through. And note the run time of the automated suite each release; when it creeps up, people start skipping it, so a slow suite is a task in its own right.
Turning fixed bugs into regression tests
The cheapest regression test to write is the one for a bug you just fixed: the steps to reproduce are already in the report, and the expected result is what the fix produces. Write the test before closing the bug, check that it fails on the code before the fix and passes after, and label it with the bug’s ref so a future failure explains itself. Playwright, for example, lets you tag tests (opens in a new tab) with names that start with @ and run only the tagged set with --grep; pytest does the same with custom markers selected by -m.
// tests/checkout.spec.ts
import { test, expect } from '@playwright/test';
test('double click on Pay charges once (BUG-311)', { tag: ['@regression'] }, async ({ page }) => {
await page.goto('/checkout');
await page.getByRole('button', { name: 'Pay' }).dblclick();
await expect(page.getByText('Order confirmed')).toHaveCount(1);
});
// run just the regression set:
// npx playwright test --grep @regressionOver time the tagged tests become a map of where your product has broken before, which is also the best guide to where it will break next.
AI agents running the suite and filing what fails
An AI coding agent is good at the tedious part: running the suite, reading the failures, grouping several failures with one cause, and writing each up with the failing test, the error and the last commit that touched the area. Claude Code can do this unattended: its documentation on running Claude Code programmatically (opens in a new tab) shows claude -p with --allowedTools, which auto-approves only the tools you list, using the same rule syntax as permissions. Allow the test command and the tools to search and file tasks, and nothing that edits code.
claude -p "Run the regression suite with: npx playwright test --grep @regression. For each distinct failure, search the board for an open bug first; if one exists, comment on it with this run's result. Otherwise file a bug with the test name, the error, the steps from the test, and the build. Do not change any code." \ --allowedTools "Bash(npx playwright test *),mcp__fenbs__fenbs_search,mcp__fenbs__fenbs_comment,mcp__fenbs__fenbs_create_item"
The tool names assume the fenbs MCP server is registered under the name fenbs. Keep a person in charge of the sign-off: the agent can say what failed, but whether a failure blocks the release is a priority decision, and it belongs to whoever orders the work.
Regression results on a fenbs board
fenbs does not run tests or store the suite; keep that in your repository and CI. It holds what the run produces. Each failure becomes a task of kind bug with a BUG- ref, the failing test and error in its note, and a priority from 1 to 10, 1 the most urgent, so the release blocker goes in at the top. When an assistant files over MCP, fenbs checks open tasks and those finished in the last 14 days and returns a likely match instead of filing a second copy, and an automated source can pass a key, such as the test name, so the same failure never files twice. When a fix lands, its test status records Tested, Partly tested or Failed with notes on the run, and filtering the board for tasks that are Completed but not tested is the question to ask before every release. If the team agrees a rule such as “every fixed bug gets a regression test before it is closed”, record it on the Decisions and rules page; every connected AI assistant reads the rules first.
Related
The whole pre-release check: QA checklist. Writing one case well: test case template. The states a bug goes through, including reopening: bug life cycle. A board set up for bugs: bug tracker template. Checking an assistant’s output before it counts: verifying AI-generated work.