Smoke Testing vs Sanity Testing: What Each Checks and When

A smoke test asks whether a build or deploy is alive at all; a sanity test asks whether one change works well enough to keep testing. Plain definitions, where sources disagree, a smoke checklist to run after every deploy, who owns it, and how an AI agent can run it with Playwright.

7 min read

Smoke testing is a short, broad, shallow check that a new build or deploy works at all: the app starts, the home page loads, a test user can sign in, and the one action the product exists for completes. It takes minutes, runs after every deploy, and its only job is to say “this is worth testing further” or “roll back now.” Sanity testing, as most teams use the word, is narrower: after a small fix, check that the fixed area and the things right next to it behave sensibly before spending time on a full regression run. Smoke goes wide and shallow; sanity goes narrow. The checklist to run after every deploy is below.

What a smoke test is

The name comes from hardware: power the device on and see whether smoke comes out. The ISTQB glossary defines smoke testing (opens in a new tab) as “a test type to gain sufficient confidence that a test object is ready for planned testing.” In practice that means a fixed, small set of checks covering the paths whose failure would make everything else pointless: if sign-in is broken, there is no reason to run 400 regression tests that all start by signing in.

A smoke test is not a quality gate for the release. It does not look for subtle bugs, edge cases or layout problems; that is the job of the regression testing checklist and the QA checklist. It answers one yes-or-no question quickly, and it should always be green. A smoke test that fails sometimes for no reason is worse than none, because people learn to ignore it.

What sanity testing is

Sanity testing is a quick, focused check after a minor change, usually a bug fix or a small configuration change: does the fix work, and does the screen it lives on still behave? It is often done by hand, by the person who made the change or a tester, and it decides whether the build goes on to deeper testing. Where a smoke test is the same list every time, a sanity check is chosen for the change in front of you.

Smoke vs sanity testing: where the sources disagree

The two terms are not used consistently, so check what a person means before arguing about it.

  • The current ISTQB glossary has an entry for smoke testing but none for sanity testing. Its older version 3.01 listed “sanity test” and “confidence test” as synonyms of smoke test, so a team trained on that glossary may use the two words for the same thing.
  • Common industry usage, followed in this guide, splits them: smoke is broad and shallow on every new build, sanity is narrow and a bit deeper on one change. This split is a convention, not a standard.
  • Load testing tools use “smoke test” differently again. Grafana k6’s guide to load test types (opens in a new tab) calls a smoke test a run with minimal load to validate that the script works and the system performs adequately, before any bigger load test. Same name, different job; see performance and load testing.

The practical answer: name your lists by what they do. “Post-deploy smoke” and “fix check” leave less room for argument than “smoke” and “sanity.”

A smoke test checklist for every deploy

Keep it to what a person could check in five minutes, and automate it so nobody has to. Run it against the environment you just deployed to, staging first and then production, using a test account that exists only for this.

Post-deploy smoke test checklist
Deploy: [version / commit]   Environment: [staging / production]
Run by: [pipeline / name]    Time: [Month day, year, hh:mm]
Result per line: PASS / FAIL. Any FAIL = stop, investigate, consider rollback.

ALIVE
[ ] Health endpoint returns 200 and reports the new version
[ ] Home page loads over HTTPS with no server error
[ ] No spike in error rate in monitoring since the deploy

ACCESS
[ ] Test account can sign in and sign out
[ ] Password reset email is requested without an error

CORE ACTION (the one your product exists for)
[ ] Create the main thing (order, booking, task, post)
[ ] It saves, reloads and shows the same data
[ ] Payment in test mode succeeds (if you take payments)

WIRING
[ ] Main API call from the front end returns data, not an error
[ ] Background jobs / queues are processing (if any)
[ ] The area this deploy changed opens without an error

AFTER
[ ] Test data from this run cleaned up or clearly marked
[ ] Result posted where the team will see it

Who runs the smoke test

  • The pipeline, first. The verify stage of a CI/CD pipeline runs the automated smoke suite straight after each deploy and can roll back on a failure.
  • The person who deployed, second. They own the result until it is green, including the checks that cannot be automated, such as a real email arriving.
  • One named owner for the list. Someone keeps it short and current: add a line when production broke in a way the smoke test should have caught, remove lines for features that are gone.

When the smoke test fails

  1. Stop the rollout. If the deploy goes to servers or regions in stages, do not send it to the next one.
  2. Check whether it is the test or the app. Run the failing check once more by hand. If it passes by hand and fails in the script, the test is flaky, and that is a task of its own; it does not wave the deploy through.
  3. Roll back or switch off. If the core action is broken, put the previous version back, or turn off the feature flag for the change, before anyone debugs. Debugging comes after users are safe.
  4. Write it down. File the failure as a bug with the deploy version and the failing line, and if production broke in a way the checklist did not catch, add a line to the checklist.

An automated smoke suite in Playwright

Playwright lets you tag tests (opens in a new tab) with names that start with @ and run only those with --grep, so the smoke suite can live beside your other end-to-end tests. Set baseURL from an environment variable in the config, and page.goto('/') goes to whichever environment you just deployed.

tests/smoke.spec.ts
import { test, expect } from '@playwright/test';

// playwright.config.ts sets: use: { baseURL: process.env.BASE_URL }

test('health check reports OK', { tag: '@smoke' }, async ({ request }) => {
  const res = await request.get('/api/health');
  await expect(res).toBeOK();
});

test('test user can sign in', { tag: '@smoke' }, async ({ page }) => {
  await page.goto('/signin');
  await page.getByLabel('Email').fill(process.env.SMOKE_EMAIL!);
  await page.getByLabel('Password').fill(process.env.SMOKE_PASSWORD!);
  await page.getByRole('button', { name: 'Sign in' }).click();
  await expect(page.getByRole('heading', { name: 'Dashboard' })).toBeVisible();
});

// after a deploy:
//   BASE_URL=https://staging.example.com npx playwright test --grep @smoke

Keep the credentials in your CI secret store, never in the test file. Give the smoke account the least access that lets it run the checks, and if a check creates data in production, make it easy to find and remove.

AI agents running the smoke test

An AI coding agent can run the suite after a deploy, read the failures, and write them up. With Claude Code you can do that unattended: its documentation on running Claude Code programmatically (opens in a new tab) shows claude -p with --allowedTools, which auto-approves only the tools you list. Allow the test command and the board tools, and nothing that edits code or deploys.

Terminal: run the smoke suite and report
claude -p "Run: npx playwright test --grep @smoke against $BASE_URL.
If everything passes, say so in one line. For each failure, search the board for an
open bug first and comment on it if one exists; otherwise file a bug with the test name,
the error, the environment and the deploy version. Do not change code or redeploy." \
  --allowedTools "Bash(npx playwright test *),mcp__fenbs__fenbs_search,mcp__fenbs__fenbs_comment,mcp__fenbs__fenbs_create_item"

For checks that are hard to script, an agent can drive a real browser through Playwright MCP and click through the checklist itself; Playwright MCP covers setup and how to keep that browser safe. Whatever runs the checks, a person decides on a rollback. The agent reports; it does not get to decide that a broken checkout is acceptable.

On fenbs, each failure becomes a task of kind bug with a BUG- ref and a priority from 1 to 10, 1 the most urgent, so a broken sign-in goes in at the top. fenbs does not run the suite, receive CI results or track test runs; the board holds what the run found, the History page records who filed and moved each task, and the fix carries its own test status and notes once someone re-checks it.

Related

The other kinds of testing and where smoke fits: types of software testing. Small fast checks under the smoke test: unit testing. Shipping a change behind a switch so a failed smoke test is cheap: feature flags. Writing up what failed: bug report template.

Questions people ask.

What is smoke testing in simple terms?

A smoke test is a few minutes of broad, shallow checks that a new build or deploy works at all: it starts, the main page loads, a test user can sign in, and the core action completes. If it fails, there is no point testing further.

What is the difference between smoke and sanity testing?

In common usage, smoke testing is a fixed, broad check of the whole build after every deploy, while sanity testing is a narrow check of one area after a small change or fix. Some sources, including older versions of the ISTQB glossary, treat the two terms as synonyms.

Should smoke tests run in production?

Yes, a short smoke test right after a production deploy is common, using a dedicated test account with minimal access and test-mode payments. Keep production checks read-mostly and clean up any data they create.

Is a smoke test the same as a regression test?

No. A smoke test checks in minutes that a build is alive. Regression testing is the broader, planned re-testing that follows, confirming that existing features still work after a change.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.