Test Automation: What to Automate First and What to Leave Manual
Automate the checks that repeat exactly and block a release, keep people on the work that needs judgment, and treat a flaky test as a bug. An order to automate in, a starter checklist, and how to handle tests an AI assistant wrote.
7 min read
Test automation pays off fastest on checks that repeat exactly, have a clear pass or fail, and would stop a release if they broke: the smoke test after a deploy, bugs that have come back before, and logic such as prices, permissions and validation. Automate those first, mostly as fast unit and integration tests, with only a handful of end-to-end journeys through the browser. Leave people the work that needs judgment: exploratory sessions, layout and wording, a new device, and features still changing every week. Run the suite on every change in CI, fix or quarantine flaky tests the day they appear, and read every test an AI assistant writes before you trust it.
What test automation covers
Test automation, automated testing and automation testing are three names for the same thing: code that runs your software, checks the result and reports pass or fail without a person clicking through. QA automation usually means the same work done by a quality team rather than the developers. The kinds of test you can automate, from unit to acceptance, are mapped in types of software testing; this page is about the order to automate them in and how to keep the suite trustworthy.
What to automate first
Start where a missed failure costs the most and the check is cheapest to write. This order works for most small teams:
- The deploy check. Five to ten steps that prove the build is alive: the home page loads, sign-in works, the main screen shows data. Smoke testing has the checklist and a Playwright example.
- Bugs that came back. Every fixed bug gets a test that would have caught it, so it cannot return silently. The routine is in the regression testing checklist.
- Pure logic. Calculations, sales tax, discounts, permission rules, date handling and input validation. These are fast and stable to test in isolation; unit testing shows how.
- The edges of your system. Your code talking to the database, a payment provider or another service. Integration testing covers contract tests and real databases in containers.
- A few critical journeys end to end. Sign up, add to cart, check out in test mode. Keep this list short: these tests are the slowest and break most often.
What to leave manual
Manual vs automated testing is not a contest; each catches what the other misses. DORA’s guidance on test automation (opens in a new tab) says teams should still “perform manual test activities such as exploratory testing, usability testing, and acceptance testing throughout the delivery process.” Keep people on:
- Exploratory sessions: an hour spent trying to break a new feature finds the cases nobody wrote down.
- Layout, wording and feel: a screen can pass every assertion and still read wrong.
- New devices and browsers you have not seen the product on yet.
- Features still changing every week: a test written today describes a screen that will be gone on Friday.
- One-off checks: a data migration you will run once is cheaper to verify by hand than to script.
When a manual session finds a bug, fix it and then automate the check. Over time the manual list stays short and the automated suite grows from real failures rather than guesses.
The test pyramid in two rules
Ham Vocke’s practical test pyramid (opens in a new tab) reduces the idea to two rules: “Write tests with different granularity” and “The more high-level you get the fewer tests you should have.” Many unit tests at the base, fewer integration tests in the middle, a handful of end-to-end tests at the top. A suite built mostly from browser tests is slow, brittle and expensive to keep green, which is why the order above starts low and ends with a short end-to-end list.
A starter checklist for the first month
If you have no automated tests today, this is a month of work a small team can actually finish:
- Week 1: pick one test runner per language and get one passing test running in CI on every push. The tooling matters more than the test.
- Week 1: write the smoke test and run it after every deploy.
- Week 2: write a regression test for each of the last five bugs you fixed.
- Week 2: unit-test the three functions that would cost you money if they were wrong.
- Week 3: add one integration test per external boundary: database, payments, email.
- Week 4: add three end-to-end journeys, no more, and set a time budget for the whole suite.
- Every week: one hour of exploratory testing on whatever changed most, with each bug found becoming a test.
Tools by layer
Choose by language and layer, not by a ranking. Common choices, all widely used:
- Unit tests: Vitest or Jest for JavaScript and TypeScript, pytest for Python, JUnit for Java, xUnit or NUnit for .NET.
- Integration tests: the same runners, plus Testcontainers for real databases and Pact for contract tests between services.
- Browser tests: Playwright, Cypress or Selenium WebDriver.
- Mobile tests: Appium, Espresso for Android, XCUITest for iOS, or Maestro for flows written in YAML.
- API tests: your unit runner calling the API, or a collection runner such as Postman’s Newman in CI.
Run it in CI, and keep it fast
A test suite nobody runs is documentation. Run it on every change, in the order cheapest first, so a failure shows up in minutes. DORA’s target is that developers “should be able to get feedback from automated tests in less than ten minutes both on local workstations and from the continuous integration system.” Where the stages sit in a pipeline is covered in what a CI/CD pipeline is.
{
"scripts": {
"lint": "eslint .",
"typecheck": "tsc --noEmit",
"test:unit": "vitest run",
"test:integration": "vitest run --config vitest.integration.config.ts",
"test:e2e": "playwright test",
"ci": "npm run lint && npm run typecheck && npm run test:unit && npm run test:integration && npm run test:e2e"
}
}Flaky tests: treat each one as a bug
A flaky test passes and fails on the same code. It is the fastest way to teach a team to ignore red. Google’s testing team reported in 2016 that about 1.5% of all test runs (opens in a new tab) across its corpus returned a flaky result, that almost 16% of its tests had some flakiness, and that about 84% of the pass-to-fail transitions its CI saw involved a flaky test. The common causes are timing, shared state between tests, test order, real clocks and networks, and third-party services.
- Find them: Playwright’s retries documentation (opens in a new tab) sorts results into “passed”, “flaky” (failed, then passed on retry) and “failed”. Retries are off by default; turn them on in CI to label flakiness, not to hide it.
- Quarantine the same day: move the test out of the release gate and file it as a bug with the failure output attached.
- Fix the cause: wait for a condition instead of a fixed time, give each test its own data, freeze the clock, stub the network.
- Delete what you will not fix. A test nobody trusts costs more than no test.
Tests written by AI assistants
Coding assistants are quick at tests, and quick at the wrong ones. GitHub’s guide to writing tests with Copilot (opens in a new tab) says “The tests that Copilot generates may not cover all scenarios, so you should always review the generated code and add any additional tests that may be necessary.” A generated test often restates the code it is testing, so it passes whether the code is right or not.
- Give it the expected behavior, not the code: input and output pairs, the edge cases, and the rule they follow.
- Ask for the failing test first, run it, and see it fail for the right reason before any fix is written.
- Read the assertions. A test that asserts a function was called, or that the output equals what the code produces today, proves little.
- Break the code on purpose. If the new test still passes, it is not testing anything.
- Never accept a change that edits a test to make it pass unless the test itself was wrong, and say so in the review.
Cursor’s guide to reviewing and testing (opens in a new tab) puts the limit plainly: “Passing tests don’t guarantee the code works correctly. It’s possible the tests are checking the wrong behavior.” The review routine for the rest of an AI-written change is in reviewing AI-generated code.
Keeping the testing work on one board
Test automation creates work of its own: flaky tests to fix, gaps to fill, manual checks nobody has scripted yet. On fenbs each of those is a task in To Do, Next Up, In Progress or Completed, filed as a bug, feature or enhancement with a priority from 1 to 10. Each task carries a test status (untested, tested, partly, failed, or needs owner check for what only a person can confirm) and test notes saying what was run and what was not. An AI assistant connected over MCP can file a quarantined test with fenbs_create_item, and record the result with fenbs_update_item when it finishes. fenbs has no CI integration, so it does not read your pipeline; the assistant or a person writes the result in.
Related
Setting the release bar: QA checklist. Writing test cases people can follow: test case template. A ready-made board for bugs: the bug tracker template. Connecting an assistant: the MCP docs.