End-to-End Testing: What It Is and How Much You Need
End-to-end testing drives a whole user journey through the real app, from the first click to the last database write. What it catches, how few of these tests you can get away with, how Playwright, Cypress and Selenium differ, and how to keep the suite from going flaky.
8 min read
End-to-end testing, usually shortened to e2e testing, runs a complete user journey through the real system: a browser opens the app, signs in, adds an item to the cart, checks out, and the test confirms the order landed in the database and the confirmation appeared. It is the only kind of test that proves all the pieces work together the way a customer uses them. It is also the slowest, most expensive and most fragile kind, which is why the answer to “how much do you need” is: a handful of tests for the journeys that would cost you money or trust if they broke, and no more. Everything else belongs lower down, in unit and integration tests.
What end-to-end testing checks
The ISTQB glossary defines end-to-end testing (opens in a new tab) as “A test type in which business processes are tested from start to finish under production-like circumstances.” Two phrases in that carry the weight. “Start to finish” means a whole process, not one screen. “Production-like” means a real browser, the real front end, the real API and a real database, deployed the way production is, with only true outsiders such as a payment provider swapped for their sandbox.
That is what separates it from its neighbors. A unit test checks one function in isolation. An integration test checks two or three pieces together, such as your API and its database. An end-to-end test checks the chain a user actually walks. It is black box by nature: it knows the screens and the expected outcome, not the code (the difference is explained in black box vs white box testing). Where it sits among the other levels and types is on the map in types of software testing.
How much end-to-end testing you need
Less than you think. The test pyramid puts many fast unit tests at the bottom, fewer integration tests in the middle and a small number of end-to-end tests at the top. Google’s testing team, in its post Just Say No to More End-to-End Tests (opens in a new tab), wrote that “As a good first guess, Google often suggests a 70/20/10 split: 70% unit tests, 20% integration tests, and 10% end-to-end tests.” Treat that as a starting shape, not a target: the same post says the exact mix differs by team.
A practical way to size it: list the journeys where a failure would cost real money or real trust, write one end-to-end test per journey, and stop. Every extra rule, field validation or error message gets tested lower down, where it is faster and does not break when a button moves. If you catch yourself writing an end-to-end test to check that a ZIP code field rejects four digits, that is a unit test in the wrong place.
Which journeys to cover: an example
For a small US online store, a first end-to-end suite might be seven tests:
- A new customer signs up and receives the welcome email (checked in a test inbox).
- A returning customer signs in, and a wrong password is refused.
- Search finds a product by name and the product page loads with its price.
- Guest checkout: add to cart, enter a shipping address and ZIP code, see sales tax on the order summary, pay with the payment provider’s sandbox card, see the confirmation.
- A signed-in customer reorders from order history.
- Password reset: request, receive the link, set a new password, sign in with it.
- An admin refunds an order and the customer’s order page shows the refund.
Each one is a business process from start to finish, and each one, broken, is a support queue on Monday. The tax rates, shipping rules and address formats behind journey 4 are tested in unit tests; the end-to-end test only proves the pieces are wired together.
E2E testing tools: Playwright, Cypress and Selenium
Three tools cover most browser end-to-end testing. Each describes itself differently, and the differences matter more than any feature list.
- Playwright: its documentation (opens in a new tab) describes Playwright Test as “an end-to-end test framework for modern web apps” that “bundles test runner, assertions, isolation, parallelization and rich tooling.” It runs Chromium, WebKit and Firefox on Windows, Linux and macOS, with test code in TypeScript or JavaScript, Python, Java or .NET. Before each action it auto-waits for the element to be visible, stable, enabled and able to receive the click, which removes most hand-written waits.
- Cypress: its own overview (opens in a new tab) says Cypress runs “in the same run loop as your application”, inside the browser rather than driving it from outside, and that it “automatically waits for commands and assertions before moving on.” It covers end-to-end and component testing. Its browser list is Chrome-family browsers, Edge and Firefox; WebKit support is labeled experimental, and the bundled Electron browser is now marked deprecated in the docs.
- Selenium: the WebDriver documentation (opens in a new tab) says WebDriver “drives a browser natively, as a user would, either locally or on a remote machine,” and that it is a W3C Recommendation. Selenium is a browser automation library with bindings for many languages; you pair it with a test runner such as JUnit or pytest. It is the long-standing choice for large cross-browser grids and for teams whose tests live in Java or C#.
If you are starting fresh on a JavaScript or TypeScript web app, any of the three will work; pick the one your team will actually maintain, and do not mix two in one suite. A minimal Playwright test for journey 4 looks like this:
import { test, expect } from '@playwright/test';
test('guest checkout shows sales tax and confirms the order', async ({ page }) => {
await page.goto('/products/desk-lamp');
await page.getByRole('button', { name: 'Add to cart' }).click();
await page.getByRole('link', { name: 'Checkout' }).click();
await page.getByLabel('Street address').fill('1 Main St');
await page.getByLabel('City').fill('Boston');
await page.getByLabel('State').selectOption('MA');
await page.getByLabel('ZIP code').fill('02134');
await page.getByRole('button', { name: 'Continue' }).click();
await expect(page.getByText('Sales tax')).toBeVisible();
await page.getByRole('button', { name: 'Place order' }).click();
await expect(page.getByRole('heading', { name: 'Order confirmed' })).toBeVisible();
});Note the locators: roles and labels, the way a person finds a control, not CSS classes that change with every redesign. That one habit prevents more broken tests than any tool choice. How end-to-end tests fit into the rest of an automated suite, and in what order to build it, is in test automation.
Flaky end-to-end tests
A flaky test passes and fails on the same code. End-to-end tests are the worst offenders because they touch everything: network, timing, animations, shared data, third-party sandboxes. Ham Vocke’s The Practical Test Pyramid (opens in a new tab) puts it bluntly: end-to-end tests “are notoriously flaky and often fail for unexpected and unforeseeable reasons.” The usual fixes:
- Wait for a condition, never a fixed time. Assert that “Order confirmed” is visible; do not sleep two seconds and hope.
- Give every test its own data. Create the customer and the cart in the test (through the API, which is faster than the UI), so tests cannot collide or depend on order.
- Stub what you do not own, unless it is the point of the test. A shipping-rate API that is slow on Tuesdays should not fail your checkout test.
- Turn on retries in CI to label flakiness, not to hide it, and keep a trace of the first retry so you can see what happened.
- Quarantine a flaky test the day it appears and file it as a bug. Google’s numbers on flakiness and the full routine are in test automation, so they are not repeated here.
import { defineConfig } from '@playwright/test';
export default defineConfig({
retries: process.env.CI ? 2 : 0,
use: {
baseURL: process.env.BASE_URL,
trace: 'on-first-retry',
},
});Where end-to-end tests run
Run the full end-to-end suite in CI against a production-like environment before a release, and keep it fast enough that people do not skip it. After every deploy, run a smaller subset against production itself: that is a smoke test, and a couple of your end-to-end journeys often double as it. Journeys that once broke belong on the regression testing checklist. The manual part of a release, such as layout, wording and devices, is the QA checklist.
AI agents running end-to-end tests
Coding agents can now drive a real browser, which makes them useful for end-to-end work in two ways. First, reproducing: an agent with Playwright MCP can walk through a reported bug step by step and tell you where it breaks. Second, writing and repairing tests: Playwright ships three test agents, a planner, a generator and a healer that “executes the test suite and automatically repairs failing tests,” in its own documentation’s words. Setup for each client is in the Playwright MCP post above.
Review what they produce like any other change. A healer that makes a red test green may have fixed a locator, or may have changed the assertion so the test no longer checks the thing that broke. Read the diff, and never accept a repaired test whose expected result changed without a reason you agree with.
Turning failures into tasks
fenbs does not run tests and has no CI integration; your runner and pipeline do that. What it holds is what the failures mean. A real failure in the checkout journey becomes a task of kind bug in To Do, with the failing step, the trace and the environment in its note; a flaky test becomes its own bug so it is fixed rather than rerun. When the fix lands, the task’s test status and test notes say which journeys were rerun and where. An AI assistant connected over MCP can file and update those tasks itself, recorded in History as “Claude via” the person who connected it, so the board shows what the agent found without you reading its transcript.
Related
Driving a browser from an agent: Playwright MCP. What to automate first: test automation. Filing a failure well: bug report template. Connecting an assistant to the board: Claude Code integration.