Integration Testing: What It Catches That Unit Tests Miss

Unit tests check each piece alone; integration tests check the seams between them, where most surprises live. What they catch, the approaches from big-bang to contract testing, a real-database example with Testcontainers, and how to let an AI agent run them.

7 min read

Integration testing checks that pieces built and tested separately work together: your code and its database, one service and another, your app and a payment provider’s sandbox. Unit tests replace those neighbors with test doubles, so they can only confirm what you assumed about them. Integration tests replace the assumption with the real thing, which is why they catch the bugs that live at the seams: queries that fail on a real database, fields that serialize differently than expected, and configuration that was never wired up.

What integration tests catch that unit tests miss

The ISTQB glossary (opens in a new tab) defines integration testing as “A test level that focuses on interactions between components or systems.” The interactions are where these failures hide:

  • Database behavior: a unique constraint, a cascade delete, case-sensitive collation, a time zone conversion, or a migration that was never applied.
  • Serialization: orderId on one side and OrderID on the other, dates in two formats, a missing field vs a field set to null.
  • Wiring and configuration: a service never registered, an environment variable read under the wrong name, a URL pointing at the wrong place.
  • Transactions and concurrency: two requests for the same order arriving together, or a write that half succeeds.
  • Authentication between services: a token with the wrong audience or scope, rejected only by the real service.
  • Third-party behavior: how a sandbox payment provider answers a declined card, or how an API paginates past the first page.
  • Mocks that drifted: a test double that still returns last year’s response shape while the real service moved on.

None of these is a reason to write fewer unit tests. Unit tests are fast and pinpoint the broken line; unit testing covers writing good ones. Integration tests are slower and broader, and they answer a different question.

Narrow, broad and system integration testing

Martin Fowler notes in Integration Test (opens in a new tab) that the term has become blurred, and separates two meanings. Narrow integration tests exercise only the code in your service that talks to a separate service, often against a test double or a local copy of it. Broad integration tests need live versions of all the services and exercise whole paths through them.

The ISTQB vocabulary draws a similar line: component integration testing checks components within one system, and system integration testing focuses on the integration of systems, such as your app with an external payment, shipping or identity provider. Most teams want many narrow tests and a few broad ones. Where integration sits among the other levels is mapped in types of software testing.

Approaches: big-bang, top-down, bottom-up and contract tests

  • Big-bang: build every component, combine them all at once, and test the whole. Simple to organize, but when something fails you have no idea which seam caused it, and you find out late.
  • Top-down: start from the top layer, such as the API handlers, and replace the layers below with stubs. Swap each stub for the real component as it becomes ready.
  • Bottom-up: start from the lowest layer, such as the data access code against a real database, and drive it with small test harnesses (drivers). Add the layers above one at a time.
  • Sandwich or hybrid: work from both ends toward the middle. Common in practice, even when nobody calls it that.
  • Contract testing: check each side of an integration separately against a shared contract, instead of running both together. Covered next.

Incremental approaches win for most teams for one reason: when a test fails, only one new seam was added, so the cause is obvious. Continuous integration is the incremental approach taken to its limit, with every merged change tested against everything else.

Contract testing with Pact

When two services are owned by different teams, running both together for every test is slow and fragile. The Pact documentation (opens in a new tab) describes contract testing as “a technique for testing an integration point by checking each application in isolation to ensure the messages it sends or receives conform to a shared understanding that is documented in a ‘contract’.”

  1. The consumer, the side that receives data, writes tests against a Pact mock of the provider. Running them records every request it makes and every response it expects into a contract file.
  2. The contract is shared with the provider, the side that supplies the data.
  3. The provider replays the recorded requests against its real code and checks that its responses match.
  4. If the provider wants to remove a field, the verification fails while the change is still on a branch, not after the consumer breaks in production.

Pact calls itself consumer-driven: the contract contains only what consumers actually use, so the provider is free to change everything else. Contract tests do not replace a few broad tests of the real path, but they move most of the checking to fast, independent runs.

Real databases with Testcontainers

The most valuable integration test in many apps is the code against its real database engine, not an in-memory substitute. Testcontainers (opens in a new tab) is an open source library that provides “throwaway, lightweight instances of databases, message brokers, web browsers, or just about anything that can run in a Docker container.” It has implementations for Java, Go, .NET, Node.js, Python and other languages, and needs Docker or a compatible container runtime.

The example below, following the pattern in the Testcontainers for Node.js documentation (opens in a new tab), starts a real PostgreSQL container, runs the same migrations production runs, and checks something a mocked repository could never prove: that a retried checkout does not create a second order.

orders.integration.test.ts (Vitest)
import { beforeAll, afterAll, test, expect } from 'vitest';
import { PostgreSqlContainer, StartedPostgreSqlContainer } from '@testcontainers/postgresql';
import { Client } from 'pg';
import { migrate, saveOrder } from '../src/orders';

let container: StartedPostgreSqlContainer;
let db: Client;

beforeAll(async () => {
  container = await new PostgreSqlContainer('postgres:17-alpine').start();
  db = new Client({ connectionString: container.getConnectionUri() });
  await db.connect();
  await migrate(db); // the same migrations production runs
}, 120_000);

afterAll(async () => {
  await db.end();
  await container.stop();
});

test('a retried checkout does not create a second order', async () => {
  const order = { orderId: 'A-1001', zip: '30301', totalCents: 4599 };
  await saveOrder(db, order);
  await saveOrder(db, order); // the browser retried

  const { rows } = await db.query(
    'SELECT count(*)::int AS n FROM orders WHERE order_id = $1',
    ['A-1001'],
  );
  expect(rows[0].n).toBe(1);
});

If saveOrder relies on a unique index that a migration forgot to create, a unit test with a mocked repository passes and this test fails. That gap is the whole argument for integration testing.

Keeping integration tests fast and trustworthy

  • Each test creates the data it needs and does not depend on another test’s leftovers.
  • Run the real migrations, not a hand-written schema that drifts from production.
  • Start containers once per test file or suite, not once per test.
  • Wait for readiness instead of sleeping for a fixed number of seconds.
  • Point third-party calls at the vendor’s sandbox, never at production accounts.
  • Run the suite on every pull request; what a CI/CD pipeline is shows where it sits among the stages.
  • Fix or quarantine a flaky test the week it appears. A suite people rerun until it passes has stopped testing anything.
  • After a deploy, a short smoke test checks the live system; it does not replace this suite.

AI agents running integration tests

Coding agents such as Claude Code, Codex CLI and Cursor’s agent can run your test commands in a terminal, read the failures and try a fix. Integration tests are where that helps most and where it needs the clearest instructions, because the failures are often about the environment rather than the code.

  • Put the exact command in the agent’s instructions file, such as npm run test:integration, along with what it needs: Docker running, which environment file to use.
  • Say that it may not change an assertion, delete a test or skip a suite to make a run pass without asking first.
  • Ask it to report what ran, what passed, what failed and what it could not run, such as a payment sandbox it has no credentials for.
  • Never give it production credentials so that a test can “just work.”

fenbs is a simple task board shared by people and AI assistants. It does not run tests and has no CI integration; your test runner and pipeline do that. What it keeps is the record on each task: a test status (Not tested, Tested, Partly tested, Failed, or Needs owner check) and test notes. An assistant connected over MCP sets them as testStatus and testNotes, for example Partly tested with “integration suite 42/42 against Postgres 17 in Testcontainers; payment sandbox not exercised.” A rule such as “a task that touches the database is not Completed until the integration suite passes” belongs on the Decisions and rules page, which every connected AI assistant reads first.

Related

The full map of test levels: types of software testing. The level below this one: unit testing. The check after a deploy: smoke testing. Where the suite runs: CI/CD pipeline. Deciding when a dependency change is breaking: semantic versioning.

Questions people ask.

What is the difference between unit testing and integration testing?

A unit test checks one piece of code on its own, with its dependencies replaced by test doubles. An integration test checks pieces working together, such as your code with a real database or another service, so it catches wrong assumptions about those neighbors that a unit test cannot see.

What is system integration testing?

Integration testing between whole systems rather than components inside one system, for example your application with an external payment, shipping or identity provider. The ISTQB glossary defines it as a test level that focuses on the integration of systems.

Should integration tests use mocks?

For the seam under test, no: the point is to use the real database or service, or a contract checked against the real one. Mocking other, unrelated dependencies to keep a test focused is fine.

How many integration tests do I need?

Enough to cover every seam that has bitten you or would hurt if it broke: each database table your code writes to, each external API, each message between services. Keep them fewer and broader than unit tests, and make every one of them reliable.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.