Unit Testing: A Practical Guide With Examples
A unit test checks one small piece of code, on its own, in milliseconds. What makes a good one, test doubles in plain words, why a coverage number is not the goal, worked examples in Vitest and pytest, and how to review the tests an AI coding assistant writes for you.
7 min read
Unit testing means writing small, automated checks for the smallest pieces of your code, usually one function or one class, and running them on every change. A unit test calls the code with a known input and asserts the output: “a subtotal of 4,999 cents pays the flat shipping fee; 5,000 cents ships free.” Good unit tests are fast enough to run hundreds of times a day, isolated from databases, networks and the clock, focused on one behavior each, and named so that a failure explains itself. They are the base of the test pyramid and the first tests a team should write. Examples in TypeScript and Python are below.
What is unit testing, exactly?
The ISTQB glossary defines unit testing (opens in a new tab) as “a test level that focuses on evaluating the smallest part of code testable in isolation.” What counts as a unit is a team choice. Martin Fowler, in his Unit Test (opens in a new tab) article, notes that object-oriented code often treats a class as the unit and functional code a single function, and that the team decides what makes sense. He also names two styles: solitary tests replace every collaborator with a stand-in, while sociable tests let the unit use its real collaborators as long as they are fast. Both are unit tests; pick one per codebase and be consistent.
Unit tests sit at the bottom of the levels described in types of software testing. Above them, integration tests check the unit with its real database or API, and system tests check the whole product.
What makes a good unit test
- Fast. Milliseconds each. If the suite takes long enough to make you reach for your phone, people stop running it before they push.
- Isolated. No real network, database, file system or current time. The test controls every input, so the same code gives the same result on every machine and every day.
- One behavior. Each test checks one rule. Several assertions are fine if they describe the same behavior; two rules in one test means the name cannot describe both.
- A clear name. Name the rule, not the function: “ships free at exactly the threshold” beats “test shipping 2”. When it fails in CI, the name is the bug report.
- Arrange, act, assert. Set up the inputs, call the code once, check the result. If the setup is 40 lines, the code under test probably has too many dependencies.
- Fails for the right reason. Before trusting a new test, break the code on purpose and watch it go red. A test that cannot fail is not testing anything.
A unit test example in TypeScript with Vitest
The function decides a cart’s shipping fee: free at 5,000 cents or more, a flat 599 cents below that, and an error for a negative or fractional amount. The tests check both sides of the line and the error. Vitest looks for files with .test. or .spec. in the name and runs in watch mode by default; npx vitest run runs once, which is what CI wants.
// shipping.ts
export function shippingFeeCents(subtotalCents: number): number {
if (!Number.isInteger(subtotalCents) || subtotalCents < 0) {
throw new RangeError('subtotal must be a whole number of cents, 0 or more');
}
return subtotalCents >= 5000 ? 0 : 599;
}
// shipping.test.ts
import { describe, expect, test } from 'vitest';
import { shippingFeeCents } from './shipping';
describe('shippingFeeCents', () => {
test('charges the flat fee one cent under the free-shipping line', () => {
expect(shippingFeeCents(4999)).toBe(599);
});
test('ships free at exactly the line', () => {
expect(shippingFeeCents(5000)).toBe(0);
});
test('rejects a negative subtotal', () => {
expect(() => shippingFeeCents(-1)).toThrow(RangeError);
});
});
// run once: npx vitest runThe same test file runs under Jest with one change: Jest provides describe, test and expect as globals, so drop the import from vitest. Jest’s getting-started guide uses the same expect(...).toBe(...) style. On the JVM, JUnit (opens in a new tab) plays the same role with methods marked @Test; its user guide is at version 6 as of October 1, 2026.
A unit test example in Python with pytest
The pytest documentation (opens in a new tab) says it runs files named test_*.py or *_test.py, collects functions that start with test_, and reports intermediate values when a plain assert fails, so you do not need special assertion methods. pytest.mark.parametrize runs one test over a table of inputs, which suits boundary checks.
# shipping.py
def shipping_fee_cents(subtotal_cents: int) -> int:
if subtotal_cents < 0:
raise ValueError("subtotal must be 0 or more")
return 0 if subtotal_cents >= 5000 else 599
# test_shipping.py
import pytest
from shipping import shipping_fee_cents
@pytest.mark.parametrize(
"subtotal, expected",
[(0, 599), (4999, 599), (5000, 0)],
)
def test_fee_around_the_free_shipping_line(subtotal, expected):
assert shipping_fee_cents(subtotal) == expected
def test_negative_subtotal_is_rejected():
with pytest.raises(ValueError):
shipping_fee_cents(-1)
# run: pytest -qTest doubles in plain words
A test double is a stand-in for something your unit depends on, the way a stunt double stands in for an actor. Fowler’s Test Double (opens in a new tab) article credits the term to Gerard Meszaros and lists five kinds:
- Dummy: passed in to fill a parameter, never used.
- Fake: a working but simplified version, such as an in-memory database instead of a real one.
- Stub: returns canned answers, such as a tax service that always says 8%.
- Spy: a stub that also records how it was called, such as an email sender that counts the messages.
- Mock: set up with expectations about the calls it should receive, and fails the test if they do not happen.
Most frameworks blur these names; Vitest calls nearly everything a mock (vi.fn, vi.spyOn, vi.mock) and pytest offers the monkeypatch fixture, which swaps an attribute or environment variable for one test and undoes it afterward. The practical rule: replace what is slow, random or outside your control (the network, the clock, a payment provider), and keep real what is cheap and yours. A test that mocks the code it is supposed to test passes no matter what that code does.
Coverage myths
Coverage tools report which lines your tests executed; vitest run --coverage prints it. Fowler’s Test Coverage (opens in a new tab) article puts the limits plainly: coverage is “a useful tool for finding untested parts of a codebase” but “of little use as a numeric statement of how good your tests are.”
- Myth: 100% coverage means no bugs. A line can run without anything checking its result. Coverage measures what executed, not what was asserted.
- Myth: a coverage target improves quality. Fowler warns that people will hit a target, and high numbers are easy to reach with weak tests. Use the report to find untested risky code, not as a gate.
- Myth: low coverage means bad code. Glue code and generated code may not be worth testing. A payments module at 60% is a problem; a settings screen at 60% may be fine.
Reviewing unit tests written by AI coding assistants
Coding assistants write tests quickly, and the failure modes are predictable. The general review method is in how to review AI-generated code; for tests specifically, check these before you merge.
- Break the code and rerun. Change
>=to>in the function under test. If every test still passes, the tests do not check the rule. - Look for expected values copied from the output. A test that asserts whatever the code currently returns locks in today’s bugs. Expected values should come from the requirement, not the implementation.
- Check what is mocked. If the assistant mocked the function under test, or most of its logic, the test checks the mock.
- Check that existing tests were not weakened. A diff that changes an assertion to match new behavior needs a reason in the task, not just a green run.
- Read the names. Names like “should work correctly” hide what the test claims. Each name should state one rule a person would agree with.
On fenbs, an assistant connected over MCP records how a task was tested in the task itself: a test status (Tested, Partly tested, Failed, or Needs owner check) and test notes saying what ran, where, and what was not checked. fenbs does not run your tests or read CI results, so those notes are only as good as what the assistant actually ran; ask for the command and the pass count in the notes, and filter for Completed tasks that are not tested before a release.
Related
Where unit tests fit among the others: types of software testing. The quick check after each deploy: smoke testing. Running tests on every pull request: GitHub Actions for CI/CD. Checking an assistant’s claims before they count: verifying AI-generated work.