Black Box vs White Box Testing, With Examples
Black box testing designs tests from what the software should do; white box testing designs them from how the code is written. The techniques behind each, where gray box fits, who does which, and one ZIP code function tested all three ways.
8 min read
Black box testing designs tests from the specification: you know what the software should do with an input, and you check that it does, without looking at the code. White box testing designs tests from the code itself: you read the branches and conditions and write tests that make each one run. Gray box testing mixes the two, using some knowledge of the inside (a database schema, an API contract) to choose better outside tests. Black box testing finds missing and misunderstood behavior; white box testing finds code that no test ever ran. Most teams need both, usually from different people: developers write white box unit tests, and testers, product owners and users test black box. Below are the techniques for each, and one ZIP code function tested all three ways.
What is black box testing?
The ISTQB glossary, the vocabulary most testing certifications use, defines black-box testing (opens in a new tab) as “A test approach based on the specification of a component or system.” The tester treats the software as a closed box. The inputs come from the requirements, the user story or the acceptance criteria, and the expected results come from the same place. If the spec says a ZIP code field accepts five digits, the test types five digits and checks the form accepts them; how the check is coded does not matter.
Black box testing works at every level, from a single function called through its public interface to a whole checkout driven through the browser. Most acceptance, system and end-to-end testing is black box by nature, and so is most of what the types of software testing map calls functional testing.
What is white box testing?
ISTQB defines white-box testing (opens in a new tab) as “A test approach based on the internal structure of a component or system.” You open the box: you read the code, see that a function has three if statements, and design tests so that each statement runs and each decision goes both ways. It is also called structural, glass box or clear box testing. Its natural home is the unit test, written by the developer who knows the code, and its usual measure is code coverage.
Black box vs white box testing at a glance
BLACK BOX WHITE BOX
-------------- ------------------------------- -------------------------------
Tests come the spec, story or criteria the code: statements, branches
from
Needs code no yes
knowledge
Typical level system, acceptance, end-to-end unit, some integration
Usually by testers, product owner, users developers
Techniques equivalence partitioning, statement coverage,
boundary values, decision branch coverage
tables, error guessing
Finds missing or misread behavior code no test ever runs,
dead and unreachable code
Misses code paths the spec never features nobody wrote
mentions (no code, so no branch)Black box testing techniques
Three techniques do most of the work, and ISTQB names each one as a black-box test technique.
- Equivalence partitioning: split the possible inputs into groups the software should treat the same way, then test one value from each group. Five-digit ZIP codes are one partition; four-digit strings are another. Testing 90210 and 10001 adds nothing that one of them does not already prove.
- Boundary value analysis: bugs cluster at the edges of partitions, where someone wrote
<instead of<=. ISTQB describes boundary value analysis (opens in a new tab) as “A black-box test technique in which the test conditions are boundary values.” For a five-digit rule, test lengths 4, 5 and 6. - Decision table testing: when the result depends on a combination of conditions, list every combination and its expected action in a table, then test each column. It finds the combination nobody thought about.
- Experience-based techniques sit beside these: error guessing (trying the inputs that broke similar software before) and exploratory testing, where a person designs the next test from what the last one showed.
Worked example: a ZIP code field, tested black box
The spec for a checkout form says: “The ZIP code field accepts a five-digit ZIP code, or ZIP+4 written as five digits, a hyphen and four digits. Spaces before and after are ignored. Anything else shows an error.” From that sentence alone, without the code, the partitions and boundaries give this test list.
INPUT PARTITION / BOUNDARY EXPECTED -------------- ------------------------------ -------- "10001" valid 5-digit valid "10001-2345" valid ZIP+4 valid " 10001 " valid, with spaces valid "1000" length 4 (below boundary) invalid "100010" length 6 (above boundary) invalid "10001-234" ZIP+4 with 3 digits after invalid "10001-23456" ZIP+4 with 5 digits after invalid "1000A" 5 characters, not all digits invalid "10001 2345" space instead of hyphen invalid "" empty invalid
Ten tests, each with a reason. Notice that the list is only as good as the spec: if the spec forgot to say what happens with a nine-digit ZIP typed without the hyphen, no black box test will cover it until someone asks. That question is itself a finding; file it against the story.
A decision table for sales tax at checkout
The same checkout charges sales tax. Suppose the product owner’s rule, simplified for the example and not tax advice, is: charge tax when the ship-to state charges sales tax, the item is taxable, and the buyer has no exemption certificate on file. Three yes-or-no conditions make eight combinations, which collapse to four rules.
R1 R2 R3 R4 State charges sales tax Y Y Y N Item is taxable Y Y N - Buyer has exemption certificate N Y - - ---------------------------------------------------- Charge sales tax yes no no no "-" means the condition does not change the result. One test per column: 4 tests cover all 8 combinations.
The same ZIP code check, tested white box
Now open the box. This is the function behind the field.
export function checkZip(input: string): 'valid' | 'invalid' {
const s = input.trim();
if (s.length === 5) {
return /^\d{5}$/.test(s) ? 'valid' : 'invalid';
}
if (s.length === 10 && s[5] === '-') {
return /^\d{5}-\d{4}$/.test(s) ? 'valid' : 'invalid';
}
return 'invalid';
}Statement coverage asks whether every statement ran. Three inputs reach every return: "10001", "10001-2345" and "1000". That is 100% statement coverage, and it still leaves holes: neither regular expression has been seen to fail. Branch coverage (opens in a new tab), “The coverage of branches in a control flow graph” in ISTQB’s words, asks whether every decision went both ways. Add "1000A" (five characters, regex fails) and "10001-234A" (ten characters with a hyphen, regex fails), and each if and each ? : has been taken in both directions: five tests, 100% branch coverage.
You do not count this by hand. A coverage tool reports it while the unit tests run; with Vitest, npx vitest run --coverage prints statements, branches, functions and lines per file, and the coverage configuration (opens in a new tab) lets you set minimum thresholds for each so the build fails if they drop. The white box view also shows which black box cases do real work. Only "10001 2345" reaches the path where the length is ten but the sixth character is not a hyphen; the five branch tests above never try it, so keep it even though it adds no branch. How to write the tests themselves is covered in unit testing.
One warning about coverage: it tells you what ran, not what was checked. A test that calls checkZip("1000A") and asserts nothing still counts toward 100%. Read the assertions.
Gray box testing (“gray box” testing)
ISTQB spells it “grey-box testing” and, in its glossary entry (opens in a new tab), defines it as “A test approach that combines elements of black-box testing and white-box testing.” US teams usually write gray box; it is the same thing. The tester drives the software from the outside, like a user, but chooses inputs using knowledge of the inside: the database schema, the API contract, the logs, an architecture diagram.
Example: the tester knows the orders table stores ZIP codes in a numeric column. Nothing in the spec mentions that, so a pure black box tester has no reason to try a ZIP code with a leading zero. The gray box tester enters 02134, places the order, and checks the confirmation email and the admin screen. If they show 2134, the numeric column dropped the zero, and every customer with a ZIP code that starts with 0 has a broken shipping label. Most integration testing and security testing is gray box in this sense.
Who does which
- Developers: white box unit tests on their own code, with branch coverage on the logic that matters most (money, dates, permissions, validation). They also run gray box integration tests against a real database.
- Testers and QA: black box tests designed from the spec with partitions, boundaries and decision tables, plus exploratory sessions. Gray box when they can read the schema or the logs.
- Product owners and users: black box acceptance testing against the criteria they agreed.
- Automation in CI: both kinds, every change. Anything that once broke becomes a test on the regression testing checklist.
- AI coding assistants: they can write either kind quickly, but a white box test generated from the code tends to restate the code. Give the assistant the spec and the boundary list, not just the function, and read every assertion.
Recording the result on a board
fenbs does not run tests or read coverage reports; your test runner does that. What it holds is the record on each task. Every task has a test status (Not tested, Tested, Partly tested, Failed, or Needs owner check) and test notes, so the ZIP code task can say “Black box: 10 cases from the partition table, all pass. White box: branch coverage 100% on zip.ts. Not tested: international addresses.” The leading-zero defect becomes its own task of kind bug, and an AI assistant connected over MCP sets the same fields as testStatus and testNotes when it finishes work. If your team decides that “every validation function ships with boundary tests”, put that on the Decisions and rules page so every connected assistant reads it first.
Related
The full map of test levels and types: types of software testing. Writing each case down: test case template. Testing whole user journeys from the outside: end-to-end testing. Filing what you find: bug report template and the bug tracker template.