Shift-Left Testing: Catching Bugs Before They Are Bugs
Shift-left testing moves checks to the moment a requirement is written and the moment code is typed, instead of a test phase at the end. What it means, the practices that make it real, what NIST asks for on security, and where AI coding agents fit.
8 min read
Shift-left testing means doing the checking earlier: when the requirement is written, when the design is sketched and when the code is typed, rather than in a test phase after development is “finished”. In practice it is six habits: requirements you can test, tests written with the code, static analysis in the editor and in CI, small changes reviewed quickly, a fast pipeline on every push, and security requirements from the first day. It does not remove later testing. It means the later testing finds fewer surprises, because most mistakes were caught by the person who made them, minutes after making them.
The name comes from a timeline drawn left to right: requirements, design, build, test, release. DORA’s page on pervasive security (opens in a new tab) puts it plainly: the idea “is also known as shifting left, because concerns, including security concerns, are addressed earlier in the software development lifecycle (that is, left in a left-to-right schedule diagram).”
What shift left means, and what it does not
- It means moving a check to the earliest point where it can run. A type error belongs in the editor, not in a QA ticket two weeks later.
- It does not mean developers do all testing and testers go away. Testers move left too: into refinement, writing acceptance criteria and spotting the case nobody thought of before code exists.
- It does not mean skipping end-to-end tests, exploratory testing or a release check. Those still run; they just stop being the first place a bug is seen.
- It is not only about tests. Reviews, linters, threat modeling and a clear definition of done are all shift-left practices.
- It pairs with “shift right”: watching production with monitoring and feature flags. Left catches what can be predicted; right catches what cannot.
Why earlier is cheaper
A mistake found by its author a minute after typing it costs a minute. The same mistake found in a release test costs a bug report, a context switch back into code the developer has forgotten, a fix, a re-test and sometimes a delayed release. DORA describes the worst case: when testing happens only after development is complete, teams discover significant problems, “including architectural flaws, that are expensive to fix.” The cost is not only time. Late bugs arrive in batches, right before a deadline, when the pressure to ship with known issues is highest.
Six shift-left testing practices
- Testable requirements. Write acceptance criteria before the work starts, each one something a person or a test can check. “Fast” is not testable; “the search page returns in under one second for 10,000 products” is. Acceptance criteria examples has twenty to copy.
- Tests written with the code. The change and its tests arrive in the same pull request. Writing the failing test first, test-driven development, is the strictest form, but the rule that matters is “no behavior change without a test”. How to write good ones: unit testing.
- Static analysis where the code is written. Type checks, linters and security scanners run in the editor and again in CI. OWASP’s page on source code analysis tools (opens in a new tab) notes that these tools “can be added into your IDE” to catch issues during development, and is honest about the limits: they struggle with authentication and access control flaws and produce false positives.
- Small changes, reviewed quickly. A 50-line pull request gets a real review; a 2,000-line one gets a skim. Review the design before the code when the change is large. A list of what to look at: code review checklist.
- A fast pipeline on every push. Build, lint, unit and integration tests run on every branch, and a red build blocks the merge. Keep it under ten minutes or people stop waiting for it. The stages: CI/CD pipeline.
- Security from the first day. Security requirements are written with the feature, the design is checked against them, and code is reviewed and tested for vulnerabilities before release. NIST spells this out, below.
Deciding which checks to automate first, and which to leave to a person, is covered in test automation.
Shift-left security: what NIST asks for
The US reference for building security in early is NIST’s Secure Software Development Framework, SP 800-218 (opens in a new tab). Version 1.1 was published in February 2022; a draft of version 1.2 followed on December 17, 2025, and as of October 1, 2026 NIST still lists it as an initial public draft. The framework groups its practices into four areas: Prepare the Organization, Protect the Software, Produce Well-Secured Software, and Respond to Vulnerabilities.
- PO.1, define security requirements for software development, so they are known before anyone designs or builds.
- PW.1, design software to meet security requirements and mitigate security risks: threat modeling, in plain terms, before the code.
- PW.7, review and/or analyze human-readable code to identify vulnerabilities: peer review, static analysis or both.
- PW.8, test executable code to identify vulnerabilities, including adding tests for previously reported vulnerabilities so they cannot come back.
One task under PW.7 is worth quoting for small teams, because it is about where findings go: “record and triage all discovered issues and recommended remediations in the development team’s workflow or issue tracking system.” A scanner warning that lives only in a CI log is not shift left. It is a finding nobody owns.
Where AI coding agents fit
AI coding agents such as Claude Code, Copilot and Cursor write code faster than anyone can review it line by line, so they push in both directions at once. Used well, they shift testing further left: an agent can write the failing test, run the suite and fix its own type errors before you see the diff. Used carelessly, they shift bugs right: a large change that “looks fine” lands untested and the problems surface in review, or in production.
The difference is whether the checks are part of the agent’s job or left to its judgment. Put the commands in the instructions file every session reads, and make “done” mean the checks passed:
## Before you call a change done - Write or update a test that fails without the change, then make it pass. - Run: npm run lint && npm run typecheck && npm test - Run the security scan the project uses (npm run scan) on changed files. - If a step fails, fix it or stop and report the failure. Never call a change done on a red run, and never weaken or delete a test to make it pass. - List anything you could not check, so a person can.
Instructions can be ignored; some tools let you go further and make a check run every time. Anthropic’s guide to Claude Code hooks (opens in a new tab) describes them as shell commands that run at fixed points, so “certain actions always happen rather than relying on the LLM to choose to run them.” A hook on the PostToolUse event with an Edit|Write matcher can run a formatter or linter on every file the agent edits.
An agent that passes its own checks still needs a person to read the change, because it can write a test that confirms its own mistake. What to look for: how to review AI-generated code.
A shift-left checklist for a small team
- Every task has acceptance criteria before work starts, and someone other than the author has read them.
- Every behavior change arrives with a test in the same pull request.
- Type checks and the linter run on save in the editor and on every push in CI.
- A dependency scanner and a static analysis tool run in CI, and their findings are triaged, not muted.
- Pull requests are small enough to review in one sitting; large changes get a design review first.
- A red build blocks the merge. Flaky tests are filed as bugs, not retried until green.
- Security requirements are written with the feature: who may do this, what data it touches, what happens on bad input.
- AI assistants have the test, lint and scan commands in their instructions file, and are told never to call work done on a red run.
- Findings from every check land in one place, with an owner.
What still happens on the right
Shift left changes the order, not the list. End-to-end tests, exploratory sessions, accessibility checks and a pass over the release still run before you ship; they just find less. Where each kind of test sits is mapped in types of software testing, and the pass before a release is the QA checklist.
Recording what the early checks find
NIST’s point about an issue tracking system is the part small teams skip, and it is where a board helps. On a fenbs board, each finding a person or an AI assistant cannot fix on the spot becomes a task of kind bug, with a BUG- ref, a priority from 1 to 10 (1 the most urgent), a note saying what is wrong and where, and a plan for the fix. When the fix is done, the task’s Testing section records how it was checked: Tested, Partly tested, Failed, or Needs owner check for what only a person can confirm.
Two other parts carry the shift-left habits to every assistant that connects. AI context holds notes every assistant reads before it starts, which is the place for the test, lint and scan commands. A team rule such as “every behavior change ships with a test” goes on the Decisions and rules page, where it is recorded as decided by a person and connected assistants read the rules first. fenbs does not run tests and has no CI or GitHub integration; the checks stay in your pipeline, and the board holds what they found.
Related
Writing tasks an assistant can finish and verify: how to write a task for an AI agent. What “done” means across the team: definition of done examples. Filing what the checks find: bug report template. Connecting an assistant to the board: MCP docs.