What Is Spec-Driven Development? A Practical Guide
Spec-driven development means writing down what you want and why before an AI agent writes any code, then planning, splitting into tasks and building against that document. What a good spec holds, the loop, the tools that formalise it, and when it is more ceremony than the job needs.
7 min read
Spec-driven development is a way of working with AI coding agents in which a written specification comes first and everything else is checked against it. You describe what should be built and why, agree a technical plan, break the plan into small tasks, and only then let the agent implement, one task at a time. The spec is not a throwaway prompt: it is a file that stays in the project, gets reviewed like code, and is the thing you point at when the result is wrong. The term is used loosely, and tools draw the steps slightly differently, but they share that order: what, then how, then the work.
A definition, and where the idea comes from
Writing requirements before code is not new. What is new is who reads them. When the implementer is an agent that will do exactly what the context says, the quality of the written intent decides the quality of the result, and a conversation that scrolls away is a poor place to keep it. GitHub’s Spec Kit (opens in a new tab) puts the method in one line: define what and why before deciding how to build it, then turn the requirements into a specification, a technical plan and actionable tasks, and guide implementation against those artifacts.
Usage varies. Some people mean a full product requirements document; others mean a one-page note with acceptance criteria. The common thread is that the document is the source of truth for the agent, not the chat history, and that a person approves it before any code is written. It is the opposite end of the scale from vibe coding, and close kin to the habits in context engineering vs vibe coding.
The loop: spec, plan, tasks, implement
- Spec. What the feature does, for whom, and how you will know it works. No frameworks, no file names. A person writes it or edits the agent’s draft, and signs it off.
- Plan. How it will be built: the stack, the data model, the parts of the codebase touched, the risks. The agent usually drafts this from the spec plus the repository; a person reviews it.
- Tasks. The plan cut into small pieces, each one finishable and checkable on its own, in an order that respects dependencies.
- Implement. The agent works through the tasks, one at a time, and each result is checked against the spec’s acceptance criteria, not against how plausible the code looks.
The loop runs backwards too. When implementation shows the plan was wrong, you change the plan; when it shows the spec was wrong, you change the spec, then regenerate what depends on it. That is the discipline that separates this from writing a long prompt once: the documents are kept true while the work runs.
What a good spec contains
- The problem and who has it, in a sentence or two.
- What is in scope and, just as important, what is out.
- User-visible behaviour, as short stories or rules.
- Acceptance criteria a stranger could check: inputs, outputs, error cases.
- Constraints that are not negotiable: performance, privacy, compatibility, things that must not change.
- Open questions, marked as open, so nobody builds on a guess.
Acceptance criteria are where most specs are weak. Kiro’s feature specs (opens in a new tab) write requirements in EARS notation, a fixed pattern of “WHEN [condition or event] THE SYSTEM SHALL [expected behaviour]”, which forces each requirement to name a trigger and a checkable outcome. You do not need the tool to borrow the pattern.
# Export tasks to CSV ## Problem Team leads copy tasks into spreadsheets by hand for monthly reports. ## In scope - Export the tasks visible in the current filter as one CSV file. ## Out of scope - Scheduled exports, other formats, attachments. ## Acceptance criteria - WHEN a user presses Export THE SYSTEM SHALL download a CSV of the tasks in the current filter, one row per task, with a header row. - WHEN a title contains a comma or quote THE SYSTEM SHALL escape it so the file opens correctly in a spreadsheet. - WHEN the filter matches no tasks THE SYSTEM SHALL say so and not download an empty file. ## Constraints - No new permissions: a user exports only what they can already see. ## Open questions - Include comments? (Owner to decide.)
That is enough. It says nothing about which library parses CSV or where the button goes in the code; those belong in the plan. The breakdown into tasks is its own skill, covered in AI agent task decomposition, and what makes each task finishable is in how to write a task for an AI agent.
When it pays off, and when it is overkill
It pays off when the work outlasts one session, when more than one person or agent will touch it, when a mistake is expensive to undo, or when the person who wants the feature is not the person reviewing the code. In those cases the spec is what lets a reviewer say “this does not meet criterion three” instead of “this feels wrong”.
It is overkill for a one-line fix, a spike you will throw away, or a bug whose report already says what correct looks like. A four-document ritual around a typo is ceremony, and ceremony teaches people to skip the process when it matters. A sensible rule: if you could write the acceptance criteria in the task itself, the task is the spec.
The tools that formalise it
- GitHub Spec Kit: an open-source toolkit of agent commands for Copilot, Claude Code and other agents. A constitution of project principles is set once; each feature then runs specify, plan, tasks, implement and converge, and the command reference (opens in a new tab) adds optional clarify, checklist and analyze steps, writing
spec.md,plan.mdandtasks.md. - Kiro: an AI-powered development environment whose specs (opens in a new tab) are three files,
requirements.md(orbugfix.md),design.mdandtasks.md, produced in phases with approval between them, or all at once with its Quick Spec option. - Taskmaster: turns a requirements document into a dependency-ordered task file an agent works through; see Taskmaster and Claude Code.
Spec Kit installs with uv and needs Python 3.11 or later. Its README uses Copilot in the examples and tells you to swap in your agent’s integration key; for Claude Code that key is claude, which installs the commands as skills in .claude/skills.
uv tool install specify-cli specify init my-project --integration claude # in Claude Code, one at a time, reviewing each result /speckit-constitution Principles: tests for every change, no new dependencies without asking. /speckit-specify Export the tasks in the current filter to CSV. (paste the spec) /speckit-plan Use the existing export service; no new libraries. /speckit-tasks /speckit-implement /speckit-converge
Two details from Spec Kit’s reference are worth copying even without the tool. Its clarify step asks up to five targeted questions about under-specified areas and writes the answers back into the spec. Its converge step checks the code against the spec, plan and tasks after implementation and appends any gaps as new tasks; you repeat implement and converge until it reports converged.
How it relates to plan modes
Plan modes, such as Claude Code’s plan mode or Copilot’s Plan, make an agent propose before it edits. That is the plan step of the loop, inside one session. Spec-driven development adds the step before it and the record after it. Plan mode does not ask what the feature is for or how it will be accepted; it takes your prompt as the spec. And its plan is written for the session that made it.
They combine well. Write the spec first, then open plan mode with the spec as the prompt, so the plan answers to written criteria rather than to a sentence you typed from memory.
Where the tasks live
A tasks.md in the repository is right for the agent that is implementing. It is less good for the people who decide what gets built next and check what came out, who are often not in the repository at all. The split that works is by level: the spec and its task file stay in the repo, and each outcome someone would check becomes one task on a shared board.
On fenbs each task has a Problem box and a Plan box, which map onto the spec and the plan: the note says what is wrong or missing and where, written once; the plan says how it will be done and is rewritten as the agent learns. An assistant connected over MCP files the tasks with fenbs_create_item, moves each through To Do, Next Up and In Progress, and records how it was tested. A person moves it to Completed after checking it against the acceptance criteria. When the owner settles an open question from the spec, it goes on the Decisions page with who decided and why, so the next session does not reopen it.
Related
Connect an agent to the board: Claude Code integration and the MCP tool reference. How a card, its plan and its testing fit together: how it works. Keeping a day of agent work reviewable: a task-tracking workflow for Claude Code.