Spec-Driven Development vs Vibe Coding vs TDD
Three ways to build with an AI coding agent, tried on the same small project: where each one breaks, which to reach for when, and how they fit together rather than compete.
8 min read
Vibe coding, spec-driven development and test-driven development answer different questions, so choosing between them is less of a contest than it looks. Vibe coding asks “what do I want next?” and lets the conversation be the spec. Spec-driven development asks “what are we building, for whom, and how will we know it works?” and writes the answer down before any code. Test-driven development asks “what should this one piece of code do?” and answers with a failing test. Vibe coding is the fastest to start and the hardest to change later; a spec costs an hour up front and pays for itself once the work outlasts a session; TDD works inside either. The combination that holds up is a spec for the feature, tests first inside each task, and vibe coding kept for experiments you will throw away.
The three, in one line each
- Vibe coding: describe what you want, run what the AI produces, say what is wrong, repeat. The full story, and how to do it without losing control, is in vibe coding with a backlog.
- Spec-driven development: a written spec, then a plan, then small tasks, then implementation checked against the spec. The loop and what a good spec holds are in what is spec-driven development.
- Test-driven development: the Agile Alliance glossary describes TDD (opens in a new tab) as a style in which coding, testing and design are tightly interwoven: write one unit test, watch it fail, write just enough code to pass, refactor, repeat.
Notice the scale of each. Vibe coding and a spec work at the level of a feature. TDD works at the level of a function or a behaviour. That is why the useful question is rarely “which one?” and more often “which one at which level?”.
One small project, three ways
The project: a waitlist page for a product launch. A visitor enters an email address and joins the list; the same address cannot join twice; each new sign-up gets a confirmation email; the owner can see how many people have joined.
Vibe coded
You ask for a page with an email form and a list, run it, ask for the duplicate check, run it, ask for the confirmation email, run it. Twenty minutes later it works on your machine. Then real use finds the gaps nobody asked about. Ann@example.com and ann@example.com both join, because the duplicate check is case-sensitive. A double click sends two confirmation emails. Nobody decided whether the form needs a consent line for marketing, so the AI either invented one or left it out, and you cannot tell which from the chat. None of these is hard to fix. The trouble is that each was a decision, and the conversation made it by accident.
Spec first
You spend half an hour on a one-page spec before the agent writes anything. Writing the acceptance criteria is where the questions surface: what counts as the same address, what happens on a double submit, what the consent line says.
## Acceptance criteria 1. WHEN a visitor submits a valid email THE SYSTEM SHALL add it and show "You are on the list". 2. WHEN the address is already on the list, in any letter case, THE SYSTEM SHALL not add it again and show the same message. 3. WHEN the form is submitted twice within a second THE SYSTEM SHALL send one confirmation email. 4. WHEN the owner opens /admin THE SYSTEM SHALL show the total count. ## Out of scope Unsubscribe links, export, double opt-in. ## Open questions - Consent wording for marketing email. (Owner to decide.)
The agent plans from that, splits it into four tasks and builds them one at a time. It takes longer to reach the first working page. It does not ship the case bug, the double email or an invented consent line, because each was written down before the code existed. The one open question waits for the person who can answer it.
Tests first
With TDD you start from the smallest behaviour: a test that “ann@example.com” is refused after “Ann@example.com” has joined. It fails. The agent writes the normalisation, the test passes, and the code is tidied with the test as a safety net. Then the next test. Each piece of logic ends up correct and covered. What TDD cannot tell you is whether you wrote the right tests. It will not raise the consent question, and it will not notice that nobody asked for an admin count, because tests describe the behaviour you already thought of.
Where each one breaks down
- Vibe coding breaks on anything that lasts. Decisions are made implicitly, the reasons live in a chat that scrolls away, and the next session, or the next person, cannot tell a deliberate choice from an accident.
- Spec-driven development breaks on small or uncertain work. A spec for a one-line fix is ceremony, and a spec written before you understand the problem is confident fiction. It also breaks when the spec is written once and never updated, so the code and the document drift apart.
- TDD breaks where the question is not “does this function behave?”. User interface feel, copy, layout and anything whose right answer is a product decision are hard to express as a failing unit test. And an agent asked to make tests pass can make them pass in ways you did not intend, such as by weakening an assertion, so the tests need reviewing too.
Which to use when
- Vibe coding: a throwaway prototype, a personal tool on data that is not sensitive, or learning how something could be built before you build it properly.
- A spec: the work will outlast one session, more than one person or agent will touch it, a mistake is expensive to undo, or the person who wants the feature is not the person reviewing the code.
- TDD: logic with clear inputs and outputs, and nearly every bug fix. Anthropic’s Claude Code best practices (opens in a new tab) suggest describing the symptom and asking the agent to write a failing test that reproduces the issue, then fix it. The test proves the bug existed and stays behind to stop it coming back.
A rule of thumb that covers most cases: if you could write the acceptance criteria in the task itself, skip the separate spec. If you cannot say what correct looks like, you are not ready for either a spec or a test, and a short vibe-coded spike to find out is the honest next step.
How they combine
- Spike. Vibe code a rough version to learn what the problem really is. Keep what you learnt, not the code.
- Spec. Write the spec with what the spike taught you. The acceptance criteria are numbered, so tests and reviews can refer to them.
- Plan and tasks. The agent drafts a plan and splits it into tasks you can finish and check one at a time.
- TDD inside each task. Each acceptance criterion becomes at least one test, written and seen failing before the code. The same guide suggests one session writing the tests and another writing the code to pass them, so the implementer cannot quietly shape the tests to fit.
- Check against the spec. A task is done when its criteria pass and a reviewer, human or a fresh agent, has compared the result with the spec rather than with how plausible the code looks.
Spec-driven development vs agile
People sometimes read spec-first as a return to big design up front. The Agile Manifesto (opens in a new tab) values working software over comprehensive documentation, and responding to change over following a plan. A spec-driven workflow fits those values when the spec is small, written per feature just before the work, and changed as soon as the work shows it is wrong. It stops fitting when the spec is a large document written months ahead and treated as a contract.
How a spec should change is a choice teams rarely make on purpose. GitHub’s Spec Kit documentation names three spec persistence models (opens in a new tab): flow-back, where any document can be edited and the team reconciles them; flow-forward, where finished specs stay as history and a change gets a new one; and living spec, where the spec is the source and the plan and tasks are regenerated from it. Pick one, write it down, and the agile objection mostly goes away.
Keeping all three on one board
Each approach leaves something different behind, and a board can hold all of it without much ceremony. On fenbs, vibe coding fills To Do with the one-line ideas you noticed and did not chase. A spec becomes one task per outcome someone could check: the note says what is missing and where the spec file lives, and the plan box holds how it will be built, rewritten as the agent learns. TDD shows up in the task’s testing status and test notes, which say what was checked and what was not, so “Tested” means something.
The spec’s open questions go on the Decisions page, where the decider is always a person, and an assistant connected over MCP can ask one with fenbs_add_decision but never answer it. Tasks move through To Do, Next Up, In Progress and Completed. fenbs does not run your tests, and it has no sprints or story points; the lanes and the testing field are all the structure it adds.
Related
Turning a spec into small tasks: AI agent task decomposition. Writing criteria an agent can check: acceptance criteria examples. The plan step inside one session: Claude Code plan mode. How a task’s problem, plan and testing fit together: how it works.