Spec-Driven Development vs Vibe Coding vs TDD

Three ways to build with an AI coding agent, tried on the same small project: where each one breaks, which to reach for when, and how they fit together rather than compete.

8 min read

Vibe coding, spec-driven development and test-driven development answer different questions, so choosing between them is less of a contest than it looks. Vibe coding asks “what do I want next?” and lets the conversation be the spec. Spec-driven development asks “what are we building, for whom, and how will we know it works?” and writes the answer down before any code. Test-driven development asks “what should this one piece of code do?” and answers with a failing test. Vibe coding is the fastest to start and the hardest to change later; a spec costs an hour up front and pays for itself once the work outlasts a session; TDD works inside either. The combination that holds up is a spec for the feature, tests first inside each task, and vibe coding kept for experiments you will throw away.

The three, in one line each

  • Vibe coding: describe what you want, run what the AI produces, say what is wrong, repeat. The full story, and how to do it without losing control, is in vibe coding with a backlog.
  • Spec-driven development: a written spec, then a plan, then small tasks, then implementation checked against the spec. The loop and what a good spec holds are in what is spec-driven development.
  • Test-driven development: the Agile Alliance glossary describes TDD (opens in a new tab) as a style in which coding, testing and design are tightly interwoven: write one unit test, watch it fail, write just enough code to pass, refactor, repeat.

Notice the scale of each. Vibe coding and a spec work at the level of a feature. TDD works at the level of a function or a behaviour. That is why the useful question is rarely “which one?” and more often “which one at which level?”.

One small project, three ways

The project: a waitlist page for a product launch. A visitor enters an email address and joins the list; the same address cannot join twice; each new sign-up gets a confirmation email; the owner can see how many people have joined.

Vibe coded

You ask for a page with an email form and a list, run it, ask for the duplicate check, run it, ask for the confirmation email, run it. Twenty minutes later it works on your machine. Then real use finds the gaps nobody asked about. Ann@example.com and ann@example.com both join, because the duplicate check is case-sensitive. A double click sends two confirmation emails. Nobody decided whether the form needs a consent line for marketing, so the AI either invented one or left it out, and you cannot tell which from the chat. None of these is hard to fix. The trouble is that each was a decision, and the conversation made it by accident.

Spec first

You spend half an hour on a one-page spec before the agent writes anything. Writing the acceptance criteria is where the questions surface: what counts as the same address, what happens on a double submit, what the consent line says.

spec.md, the part that matters
## Acceptance criteria
1. WHEN a visitor submits a valid email THE SYSTEM SHALL add it and
   show "You are on the list".
2. WHEN the address is already on the list, in any letter case,
   THE SYSTEM SHALL not add it again and show the same message.
3. WHEN the form is submitted twice within a second THE SYSTEM SHALL
   send one confirmation email.
4. WHEN the owner opens /admin THE SYSTEM SHALL show the total count.

## Out of scope
Unsubscribe links, export, double opt-in.

## Open questions
- Consent wording for marketing email. (Owner to decide.)

The agent plans from that, splits it into four tasks and builds them one at a time. It takes longer to reach the first working page. It does not ship the case bug, the double email or an invented consent line, because each was written down before the code existed. The one open question waits for the person who can answer it.

Tests first

With TDD you start from the smallest behaviour: a test that “ann@example.com” is refused after “Ann@example.com” has joined. It fails. The agent writes the normalisation, the test passes, and the code is tidied with the test as a safety net. Then the next test. Each piece of logic ends up correct and covered. What TDD cannot tell you is whether you wrote the right tests. It will not raise the consent question, and it will not notice that nobody asked for an admin count, because tests describe the behaviour you already thought of.

Where each one breaks down

  • Vibe coding breaks on anything that lasts. Decisions are made implicitly, the reasons live in a chat that scrolls away, and the next session, or the next person, cannot tell a deliberate choice from an accident.
  • Spec-driven development breaks on small or uncertain work. A spec for a one-line fix is ceremony, and a spec written before you understand the problem is confident fiction. It also breaks when the spec is written once and never updated, so the code and the document drift apart.
  • TDD breaks where the question is not “does this function behave?”. User interface feel, copy, layout and anything whose right answer is a product decision are hard to express as a failing unit test. And an agent asked to make tests pass can make them pass in ways you did not intend, such as by weakening an assertion, so the tests need reviewing too.

Which to use when

  • Vibe coding: a throwaway prototype, a personal tool on data that is not sensitive, or learning how something could be built before you build it properly.
  • A spec: the work will outlast one session, more than one person or agent will touch it, a mistake is expensive to undo, or the person who wants the feature is not the person reviewing the code.
  • TDD: logic with clear inputs and outputs, and nearly every bug fix. Anthropic’s Claude Code best practices (opens in a new tab) suggest describing the symptom and asking the agent to write a failing test that reproduces the issue, then fix it. The test proves the bug existed and stays behind to stop it coming back.

A rule of thumb that covers most cases: if you could write the acceptance criteria in the task itself, skip the separate spec. If you cannot say what correct looks like, you are not ready for either a spec or a test, and a short vibe-coded spike to find out is the honest next step.

How they combine

  1. Spike. Vibe code a rough version to learn what the problem really is. Keep what you learnt, not the code.
  2. Spec. Write the spec with what the spike taught you. The acceptance criteria are numbered, so tests and reviews can refer to them.
  3. Plan and tasks. The agent drafts a plan and splits it into tasks you can finish and check one at a time.
  4. TDD inside each task. Each acceptance criterion becomes at least one test, written and seen failing before the code. The same guide suggests one session writing the tests and another writing the code to pass them, so the implementer cannot quietly shape the tests to fit.
  5. Check against the spec. A task is done when its criteria pass and a reviewer, human or a fresh agent, has compared the result with the spec rather than with how plausible the code looks.

Spec-driven development vs agile

People sometimes read spec-first as a return to big design up front. The Agile Manifesto (opens in a new tab) values working software over comprehensive documentation, and responding to change over following a plan. A spec-driven workflow fits those values when the spec is small, written per feature just before the work, and changed as soon as the work shows it is wrong. It stops fitting when the spec is a large document written months ahead and treated as a contract.

How a spec should change is a choice teams rarely make on purpose. GitHub’s Spec Kit documentation names three spec persistence models (opens in a new tab): flow-back, where any document can be edited and the team reconciles them; flow-forward, where finished specs stay as history and a change gets a new one; and living spec, where the spec is the source and the plan and tasks are regenerated from it. Pick one, write it down, and the agile objection mostly goes away.

Keeping all three on one board

Each approach leaves something different behind, and a board can hold all of it without much ceremony. On fenbs, vibe coding fills To Do with the one-line ideas you noticed and did not chase. A spec becomes one task per outcome someone could check: the note says what is missing and where the spec file lives, and the plan box holds how it will be built, rewritten as the agent learns. TDD shows up in the task’s testing status and test notes, which say what was checked and what was not, so “Tested” means something.

The spec’s open questions go on the Decisions page, where the decider is always a person, and an assistant connected over MCP can ask one with fenbs_add_decision but never answer it. Tasks move through To Do, Next Up, In Progress and Completed. fenbs does not run your tests, and it has no sprints or story points; the lanes and the testing field are all the structure it adds.

Related

Turning a spec into small tasks: AI agent task decomposition. Writing criteria an agent can check: acceptance criteria examples. The plan step inside one session: Claude Code plan mode. How a task’s problem, plan and testing fit together: how it works.

Questions people ask.

Is spec-driven development the opposite of vibe coding?

Mostly, at the level of a feature. Vibe coding lets the conversation decide what gets built; spec-driven development writes that down and gets it approved first. They still combine well: a short vibe-coded spike is often the best way to learn enough to write the spec.

What is the difference between spec-driven development and TDD?

Scale and question. A spec describes a whole feature: who it is for, what is in and out of scope, and how it will be accepted. TDD works one behaviour at a time, writing a failing test before the code. Acceptance criteria in a spec are a natural source of tests.

Can you use TDD while vibe coding?

Yes, and it is one of the cheapest ways to add safety. Ask the AI for a failing test before each change, check that it fails for the right reason, then ask for the code. You keep the pace and gain a record of what the code is supposed to do.

Is spec-driven development compatible with agile?

It can be. Keep specs short, write them per feature just before the work, and change them when the work shows they are wrong. It conflicts with agile values when specs become large documents written far ahead and treated as fixed contracts.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.