GitHub Spec Kit: Specs, Plans and Tasks for Coding Agents
Spec Kit is GitHub’s open-source toolkit for spec-driven development with an AI coding agent. How to install it, what each command does and which files it writes, a walk-through on one feature, its limits, and where the tasks should live once they exist.
8 min read
GitHub Spec Kit is an open-source, MIT-licensed toolkit that gives an AI coding agent a fixed process for building a feature from a written specification. You install a command-line tool called specify, run specify init in your project, and it adds templates, scripts and a set of speckit skills or commands for the agent you use. Then, in the agent’s chat, you set a constitution once and run specify, plan, tasks, implement and converge for each feature, with clarify, checklist and analyze as optional checks. Each feature gets its own folder holding spec.md, plan.md and tasks.md. It is thorough, it works with many agents, and it is more process than a small change needs.
What you need, and how to install it
The Spec Kit README (opens in a new tab) asks for Python 3.11 or later, the uv package manager and a supported coding agent, on Linux, macOS or Windows. The CLI is published on PyPI as specify-cli.
uv tool install specify-cli specify version # a new project specify init my-app --integration claude # or an existing repository: commit first, then specify init --here --force --integration claude
--here initialises the current folder and --force lets it merge into one that already has files, which is why the existing-project guide says to commit or stash your work first, so every generated file shows up in an ordinary review. If you leave out --integration, an interactive terminal asks which agent you use; a non-interactive run, such as CI, falls back to GitHub Copilot. Git is optional: numbered feature branches come from a git extension you add with specify extension add git.
Which agents it supports
The integrations reference (opens in a new tab) lists a long and growing set of agents, each with a key you pass to --integration. A few examples: copilot installs skills under .github/skills/, claude installs skills in .claude/skills, codex uses .agents/skills, and there are keys for Gemini CLI, Cursor, Kiro CLI, opencode and many more, plus a generic option for an agent that is not listed. specify integration list shows what your installed version supports.
The steps are the same everywhere; only the spelling changes. Copilot and Claude Code invoke them as /speckit-specify, Codex as $speckit-specify, and the reference pages write them in dotted form, /speckit.specify. They run in the agent’s chat, not in your terminal.
The commands, and the files each one writes
constitution: the project’s principles, which every later step is checked against. Run once, update when your principles change. Writes.specify/memory/constitution.md.specify: the feature specification, what and why, no tech stack. Creates a numbered folder such asspecs/001-password-reset/withspec.mdand a built-in quality checklist,checklists/requirements.md. It may leave up to three[NEEDS CLARIFICATION]markers for decisions it will not guess.clarify(optional): asks up to five targeted questions about under-specified areas and writes the answers back intospec.md.plan: the technical plan, where your stack and architecture go. Writesplan.md, and alongside itresearch.md,data-model.md, acontracts/folder when there are external interfaces, andquickstart.md. It checks the plan against the constitution before and after design.checklist(optional): “unit tests for your requirements”, custom checklists underchecklists/that ask whether the spec is complete and unambiguous. They belong to the reviewer: the agent may help evaluate them when asked, but must not tick them off on its own.tasks: a dependency-orderedtasks.md, in phases: Setup, Foundational, one phase per user story in priority order, then Polish. Tasks carry ids such asT001, and[P]marks ones that can run in parallel.analyze(optional): a read-only consistency check acrossspec.md,plan.mdandtasks.md. It writes nothing; it reports.implement: works throughtasks.mdin dependency order. It reads the checklists first and asks before going on if any item is unticked.converge: compares the code with the spec, plan and tasks. Either it reports Converged, or it appends the gaps as new tasks in a Convergence phase oftasks.md. You repeat implement and converge until it converges.taskstoissues(optional): turnstasks.mdinto GitHub issues. It needs a GitHub remote and the GitHub MCP tools.
The command reference (opens in a new tab) says only specify is strictly required before plan; the shorter path for small features is specify, plan, tasks, implement, converge. The full path adds clarify, checklist and analyze for production work. The folder that commands act on is recorded in .specify/feature.json, not taken from your git branch.
A walk-through on one feature
Say you are adding password reset by email to an existing web app, with Claude Code as the agent. After specify init --here --force --integration claude, open Claude Code in the project and work through the steps one at a time, reading each result before you run the next.
/speckit-constitution Every change has tests. No new dependencies without asking. Never log tokens. /speckit-specify Users who forget their password can reset it by email. The link works once and expires after an hour. Out of scope: SMS, security questions. /speckit-clarify /speckit-plan Use the existing mailer and the users table. Store only a hash of the token. /speckit-tasks /speckit-analyze /speckit-implement Only the Setup and Foundational phases. /speckit-implement Now the first user story. /speckit-converge
- After specify, read
spec.mdas the product owner would. Are the user stories right, and is anything missing from out of scope? A clarification marker such as “what should the page say when the address is not registered?” is a decision for a person, not the agent. - Clarify turns vague spots into questions. Answer them; the answers go into the spec, not into the chat.
- After plan, read
plan.mdanddata-model.mdas the tech lead would. This is where “store only a hash of the token” should now appear as a design choice, checked against the constitution’s rule about never logging tokens. - After tasks, skim
tasks.mdfor size. A task that says “implement reset flow” is too big; ask for it to be split before you implement. - Implement in stages. The reference recommends scoping large features by phase so the agent’s context is not overwhelmed, and checking each stage works before moving on.
- Converge last. If it appends tasks, implement them and converge again.
Limits worth knowing
- The checks are the agent checking itself. Analyze and converge are useful, but they are the model reading its own artifacts. Your tests and a person’s review are still what tell you the feature works.
- It is heavy for small changes. Nine steps around a copy change is ceremony. For a one-line fix, a good task description is the spec.
- Specs drift unless you choose how they change. The documentation deliberately leaves this to you and names three persistence models (opens in a new tab): flow-back, flow-forward and living spec. Pick one per project.
- Large features strain context. That is why implement can be scoped to a phase or a story.
- It does not rewrite your existing code or infer specs from it. Adopting it in an existing repository adds the
.specify/files and the agent’s skills; specs for behaviour that already exists are yours to write if you want them. - The task list is a file.
tasks.mdis right for the agent at work, merges with branches like any file, and is invisible to anyone who is not in the repository. The built-in export goes to GitHub issues only.
Spec Kit also ships two other processes you can use on their own: a bug-fixing extension (assess, fix, test, with reports under .specify/bugs/) and an idea assessment extension that ends in a go, clarify or stop decision. Both are added with specify extension add.
Keeping the tasks on a board
The tasks in tasks.md are implementation steps: “T014 add the token table”, “T015 hash the token before saving”. The people who decide what gets built next and check what came out care about a different unit: the user story. The split that works is one board task per user story in spec.md, with the T0xx steps staying in the file.
On fenbs each board task has a note for the problem and a plan for how it will be done. Put the user story and the path to its spec.md in the note, and a short summary of the relevant part of plan.md in the plan. Order Next Up by the story priorities the spec gives them. A [NEEDS CLARIFICATION] marker becomes an open question on the Decisions page, which only a person can decide. When a story’s tasks are implemented and converge reports no gaps, the agent sets the task’s testing status and notes and a person moves it to Completed.
## Spec Kit and the board - One fenbs task per user story in specs/<feature>/spec.md. Never one per T0xx task. - Note: the story and the spec path. Plan: the part of plan.md that covers it. - A [NEEDS CLARIFICATION] marker: fenbs_add_decision as an open question. Do not guess. - After /speckit-converge reports Converged: set testStatus and testNotes, then comment with the commits. Leave the move to Completed to a person.
fenbs does not read tasks.md or sync with it; the assistant files and updates the board tasks over MCP with fenbs_create_item and fenbs_update_item. Related tasks can be linked with relatesTo, but a link does not block anything, and fenbs has no epics or sprints to mirror Spec Kit’s phases.
Related
The method behind the tool: what is spec-driven development. Doing it with Claude Code’s own tools, with or without Spec Kit: spec-driven development with Claude Code. Another spec-first tool: Kiro vs Claude Code. Connecting an agent to the board: Claude Code integration and the MCP docs.