Sequential Thinking MCP: What It Does and Whether You Need It
The Sequential Thinking server is one of the MCP project’s reference servers, and one of the most installed. What its single tool really does, its status today, how it compares with a model’s built-in thinking and with plan mode, and when it earns its tokens.
7 min read
The Sequential Thinking MCP server gives a model one tool, sequentialthinking, that it calls once per step of its reasoning: it writes a thought, numbers it, says how many it expects in total, and marks a thought as a revision or a branch when it changes course. The server itself does no thinking. It keeps the thoughts in memory, prints them to its log and hands back a small status message. The value is structure: reasoning that is broken into numbered steps and shows up as visible tool calls. It is still maintained in the MCP project’s reference servers, not archived. Whether you need it depends on your model. Current Claude models already think before they answer and between tool calls, and Anthropic now recommends that built-in thinking over a dedicated “think” tool in most cases. On those models, Sequential Thinking mostly adds tokens. On a model with no thinking of its own, or when you want the steps on the record, it can still help.
What the one tool actually does
The server registers a single tool. The code names it sequentialthinking; the heading in the Sequential Thinking README (opens in a new tab) spells it sequential_thinking, so look for either in your client’s tool list. Its inputs are the whole design:
thought: the text of the current step.thoughtNumberandtotalThoughts: where the model is and how many steps it now expects. The total can go up or down as it learns.nextThoughtNeeded: whether to keep going. The model sets it to false when it is done.isRevisionandrevisesThought: this step reconsiders an earlier one.branchFromThoughtandbranchId: this step explores an alternative from an earlier point.needsMoreThoughts: the model reached what it thought was the end and wants more room.
Read the source and the mechanism is plain. The server appends each thought to a list in memory, records branches, prints a formatted box to standard error unless DISABLE_THOUGHT_LOGGING is set to true, and returns the thought number, the total, whether another is needed, the branch names and the history length. All the reasoning lives in the arguments the model writes. The history disappears when the server process stops, and it is not shared between sessions or clients. It is not memory; for that, see agent memory over MCP.
The tool’s long description does the real work. It tells the model to break problems into steps, question and revise earlier thoughts, generate a hypothesis, check it against the earlier steps, and only stop when it is satisfied. In effect, the server is a prompt that the model reads on every request, plus a place to write each step down.
Is it maintained or archived?
Maintained. The MCP servers repository (opens in a new tab) lists Sequential Thinking among its current reference servers, beside Everything, Fetch, Filesystem, Git, Memory and Time, while servers such as GitHub, Google Drive and PostgreSQL moved to the archive. The npm package, @modelcontextprotocol/server-sequential-thinking, had a release on August 31, 2026, and the code had fixes in late August. The tool is marked read-only in its annotations, since it changes nothing outside itself.
Keep the repository’s own warning in mind: the servers there are reference implementations meant to demonstrate MCP features and SDK usage, not production-ready solutions. For a tool that only echoes its input, that matters less than it would for a server that touches files or a database, but it is the maintainers’ stated purpose.
Setting it up
It runs locally over stdio with npx or Docker and needs no keys. In Claude Desktop, add it to claude_desktop_config.json and restart the app:
{
"mcpServers": {
"sequential-thinking": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-sequential-thinking"],
"env": { "DISABLE_THOUGHT_LOGGING": "true" }
}
}
}The environment variable is optional; it stops the server printing each thought to its log. In Claude Code, one command adds it:
claude mcp add --transport stdio sequential-thinking -- npx -y @modelcontextprotocol/server-sequential-thinking
On Windows, the README wraps npx in cmd /c for Claude Desktop. To see it working, ask for something with several uncertain steps and watch for repeated sequentialthinking calls with a rising thoughtNumber. If you are new to editing the desktop config file, the Claude Desktop MCP config guide walks through it.
Compared with built-in thinking
Anthropic’s documentation on Claude’s thinking (opens in a new tab) describes the same job done inside the model. With thinking active, Claude works through the problem before it answers, tries approaches, checks intermediate results and abandons paths that do not hold up. With tools, it can think again between tool calls, reasoning about each result before the next step. On Anthropic’s newest models thinking is on by default and Claude decides how deeply to think. Thinking tokens are billed as output tokens, even when the text is not shown to you.
The closest comparison is a tool Anthropic itself published: a “think” tool with a single thought input that only appends the thought to a log, which is the Sequential Thinking idea without the numbering and branches. In a December 15, 2025 update to its think tool post (opens in a new tab), Anthropic says extended thinking has improved so much that it recommends using that feature instead of a dedicated think tool in most cases, with similar benefits and better integration and performance.
- Built-in thinking: happens inside the model, needs no server, can run between tool calls, and on some models shows you only a summary or nothing at all.
- Sequential Thinking: happens in visible tool calls, works with any model that can call tools, and leaves numbered steps in the transcript. It costs a round trip per step and the tool’s description in your context.
- Plan mode: a different job entirely. In Claude Code, plan mode (opens in a new tab) lets Claude read and explore but blocks edits until you approve its plan. It is a gate on action, not a way of reasoning, and it works alongside either kind of thinking. More in Claude Code plan mode.
When it helps
- Your model has no thinking of its own, or you run it with thinking off. Check the model you run locally, for example; if it does not reason before answering, the tool gives it a structure to follow.
- You want the reasoning on the record. Numbered thoughts in the tool log are easier to review afterward than a final answer, and some hosts hide or summarize built-in thinking.
- The work is planning with revisions: a migration, a refactor across modules, a design with an assumption that might not hold. The revision and branch fields suit that shape.
- You are teaching or debugging how an agent approaches a problem, and seeing each step as a separate call is the point.
When it only adds tokens
- Your model already thinks. Claude with thinking on, or any current reasoning model, is doing the same work internally; routing it through a tool mostly duplicates it.
- The task is simple. A lookup, a one-file fix or a short summary does not need ten numbered thoughts, but the tool description encourages them.
- Every thought is written as tool arguments, so it costs output tokens, and it then sits in the context for the rest of the session. The long tool description is loaded too, unless your client defers tool definitions; MCP token usage explains how.
- Your client asks for approval on each tool call. Approving a dozen thoughts per question gets old quickly.
- You expect it to make the model correct. It checks nothing. A wrong step written down neatly is still wrong.
The honest test is to run the same task twice, once with the server connected and once without, and compare the answers and the token counts. If the answers are the same, remove it; every connected server competes for the model’s attention.
Where the plan should end up
Sequential Thinking produces a plan, and then the plan vanishes with the server process. If the thinking was worth doing, it is worth keeping somewhere the next person or the next session can read it. That is the job a task’s plan field does on fenbs: the note says what the problem is, and the plan says how it will be done, rewritten as the work teaches you more.
With fenbs connected as a second MCP server at https://fenbs.ai/api/mcp, you can ask the assistant to finish its reasoning and then save the result: search the board for the task, and write the final steps into its plan with fenbs_update_item, or create a new task with fenbs_create_item if none exists. History records the change under the assistant’s name. Keep the plan to the steps that survived; the discarded branches belong in the transcript, not on the board. Breaking one plan into several tasks is covered in AI agent task decomposition.
Related
Other servers worth knowing: MCP examples. Where planning fits in a coding agent’s rhythm: a task-tracking workflow for Claude Code. Connecting fenbs: the MCP docs and Claude Code integration.