How to Debug MCP Tools When an Assistant Gets It Wrong
When an assistant calls the wrong tool, passes the wrong arguments or shrugs at an error, the fault is in one of four places. How to find which, using the client’s own status and logs, a replay in the MCP Inspector, and a careful read of what the model was actually shown.
Updated 8 min read
To debug an MCP tool, work out which of four layers failed before changing anything. Connection: the tool never reached the assistant. Selection: it was there, but the model did not call it, or called another. Arguments: it was called with the wrong inputs. Result: it ran, but what came back was an error the model could not use, or too much to read. Check them in that order. The client’s status panel and debug log answer the first. Replaying the exact call in the MCP Inspector separates the server from the model. And reading the tool’s name, description and schema the way the model reads them explains most of the rest, because in practice the usual cause of an assistant getting a tool wrong is how the tool describes itself.
This assumes the server exists and worked when it was built; building one and first testing it in the Inspector is covered in how to build an MCP server. This piece is for when it misbehaves in real use.
Name the failure first
- “I don’t have a tool for that”: connection. The server failed, is waiting for approval or sign-in, or returned no tools.
- It used a different tool, or did the job by hand: selection. Descriptions overlap or do not say when to use the tool.
- It called the right tool and got a refusal or validation error: arguments. The schema allowed something the server rejects, or did not say what a field means.
- It said the job was done when it was not, or gave up: result. The error came back in a form the model could not read, or the output was cut off.
Layer one: is the tool there at all?
Start with the client’s own view. In Claude Code, /mcp shows every configured server, its status and whether it is approved for this project. claude mcp list shows a health status for each server, such as connected, needs authentication or failed to connect, and claude mcp get <name> adds an Issue line with the HTTP status or error the server returned. A project server in .mcp.json shows as pending approval until someone approves it.
Claude Code’s guide to debugging your configuration (opens in a new tab) lists the common causes. Relative paths in a server’s command or arguments resolve against the folder you started Claude Code in, not the location of .mcp.json. And a server that connects but lists zero tools has started without returning a tool list: reconnect it from /mcp, and if the count stays at zero, run with MCP debugging on and read the server’s standard error in the debug log.
claude mcp list # status of every server for this project claude mcp get fenbs # one server, with an Issue line if it failed claude --debug=mcp # start a session with MCP debug logging # then read ~/.claude/debug/<session-id>.txt
The --debug flag takes categories, as the CLI reference (opens in a new tab) shows, so --debug=mcp keeps the log to the part you need, and --debug-file writes it to a path of your choosing. Two more causes are worth knowing. If a stdio server writes anything that is not protocol to standard output, it corrupts the stream; its own logging belongs on standard error. And a server slow to start may simply time out: MCP_TIMEOUT sets the startup timeout in milliseconds.
Claude Code also checks tool schemas when it loads a server. According to the Claude Code MCP documentation (opens in a new tab), a tool whose input schema would fail the API’s checks, such as a property name with characters outside letters, digits, _, . and -, is left out while the server’s other tools keep working, and the reason is recorded in the server’s log. So “one tool is missing” is often a schema problem, not a connection problem.
Layer two: replay the call without the model
Once the tool is present, take the model out of the picture. Copy the exact arguments the assistant sent from its transcript and call the tool yourself. The Inspector’s command-line client (opens in a new tab) is made for this: it connects, makes one request, prints the result and exits.
# A local stdio server: everything positional is the command that starts it
npx @modelcontextprotocol/inspector --cli node build/index.js --method tools/list
# A remote server, replaying the arguments exactly as the assistant sent them
npx @modelcontextprotocol/inspector --cli https://example.com/mcp --transport http \
--method tools/call --tool-name create_invoice \
--tool-args-json '{"customer":"C-0042","due":"2026-10-31"}' --format jsonUse --tool-args-json for replays rather than --tool-arg key=value. The latter parses each value as JSON, so "012" becomes the number 12, which can make a failing call pass or a passing one fail for reasons that have nothing to do with the assistant. The CLI’s exit code tells you the class of failure: 3 when the server needs authentication, 4 when it cannot be reached, and 5 when the tool returned isError: true or does not exist. The Inspector can also read a client’s existing configuration file with --config, and its web client can import configurations from Claude Desktop, Cursor, Cline and VS Code, so you test the server exactly as the client starts it.
Now you know which side is at fault. If the replay fails the same way, the server is wrong: fix it and replay again. If the replay works, the server is doing what it was told, and the problem is what the model was told.
Layer three: read the tool the way the model reads it
The model chooses a tool, and fills in its arguments, from the name, the description and the input schema. Nothing else. Print the tool list and read it cold, as if you had never seen the code.
- Does the description say when to use the tool, not only what it does? “Search tasks” loses to “Search tasks by text across every board you can see; use this before creating one.”
- Do two tools overlap?
searchandfind_itemson the same server, or two servers each with asearch, is a coin toss. Prefix names with the product and merge near-duplicates. - Is anything important late in a long description? Claude Code truncates each tool description, and each server’s instructions, at 2,048 characters by default, so the rule in the last paragraph may never reach the model.
- Does every field say what it means and in what form? A
datewith no format, anidthat could be three different kinds of id, a free string where an enum belongs. - Is the tool found at all? With tool search, which is on by default in Claude Code, only tool names and server instructions load at the start, so a server’s instructions should say what kind of task its tools are for.
Layer four: read the error the model read
The MCP specification (opens in a new tab) separates two kinds of failure. Protocol errors, such as an unknown tool or a malformed request, come back as JSON-RPC errors, which models are less likely to recover from. Tool execution errors, including API failures, business rules and input validation errors such as a date in the wrong format, come back as a normal result with isError: true and text explaining what went wrong. Clients should pass those to the model so it can correct itself and retry.
So a server that throws an exception on a bad date, rather than returning an isError result that says what format it wanted, has given the model nothing to work with. When an assistant gives up or loops, look at the raw result. Is it a sentence saying what was wrong and what to do next, or a stack trace, an empty string, or a bare “error”?
Size is the other result problem. Claude Code warns when a tool’s output exceeds 10,000 tokens and limits it to 25,000 by default; MAX_MCP_OUTPUT_TOKENS raises the limit. If a tool returns everything it has, the model may be reading a truncated answer. Pagination and filters fix that at the source.
Logs on the server side
The protocol’s debugging guide (opens in a new tab) sets out where server logs go. A stdio server’s standard error goes to the client, which may capture, forward or ignore it, so check where yours puts it. A remote server’s does not reach the client at all, so use your own log aggregation or OpenTelemetry, and ordinary HTTP tools to inspect requests. Logging through the protocol itself, with notifications/message, is deprecated as of the 2026-07-28 revision. Whatever you use, log each tool call with its arguments, the caller, the duration and the outcome, with secrets and personal data masked.
Keep the fix
A misbehaving tool is a bug, so file it like one, with the exact call, the result the model saw and the fix. Then turn the replay into a check that runs on every change to the server: the Inspector CLI’s JSON output and exit codes are built for CI, and its --stored-auth-only flag makes a run fail at once when no token is stored instead of waiting for a browser sign-in.
fenbs, a task board with its own MCP server, is built around layer four. A permission refusal comes back as a tool result with isError: true, not a protocol error, and names the permission that was missing and the role the caller holds, so the assistant can relay it rather than hunt for a workaround. Not-found errors and rate limits come back the same way, the latter telling the assistant to wait and try again. Tools say when to call them: fenbs_whoami is described as the one to call first. And the bug itself belongs on the board, filed with fenbs_create_item as kind bug, with the replay command in the note.
Related
Designing tools that are hard to misuse: how to build an MCP server. When the assistant itself stalls rather than a tool: Claude Code task stuck. The fenbs tools and how to connect them: MCP docs.