How to Debug MCP Tools When an Assistant Gets It Wrong

When an assistant calls the wrong tool, passes the wrong arguments or shrugs at an error, the fault is in one of four places. How to find which, using the client’s own status and logs, a replay in the MCP Inspector, and a careful read of what the model was actually shown.

Updated 8 min read

To debug an MCP tool, work out which of four layers failed before changing anything. Connection: the tool never reached the assistant. Selection: it was there, but the model did not call it, or called another. Arguments: it was called with the wrong inputs. Result: it ran, but what came back was an error the model could not use, or too much to read. Check them in that order. The client’s status panel and debug log answer the first. Replaying the exact call in the MCP Inspector separates the server from the model. And reading the tool’s name, description and schema the way the model reads them explains most of the rest, because in practice the usual cause of an assistant getting a tool wrong is how the tool describes itself.

This assumes the server exists and worked when it was built; building one and first testing it in the Inspector is covered in how to build an MCP server. This piece is for when it misbehaves in real use.

Name the failure first

  • “I don’t have a tool for that”: connection. The server failed, is waiting for approval or sign-in, or returned no tools.
  • It used a different tool, or did the job by hand: selection. Descriptions overlap or do not say when to use the tool.
  • It called the right tool and got a refusal or validation error: arguments. The schema allowed something the server rejects, or did not say what a field means.
  • It said the job was done when it was not, or gave up: result. The error came back in a form the model could not read, or the output was cut off.

Layer one: is the tool there at all?

Start with the client’s own view. In Claude Code, /mcp shows every configured server, its status and whether it is approved for this project. claude mcp list shows a health status for each server, such as connected, needs authentication or failed to connect, and claude mcp get <name> adds an Issue line with the HTTP status or error the server returned. A project server in .mcp.json shows as pending approval until someone approves it.

Claude Code’s guide to debugging your configuration (opens in a new tab) lists the common causes. Relative paths in a server’s command or arguments resolve against the folder you started Claude Code in, not the location of .mcp.json. And a server that connects but lists zero tools has started without returning a tool list: reconnect it from /mcp, and if the count stays at zero, run with MCP debugging on and read the server’s standard error in the debug log.

Terminal
claude mcp list                  # status of every server for this project
claude mcp get fenbs             # one server, with an Issue line if it failed
claude --debug=mcp               # start a session with MCP debug logging
# then read ~/.claude/debug/<session-id>.txt

The --debug flag takes categories, as the CLI reference (opens in a new tab) shows, so --debug=mcp keeps the log to the part you need, and --debug-file writes it to a path of your choosing. Two more causes are worth knowing. If a stdio server writes anything that is not protocol to standard output, it corrupts the stream; its own logging belongs on standard error. And a server slow to start may simply time out: MCP_TIMEOUT sets the startup timeout in milliseconds.

Claude Code also checks tool schemas when it loads a server. According to the Claude Code MCP documentation (opens in a new tab), a tool whose input schema would fail the API’s checks, such as a property name with characters outside letters, digits, _, . and -, is left out while the server’s other tools keep working, and the reason is recorded in the server’s log. So “one tool is missing” is often a schema problem, not a connection problem.

Layer two: replay the call without the model

Once the tool is present, take the model out of the picture. Copy the exact arguments the assistant sent from its transcript and call the tool yourself. The Inspector’s command-line client (opens in a new tab) is made for this: it connects, makes one request, prints the result and exits.

Terminal
# A local stdio server: everything positional is the command that starts it
npx @modelcontextprotocol/inspector --cli node build/index.js --method tools/list

# A remote server, replaying the arguments exactly as the assistant sent them
npx @modelcontextprotocol/inspector --cli https://example.com/mcp --transport http \
  --method tools/call --tool-name create_invoice \
  --tool-args-json '{"customer":"C-0042","due":"2026-10-31"}' --format json

Use --tool-args-json for replays rather than --tool-arg key=value. The latter parses each value as JSON, so "012" becomes the number 12, which can make a failing call pass or a passing one fail for reasons that have nothing to do with the assistant. The CLI’s exit code tells you the class of failure: 3 when the server needs authentication, 4 when it cannot be reached, and 5 when the tool returned isError: true or does not exist. The Inspector can also read a client’s existing configuration file with --config, and its web client can import configurations from Claude Desktop, Cursor, Cline and VS Code, so you test the server exactly as the client starts it.

Now you know which side is at fault. If the replay fails the same way, the server is wrong: fix it and replay again. If the replay works, the server is doing what it was told, and the problem is what the model was told.

Layer three: read the tool the way the model reads it

The model chooses a tool, and fills in its arguments, from the name, the description and the input schema. Nothing else. Print the tool list and read it cold, as if you had never seen the code.

  • Does the description say when to use the tool, not only what it does? “Search tasks” loses to “Search tasks by text across every board you can see; use this before creating one.”
  • Do two tools overlap? search and find_items on the same server, or two servers each with a search, is a coin toss. Prefix names with the product and merge near-duplicates.
  • Is anything important late in a long description? Claude Code truncates each tool description, and each server’s instructions, at 2,048 characters by default, so the rule in the last paragraph may never reach the model.
  • Does every field say what it means and in what form? A date with no format, an id that could be three different kinds of id, a free string where an enum belongs.
  • Is the tool found at all? With tool search, which is on by default in Claude Code, only tool names and server instructions load at the start, so a server’s instructions should say what kind of task its tools are for.

Layer four: read the error the model read

The MCP specification (opens in a new tab) separates two kinds of failure. Protocol errors, such as an unknown tool or a malformed request, come back as JSON-RPC errors, which models are less likely to recover from. Tool execution errors, including API failures, business rules and input validation errors such as a date in the wrong format, come back as a normal result with isError: true and text explaining what went wrong. Clients should pass those to the model so it can correct itself and retry.

So a server that throws an exception on a bad date, rather than returning an isError result that says what format it wanted, has given the model nothing to work with. When an assistant gives up or loops, look at the raw result. Is it a sentence saying what was wrong and what to do next, or a stack trace, an empty string, or a bare “error”?

Size is the other result problem. Claude Code warns when a tool’s output exceeds 10,000 tokens and limits it to 25,000 by default; MAX_MCP_OUTPUT_TOKENS raises the limit. If a tool returns everything it has, the model may be reading a truncated answer. Pagination and filters fix that at the source.

Logs on the server side

The protocol’s debugging guide (opens in a new tab) sets out where server logs go. A stdio server’s standard error goes to the client, which may capture, forward or ignore it, so check where yours puts it. A remote server’s does not reach the client at all, so use your own log aggregation or OpenTelemetry, and ordinary HTTP tools to inspect requests. Logging through the protocol itself, with notifications/message, is deprecated as of the 2026-07-28 revision. Whatever you use, log each tool call with its arguments, the caller, the duration and the outcome, with secrets and personal data masked.

Keep the fix

A misbehaving tool is a bug, so file it like one, with the exact call, the result the model saw and the fix. Then turn the replay into a check that runs on every change to the server: the Inspector CLI’s JSON output and exit codes are built for CI, and its --stored-auth-only flag makes a run fail at once when no token is stored instead of waiting for a browser sign-in.

fenbs, a task board with its own MCP server, is built around layer four. A permission refusal comes back as a tool result with isError: true, not a protocol error, and names the permission that was missing and the role the caller holds, so the assistant can relay it rather than hunt for a workaround. Not-found errors and rate limits come back the same way, the latter telling the assistant to wait and try again. Tools say when to call them: fenbs_whoami is described as the one to call first. And the bug itself belongs on the board, filed with fenbs_create_item as kind bug, with the replay command in the note.

Related

Designing tools that are hard to misuse: how to build an MCP server. When the assistant itself stalls rather than a tool: Claude Code task stuck. The fenbs tools and how to connect them: MCP docs.

Questions people ask.

Why does my assistant say it has no tool for something my MCP server provides?

Usually the server failed to start, is waiting for project approval or sign-in, or returned no tools. In Claude Code, check /mcp and claude mcp list, then run with --debug=mcp and read the debug log. A single missing tool is often one whose input schema the client rejected.

How do I reproduce an MCP tool call outside the assistant?

Use the MCP Inspector command-line client with --method tools/call, the tool name, and the exact arguments passed as JSON with --tool-args-json. It exits with code 5 when the tool returns isError, so the same command works as a CI check.

Should an MCP tool throw an error or return isError?

Return a result with isError set to true for anything the model could fix, such as bad input, a business rule or a failing API, with a sentence saying what went wrong and what to do. Protocol errors are for problems like an unknown tool or a malformed request, which models are less likely to recover from.

Where are MCP server logs?

For a stdio server, whatever it writes to standard error is captured by the client; in Claude Code it appears in the debug log. For a remote server, use your own logging or OpenTelemetry. Never log to standard output from a stdio server, because that is the protocol channel.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.