Logging for MCP Servers: What to Record and Where
A stdio MCP server that logs to standard output breaks its own connection, and the protocol’s built-in logging is now deprecated. Where an MCP server’s logs should go, what one record per tool call should hold, what to keep out of it, and how debug logs, audit records and traces differ.
7 min read
Log an MCP server the way you log any service, with three rules of its own. On stdio, never write anything but protocol messages to standard output: logs go to standard error. Do not build on the protocol’s own logging feature, notifications/message, which the 2026-07-28 revision deprecates in favour of standard error and OpenTelemetry. And write one structured record per tool call, holding the tool, the caller, the outcome and the duration, but not the raw arguments or results, which is where secrets and personal data live. Keep that operational log apart from the audit record of what changed, and let traces tie a call to the request that caused it.
For the agent side of the picture, the model calls and the runs, see AI agent observability. This page is about the server.
The stdio rule: standard output is not yours
A client that starts your server as a child process reads its standard output as the protocol channel. The stdio transport (opens in a new tab) is blunt about it: the server must not write anything to standard output that is not a valid MCP message. It may write UTF-8 text to standard error for any logging, and the client may capture it, forward it or ignore it, and should not take it as a sign of an error.
- Python: use
logging, which writes to standard error by default, neverprint. The SDK diverts flushed stray output while serving, but its logging guide warns that an unflushedprintcan still land on the protocol stream at exit. - TypeScript and Node:
console.error, neverconsole.log. Oneconsole.logputs a line no JSON-RPC parser accepts into the stream the client reads. - C#: send the console logger to standard error, for example with
LogToStandardErrorThreshold = LogLevel.Trace. - Libraries count too. A dependency that prints a banner or a progress bar corrupts the channel just as surely; check before you add one.
A remote server over Streamable HTTP has no such trap, and no client capturing its output either. Its logs go wherever the rest of your services log. Streamable HTTP also copies the method and the tool name into the Mcp-Method and Mcp-Name headers, which the specification says is so that gateways and observability tools can route and inspect requests without parsing the body.
The protocol’s logging feature is deprecated
MCP has had a way for a server to send log lines to the client: a notifications/message notification with one of eight syslog levels, from debug to emergency, and data of any shape. The logging page (opens in a new tab) of the current revision marks it deprecated under SEP-2577. It stays in the specification for at least twelve months, so the earliest it can be removed is the first revision released on or after 2027-07-28. New servers should not adopt it; existing ones should move to standard error for stdio and to OpenTelemetry for structured observability.
{
"jsonrpc": "2.0",
"method": "notifications/message",
"params": { "level": "warning", "logger": "import", "data": { "skipped": 3 } }
}Two changes in the same revision matter if you still use it. logging/setLevel is gone: a client that wants log messages puts io.modelcontextprotocol/logLevel in the request’s _meta, and a server must not send notifications/message for a request that did not. And the messages travel only on that request’s own response stream. Either way, they were never a place for your operational logs: they go to the client, not to you.
Structured, one line per event
Write JSON lines rather than sentences, so that you can filter by tool, caller or outcome instead of searching text. In Python that is a formatter and a decorator around each tool:
import functools
import json
import logging
import sys
import time
from mcp.server import MCPServer
from mcp.server.mcpserver.exceptions import ToolError
class JsonLines(logging.Formatter):
def format(self, record: logging.LogRecord) -> str:
entry = {"ts": self.formatTime(record), "level": record.levelname.lower(), "msg": record.getMessage()}
entry.update(getattr(record, "fields", {}))
return json.dumps(entry)
handler = logging.StreamHandler(sys.stderr) # never stdout
handler.setFormatter(JsonLines())
calls = logging.getLogger("tasks.calls")
calls.addHandler(handler)
calls.setLevel(logging.INFO)
calls.propagate = False
def logged(fn):
@functools.wraps(fn)
def wrapper(**kwargs):
start, outcome = time.perf_counter(), "ok"
try:
return fn(**kwargs)
except ToolError:
outcome = "tool_error"
raise
except Exception:
outcome = "crash"
raise
finally:
calls.info("tool call", extra={"fields": {
"tool": fn.__name__,
"arg_keys": sorted(kwargs), # names only, never values
"outcome": outcome,
"ms": round((time.perf_counter() - start) * 1000),
}})
return wrapper
mcp = MCPServer("tasks")
@mcp.tool()
@logged
def tasks_get(ref: str) -> str:
"""Read one task by its reference, such as T-1."""
return "Draft the release notes"Run against version 2.2.0 of the Python SDK, a call to tasks_get produces a single line on standard error with the tool, arg_keys: ["ref"], the outcome and the time taken, and nothing extra on standard output. The SDK’s own logging still reports crashes with their tracebacks, so the decorator does not need to.
What to record for every tool call
- When, and how long: a UTC timestamp and the duration in milliseconds.
- Which call: the JSON-RPC request id, the method and the tool name.
- Who: the identity your authentication established, such as a user id and the name of the token used. The client’s
clientInfoname and version are worth recording for display, but the specification notes they are self-reported and unverified, so never treat them as the caller. - Which protocol version the request declared, which tells you when a client is behind.
- What was asked, in outline: the argument names, sizes and counts, and identifiers you have decided are safe. Not the values by default.
- What happened:
ok,tool_errorfor a result withisError: true,protocol_errorwith its JSON-RPC code, orcrash. Keep the message of a tool error, which you wrote; keep the traceback of a crash in its own record. - How big the answer was: a count of items or bytes, which is how you notice a tool that returns far more than the model can use.
- A trace id, so the line can be joined to the trace of the whole request.
{
"ts": "2026-09-28T10:41:07.212Z",
"level": "info",
"msg": "tool call",
"request_id": "17",
"method": "tools/call",
"tool": "tasks_get",
"caller": "user:4821",
"token": "Claude laptop",
"client": "example-client 1.0.0",
"protocol": "2026-07-28",
"arg_keys": ["ref"],
"outcome": "tool_error",
"error": "No task T-9. Refs look like T-1.",
"result_items": 0,
"ms": 38,
"trace_id": "0af7651916cd43dd8448eb211c80319c"
}Levels follow the audience. A tool error is expected behaviour, the model misjudged and was told, so log it at INFO; the Python SDK does the same. A crash is ERROR, with its traceback. Then an alert on ERROR means something is actually broken, and a busy morning of mistyped task refs stays quiet.
Keep secrets and personal data out
Tool arguments and results are where the sensitive material is: a customer’s email in a search, a file’s contents, a token pasted into a field. The specification’s logging page says log messages must not contain credentials or secrets, personal identifying information, or internal details that could aid an attack, and the OWASP logging cheat sheet (opens in a new tab) adds access tokens, session identifiers, connection strings and encryption keys to the list to remove, mask, hash or encrypt.
- Allow-list, do not deny-list. Log the fields you chose; a new argument then stays out until someone decides it is safe.
- Never log headers wholesale.
Authorizationcarries a bearer token on every request to a remote server. - Hash identifiers you need to correlate but not read, such as an email address used as a lookup key.
- Treat tool results like arguments. A read tool returns whatever it read.
- Keep full payloads, if you need them at all, behind a flag you switch on for one investigation, with a short retention.
Debug logs are not an audit trail
The log above answers “why did that call fail?” for the person running the server. An audit record answers “who changed this, and when?” for the people who own the data. They differ in almost everything:
- Audience: operators and developers, against owners, reviewers and sometimes clients.
- Content: every call, including reads and failures, against the changes that were made, in words a person can read.
- Lifetime: long enough to investigate, against as long as your obligations require, and not rotated away with the debug noise.
- Place: your log pipeline, against the system that holds the data, where it cannot be edited by the caller.
Write the audit record where the change is made, in the same transaction if you can, not from a log line after the fact. Why it has to be kept by the tool rather than by the agent is covered in an audit trail for AI agents.
Tracing with OpenTelemetry
- The 2026-07-28 revision reserves
traceparent,tracestateandbaggagein_metafor OpenTelemetry trace context, in W3C format, so a client can pass its trace to the server inside the message. - The Python SDK’s OpenTelemetry guide (opens in a new tab) says every server already emits a span per message, named for example
tools/call tasks_get, withmcp.method.nameandmcp.protocol.version. They cost almost nothing until you install an OpenTelemetry SDK and an exporter. - OpenTelemetry’s conventions for MCP (opens in a new tab) are still marked Development. They set
error.typetotool_errorwhen a tool result hasisError: true, and make recording tool arguments and results opt-in, since they may contain sensitive information.
Put the trace id in your log line, and a slow or failing call can be followed from the assistant’s request through your server to the database and back.
What fenbs records
fenbs, a task board where AI assistants are members with roles, keeps the audit side where the data lives: every change on the board is recorded in History with who made it, an assistant’s changes carry the assistant’s name and the person it acts for, and revoking an assistant’s access leaves that record in place. Its stdio bridge, fenbs-mcp, follows the stdio rule: it writes only JSON-RPC replies to standard output and its own complaints, such as a missing FENBS_TOKEN, to standard error. Connecting an assistant is described in the MCP docs.
Related
The agent side: AI agent observability. The people-readable record: an audit trail for AI agents. Deciding which failures are tool errors: MCP error handling. What the logged messages look like: why MCP uses JSON-RPC.