Logging for MCP Servers: What to Record and Where

A stdio MCP server that logs to standard output breaks its own connection, and the protocol’s built-in logging is now deprecated. Where an MCP server’s logs should go, what one record per tool call should hold, what to keep out of it, and how debug logs, audit records and traces differ.

7 min read

Log an MCP server the way you log any service, with three rules of its own. On stdio, never write anything but protocol messages to standard output: logs go to standard error. Do not build on the protocol’s own logging feature, notifications/message, which the 2026-07-28 revision deprecates in favour of standard error and OpenTelemetry. And write one structured record per tool call, holding the tool, the caller, the outcome and the duration, but not the raw arguments or results, which is where secrets and personal data live. Keep that operational log apart from the audit record of what changed, and let traces tie a call to the request that caused it.

For the agent side of the picture, the model calls and the runs, see AI agent observability. This page is about the server.

The stdio rule: standard output is not yours

A client that starts your server as a child process reads its standard output as the protocol channel. The stdio transport (opens in a new tab) is blunt about it: the server must not write anything to standard output that is not a valid MCP message. It may write UTF-8 text to standard error for any logging, and the client may capture it, forward it or ignore it, and should not take it as a sign of an error.

  • Python: use logging, which writes to standard error by default, never print. The SDK diverts flushed stray output while serving, but its logging guide warns that an unflushed print can still land on the protocol stream at exit.
  • TypeScript and Node: console.error, never console.log. One console.log puts a line no JSON-RPC parser accepts into the stream the client reads.
  • C#: send the console logger to standard error, for example with LogToStandardErrorThreshold = LogLevel.Trace.
  • Libraries count too. A dependency that prints a banner or a progress bar corrupts the channel just as surely; check before you add one.

A remote server over Streamable HTTP has no such trap, and no client capturing its output either. Its logs go wherever the rest of your services log. Streamable HTTP also copies the method and the tool name into the Mcp-Method and Mcp-Name headers, which the specification says is so that gateways and observability tools can route and inspect requests without parsing the body.

The protocol’s logging feature is deprecated

MCP has had a way for a server to send log lines to the client: a notifications/message notification with one of eight syslog levels, from debug to emergency, and data of any shape. The logging page (opens in a new tab) of the current revision marks it deprecated under SEP-2577. It stays in the specification for at least twelve months, so the earliest it can be removed is the first revision released on or after 2027-07-28. New servers should not adopt it; existing ones should move to standard error for stdio and to OpenTelemetry for structured observability.

The deprecated notification, for recognition
{
  "jsonrpc": "2.0",
  "method": "notifications/message",
  "params": { "level": "warning", "logger": "import", "data": { "skipped": 3 } }
}

Two changes in the same revision matter if you still use it. logging/setLevel is gone: a client that wants log messages puts io.modelcontextprotocol/logLevel in the request’s _meta, and a server must not send notifications/message for a request that did not. And the messages travel only on that request’s own response stream. Either way, they were never a place for your operational logs: they go to the client, not to you.

Structured, one line per event

Write JSON lines rather than sentences, so that you can filter by tool, caller or outcome instead of searching text. In Python that is a formatter and a decorator around each tool:

Python: one JSON line per tool call, on standard error
import functools
import json
import logging
import sys
import time

from mcp.server import MCPServer
from mcp.server.mcpserver.exceptions import ToolError


class JsonLines(logging.Formatter):
    def format(self, record: logging.LogRecord) -> str:
        entry = {"ts": self.formatTime(record), "level": record.levelname.lower(), "msg": record.getMessage()}
        entry.update(getattr(record, "fields", {}))
        return json.dumps(entry)


handler = logging.StreamHandler(sys.stderr)  # never stdout
handler.setFormatter(JsonLines())
calls = logging.getLogger("tasks.calls")
calls.addHandler(handler)
calls.setLevel(logging.INFO)
calls.propagate = False


def logged(fn):
    @functools.wraps(fn)
    def wrapper(**kwargs):
        start, outcome = time.perf_counter(), "ok"
        try:
            return fn(**kwargs)
        except ToolError:
            outcome = "tool_error"
            raise
        except Exception:
            outcome = "crash"
            raise
        finally:
            calls.info("tool call", extra={"fields": {
                "tool": fn.__name__,
                "arg_keys": sorted(kwargs),  # names only, never values
                "outcome": outcome,
                "ms": round((time.perf_counter() - start) * 1000),
            }})
    return wrapper


mcp = MCPServer("tasks")


@mcp.tool()
@logged
def tasks_get(ref: str) -> str:
    """Read one task by its reference, such as T-1."""
    return "Draft the release notes"

Run against version 2.2.0 of the Python SDK, a call to tasks_get produces a single line on standard error with the tool, arg_keys: ["ref"], the outcome and the time taken, and nothing extra on standard output. The SDK’s own logging still reports crashes with their tracebacks, so the decorator does not need to.

What to record for every tool call

  • When, and how long: a UTC timestamp and the duration in milliseconds.
  • Which call: the JSON-RPC request id, the method and the tool name.
  • Who: the identity your authentication established, such as a user id and the name of the token used. The client’s clientInfo name and version are worth recording for display, but the specification notes they are self-reported and unverified, so never treat them as the caller.
  • Which protocol version the request declared, which tells you when a client is behind.
  • What was asked, in outline: the argument names, sizes and counts, and identifiers you have decided are safe. Not the values by default.
  • What happened: ok, tool_error for a result with isError: true, protocol_error with its JSON-RPC code, or crash. Keep the message of a tool error, which you wrote; keep the traceback of a crash in its own record.
  • How big the answer was: a count of items or bytes, which is how you notice a tool that returns far more than the model can use.
  • A trace id, so the line can be joined to the trace of the whole request.
A fuller record
{
  "ts": "2026-09-28T10:41:07.212Z",
  "level": "info",
  "msg": "tool call",
  "request_id": "17",
  "method": "tools/call",
  "tool": "tasks_get",
  "caller": "user:4821",
  "token": "Claude laptop",
  "client": "example-client 1.0.0",
  "protocol": "2026-07-28",
  "arg_keys": ["ref"],
  "outcome": "tool_error",
  "error": "No task T-9. Refs look like T-1.",
  "result_items": 0,
  "ms": 38,
  "trace_id": "0af7651916cd43dd8448eb211c80319c"
}

Levels follow the audience. A tool error is expected behaviour, the model misjudged and was told, so log it at INFO; the Python SDK does the same. A crash is ERROR, with its traceback. Then an alert on ERROR means something is actually broken, and a busy morning of mistyped task refs stays quiet.

Keep secrets and personal data out

Tool arguments and results are where the sensitive material is: a customer’s email in a search, a file’s contents, a token pasted into a field. The specification’s logging page says log messages must not contain credentials or secrets, personal identifying information, or internal details that could aid an attack, and the OWASP logging cheat sheet (opens in a new tab) adds access tokens, session identifiers, connection strings and encryption keys to the list to remove, mask, hash or encrypt.

  • Allow-list, do not deny-list. Log the fields you chose; a new argument then stays out until someone decides it is safe.
  • Never log headers wholesale. Authorization carries a bearer token on every request to a remote server.
  • Hash identifiers you need to correlate but not read, such as an email address used as a lookup key.
  • Treat tool results like arguments. A read tool returns whatever it read.
  • Keep full payloads, if you need them at all, behind a flag you switch on for one investigation, with a short retention.

Debug logs are not an audit trail

The log above answers “why did that call fail?” for the person running the server. An audit record answers “who changed this, and when?” for the people who own the data. They differ in almost everything:

  • Audience: operators and developers, against owners, reviewers and sometimes clients.
  • Content: every call, including reads and failures, against the changes that were made, in words a person can read.
  • Lifetime: long enough to investigate, against as long as your obligations require, and not rotated away with the debug noise.
  • Place: your log pipeline, against the system that holds the data, where it cannot be edited by the caller.

Write the audit record where the change is made, in the same transaction if you can, not from a log line after the fact. Why it has to be kept by the tool rather than by the agent is covered in an audit trail for AI agents.

Tracing with OpenTelemetry

  • The 2026-07-28 revision reserves traceparent, tracestate and baggage in _meta for OpenTelemetry trace context, in W3C format, so a client can pass its trace to the server inside the message.
  • The Python SDK’s OpenTelemetry guide (opens in a new tab) says every server already emits a span per message, named for example tools/call tasks_get, with mcp.method.name and mcp.protocol.version. They cost almost nothing until you install an OpenTelemetry SDK and an exporter.
  • OpenTelemetry’s conventions for MCP (opens in a new tab) are still marked Development. They set error.type to tool_error when a tool result has isError: true, and make recording tool arguments and results opt-in, since they may contain sensitive information.

Put the trace id in your log line, and a slow or failing call can be followed from the assistant’s request through your server to the database and back.

What fenbs records

fenbs, a task board where AI assistants are members with roles, keeps the audit side where the data lives: every change on the board is recorded in History with who made it, an assistant’s changes carry the assistant’s name and the person it acts for, and revoking an assistant’s access leaves that record in place. Its stdio bridge, fenbs-mcp, follows the stdio rule: it writes only JSON-RPC replies to standard output and its own complaints, such as a missing FENBS_TOKEN, to standard error. Connecting an assistant is described in the MCP docs.

Related

The agent side: AI agent observability. The people-readable record: an audit trail for AI agents. Deciding which failures are tool errors: MCP error handling. What the logged messages look like: why MCP uses JSON-RPC.

Questions people ask.

Where should a stdio MCP server write its logs?

To standard error. Standard output is the protocol channel, and the specification forbids writing anything there that is not a valid MCP message. The client may capture standard error, show it in a debug log or ignore it.

Is MCP logging deprecated?

Yes. The protocol logging feature, notifications/message, is deprecated as of the 2026-07-28 revision under SEP-2577. It remains for at least twelve months, and the specification suggests standard error for stdio servers and OpenTelemetry for structured observability instead.

Should an MCP server log tool arguments?

Not by default. Log the argument names, sizes and identifiers you have decided are safe, and keep full arguments and results behind a switch for a specific investigation. They are where credentials and personal data turn up.

Is a server log enough as an audit trail?

No. A debug log is for operators, is rotated, and holds every call. An audit trail records the changes themselves, in terms people can read, kept by the system that holds the data for as long as it is needed.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.