AI Agent Observability: Traces, Logs and the Record People Read
When an agent takes forty steps to finish a task, you need to see all forty when something goes wrong. What to capture, the OpenTelemetry conventions for it, what Claude Code and the agent SDKs already export, and why a machine trace is not the record your team reads.
7 min read
AI agent observability is the ability to see what an agent did on each run and why: every turn of its loop, every model request, every tool call with its result, the tokens and cost, and the errors. In practice it means exporting traces, metrics and log events, ideally in the OpenTelemetry format, to a backend where you can follow a single run from the first prompt to the last tool call. It is for engineers debugging and operating agents. It does not replace the plain record your team reads of what changed in its own tools, which is a separate thing kept for a different reader.
What to capture
- The loop as a trace. One root span per run, with a child for each model request and each tool call, so you can see order, nesting and where the time went. Subagents and handoffs appear as their own branches.
- Tool calls. The tool’s name, whether it succeeded, how long it took, and whether a person approved or refused it. Arguments and results are useful but sensitive; see the privacy section below.
- Tokens and cost. Input, output and cached tokens per request, rolled up per run, per agent and per day. Cost per completed task is the number that tells you whether a change was worth it.
- Errors. API errors, rate limits, tool failures, refusals, timeouts and retries, each attached to the span where it happened.
- Identity and context. Which agent, which model, which version of its instructions, acting for which person, on which task. Without these, a trace answers “what happened” but not “whose was it”.
The OpenTelemetry GenAI conventions
OpenTelemetry is the vendor-neutral standard for traces, metrics and logs, and it has a set of semantic conventions for generative AI, now maintained in their own OpenTelemetry GenAI repository (opens in a new tab). Check the status before you build on them: at the time of writing every GenAI convention, including the agent spans, is marked Development, which means names and attributes can still change.
The shape is already useful. An agent run is an invoke_agent span, named after the agent where its name is known. A tool call is an execute_tool span. Model requests carry attributes such as gen_ai.usage.input_tokens and gen_ai.usage.output_tokens, and there are histograms such as gen_ai.client.operation.duration. There are also conventions for MCP, so a tool call made over MCP can be traced the same way. Using these names, rather than inventing your own, means a trace from one framework can be read in any backend that understands them.
What Claude Code and the agent SDKs export
You do not have to instrument a coding agent yourself. Claude Code’s monitoring documentation (opens in a new tab) describes OpenTelemetry export switched on with environment variables. Metrics include sessions, tokens, cost, lines of code changed, commits, pull requests and edit-permission decisions. Log events include each prompt, API request, API error, tool result and tool decision. Traces are in beta and are off by default; when enabled, each interaction is a root span with the model requests, tool calls and hooks beneath it. The documentation notes that its cost figures are approximations, and that billing data comes from your API provider.
export CLAUDE_CODE_ENABLE_TELEMETRY=1 export OTEL_METRICS_EXPORTER=otlp export OTEL_LOGS_EXPORTER=otlp export OTEL_EXPORTER_OTLP_PROTOCOL=grpc export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317 # Traces are beta and need both of these as well export CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1 export OTEL_TRACES_EXPORTER=otlp
The Claude Agent SDK runs the Claude Code CLI as a child process and, according to its observability guide (opens in a new tab), produces no telemetry of its own: the same variables are passed through, and the CLI exports the same metrics, events and spans. The guide also warns that export errors fail silently by default, so check that data is actually arriving before you rely on it.
The OpenAI Agents SDK takes a different default. Its tracing documentation (opens in a new tab) says tracing is on by default and records model generations, tool calls, handoffs, guardrails and custom events, sent to OpenAI’s Traces dashboard unless you replace the default trace processors with your own. It can be switched off with OPENAI_AGENTS_DISABLE_TRACING=1, and the documentation notes it is unavailable to organisations using OpenAI’s APIs under a Zero Data Retention policy.
Privacy of logged prompts
A trace that includes prompts and tool results is a copy of everything the agent read: source code, customer emails, internal documents, sometimes a secret it should not have seen. Decide what goes into it before you switch it on. The GenAI span conventions (opens in a new tab) treat instructions, inputs and outputs as sensitive: instrumentations should not capture them by default, should offer an opt-in, and for production the recommended pattern is to store content elsewhere under its own access controls and record only a reference on the span.
- Know the defaults. Claude Code redacts prompt text, tool arguments and tool content unless you set
OTEL_LOG_USER_PROMPTS,OTEL_LOG_TOOL_DETAILSorOTEL_LOG_TOOL_CONTENT. The OpenAI Agents SDK captures model inputs and outputs by default unlesstrace_include_sensitive_datais turned off. - Turn content on where the storage is approved for it, such as a test environment, and leave it off elsewhere.
- Give traces the same retention and access rules as the data inside them. A trace store that anyone in engineering can search is a copy of every document your agents read.
- Never rely on redaction to protect a secret. Keep secrets out of what the agent can read in the first place.
Alerting
Dashboards are for investigating; alerts are for the few things someone must act on today. For agents, a short list covers most of it:
- Cost per run or per day above a set ceiling, which usually means a loop.
- A run with more tool calls or turns than any normal task needs.
- A rise in tool errors, API errors or refusals compared with last week.
- Any denied permission or blocked command, which is either a setting that is too tight or an agent attempting something it should not.
- An outbound request to a domain the agent has never contacted before, a possible sign of indirect prompt injection.
- Silence: an agent scheduled to run that sent no telemetry, which may mean the exporter failed rather than that nothing happened.
The same traces feed evaluation. A failed run found in production becomes a test task, as described in AI agent evaluation.
Machine traces are not the record people read
A trace is written for a developer with a query language and a reason to look. It is long, it is full of retries and token counts, it is often redacted, and it expires on the retention schedule you chose. Your team needs something else: a short account, in words, of what each agent changed in the tools you share, signed with its name and the person it acted for, kept by the tool and not by the agent, and kept after its access is revoked. That is an audit trail, and what a good one contains is in an audit trail for AI agents.
Keep both, and link them. Put the task reference in the agent’s instructions so it appears in the prompt and so in the trace, and put the commit or run id in the comment the agent leaves on the task. Then a question in either direction takes a minute: from a strange trace to the task it was working on, or from a task that looks wrong to the run that did it. Reading both on a schedule is part of how to audit AI agents.
Where the board fits
On fenbs the people-readable side is built in for the board. Every change is recorded in History with who made it, an assistant’s changes read “Claude via” the person it acts for, and History can be filtered to AI Assistants only. Each task shows its own trail, the assistant’s comments say what it did and name the commit, and the Testing section records what was checked. Revoking an assistant’s token stops it at once and leaves everything it did in History. None of that replaces traces of the agent’s own loop; it is the part of the record your whole team can read.
Related
The human record: an audit trail for AI agents. The monthly check that reads it: how to audit AI agents. When the traces show something has gone wrong: AI agent incident response. Connecting an assistant with its own token: the connection guide.