The Claude Agent SDK: What It Gives You Beyond Claude Code
The Agent SDK is Claude Code’s agent loop as a TypeScript or Python library. What it adds over the CLI, the options that matter, three short examples of different shapes, and when claude -p is the better tool.
9 min read
The Claude Agent SDK is the agent loop that runs Claude Code, packaged as a library for TypeScript and Python. You call query() with a prompt and an options object, and iterate over the messages it streams back: the same built-in tools, permission rules, hooks, subagents, MCP servers and sessions as the CLI, driven from your own code. What it gives you beyond Claude Code is control from inside your program: a callback that decides each tool call, hooks written as functions, custom tools that run in your process, typed results, and sessions you store and resume by ID. Below are the options that matter and three short examples of different shapes.
This is an overview and reference, not a first project. For a step-by-step build of a small agent, with testing, see how to build your first AI agent. For running the CLI itself unattended, see how to automate Claude Code tasks.
What you install, and what it runs
# TypeScript (Node.js 18+) npm install @anthropic-ai/claude-agent-sdk # Python (3.10+) pip install claude-agent-sdk export ANTHROPIC_API_KEY=your-api-key
Both packages bundle a native Claude Code binary, so most installs need nothing else; the SDK runs it for you. It authenticates with an API key, or with Amazon Bedrock, Google Cloud or Microsoft Foundry credentials. Anthropic’s Agent SDK overview (opens in a new tab) adds a condition for products: unless previously approved, third-party developers may not offer claude.ai login or its rate limits in agents built on the SDK.
Two defaults surprise people coming from the CLI. First, with no systemPrompt the SDK uses a minimal prompt that covers tool calling but leaves out Claude Code’s own instructions, safety guidance and environment context, whereas claude -p uses the full Claude Code prompt; the system prompt guide (opens in a new tab) says to set the claude_code preset if you want matching behaviour. Second, with no settingSources the SDK reads the same files the CLI does: your user and project settings, CLAUDE.md, skills, hooks and .claude/agents.
The options that matter
toolsdecides which built-in tools exist in the session at all.["Read", "Grep"]means nothing else is visible;[]removes every built-in, leaving only your MCP tools.allowedToolspre-approves calls; it does not restrict. A tool missing from it is still available and falls through to the permission mode.disallowedToolsblocks: a bare name such asBashremoves the tool, a scoped rule such asBash(rm *)denies matching calls in every mode.permissionModeisdefault,dontAsk,acceptEdits,plan,autoorbypassPermissions.dontAskrefuses anything not pre-approved instead of asking, which is what an unattended job wants.canUseToolis your approval function, called with the tool name and input only when nothing earlier settled the call. It returns allow, optionally with changed input, or deny with a message Claude reads.hooksare functions keyed by event:PreToolUsecan block or rewrite a call,PostToolUsesees the result, and there are events for sessions, subagents, compaction and more. Some events exist only in TypeScript.mcpServerstakes remote servers by URL, local ones by command, and in-process servers you build withcreateSdkMcpServer().agentsdefines subagents inline, each with adescriptionthat tells Claude when to use it, aprompt, and optionally its owntoolsandmodel.resume,continueandforkSessionpick up earlier work. Every result message carries asession_id; pass it toresumeto continue that conversation later.settingSourceschooses which ofuser,projectandlocalload.[]loads none of them.maxTurnsandmaxBudgetUsdare hard caps, andoutputFormattakes a JSON Schema so the result arrives as validated data instructured_output.
Python spells the same options in snake case: allowed_tools, permission_mode, can_use_tool, setting_sources, max_budget_usd, output_format.
Example 1: a read-only reviewer
A reviewer should read everything and change nothing. This one gets four tools plus the Agent tool for its subagent. Beyond git diff, git log and the read-only commands Claude Code already treats as safe, every shell command is refused instead of asked about. It reads the repository’s CLAUDE.md and rules, so it reviews against your conventions, and hands test coverage to a cheaper subagent.
import { query } from "@anthropic-ai/claude-agent-sdk";
const base = process.argv[2] ?? "main";
for await (const message of query({
prompt: `Review this branch against ${base}. List likely bugs as file:line and one sentence each. Change nothing.`,
options: {
tools: ["Read", "Grep", "Glob", "Bash", "Agent"], // everything else is out of reach
allowedTools: ["Read", "Grep", "Glob", "Bash(git diff *)", "Bash(git log *)"],
permissionMode: "dontAsk", // anything not allowed is refused, not asked
systemPrompt: { type: "preset", preset: "claude_code", append: "You are reviewing. Never edit files." },
settingSources: ["project"], // read the repository's CLAUDE.md and rules
agents: {
"test-reader": {
description: "Reads the tests that cover a changed file and says what they do not check.",
prompt: "You read tests. Report gaps in coverage for the files you are given, nothing else.",
tools: ["Read", "Grep", "Glob"],
model: "haiku",
},
},
maxTurns: 30,
},
})) {
if (message.type === "result") {
console.log(message.subtype === "success" ? message.result : `Stopped: ${message.subtype}`);
}
}The safety here is tools plus dontAsk, not the prompt. Edit and Write are not in the session, so the instruction never to edit is a second line of defence rather than the only one.
Example 2: a scheduled report
A job that runs every morning from cron should behave the same on every machine and hand back data, not prose. This one loads no settings files, has its own short system prompt, returns JSON that matches a schema, and stops at a turn cap and a spend cap.
import asyncio
import json
from claude_agent_sdk import ClaudeAgentOptions, ResultMessage, query
SCHEMA = {
"type": "object",
"properties": {
"merged": {"type": "array", "items": {"type": "string"}},
"risky": {"type": "array", "items": {"type": "string"}},
},
"required": ["merged", "risky"],
}
async def main() -> None:
options = ClaudeAgentOptions(
tools=["Read", "Bash"],
allowed_tools=["Read", "Bash(git log *)", "Bash(git show *)"],
permission_mode="dontAsk",
setting_sources=[], # same behaviour on every machine
system_prompt="You write a short daily engineering report. Facts from git only.",
output_format={"type": "json_schema", "schema": SCHEMA},
max_turns=15,
max_budget_usd=0.50,
)
try:
async for message in query(
prompt="Summarise what merged to main in the last 24 hours, and flag anything touching migrations or auth.",
options=options,
):
if isinstance(message, ResultMessage):
print(json.dumps(message.structured_output, indent=2))
print("session:", message.session_id)
except Exception as error:
print("Run ended:", error)
asyncio.run(main())The try is not decoration. The configuration guide (opens in a new tab) says a single-shot query() yields the capped result and then raises when a run hits max_turns or max_budget_usd, so a scheduled job without it ends in a stack trace. The printed session_id lets you resume the run later and ask a follow-up question with the same context.
Example 3: a custom tool with an audit hook
Custom tools are where the SDK does what the CLI cannot. This agent gets one tool of your own, the live list of feature flags from an internal service, alongside read-only file search, and looks for flags the code still checks that no longer exist. A PostToolUse hook writes one audit line for every tool call.
import { query, tool, createSdkMcpServer, type HookCallback } from "@anthropic-ai/claude-agent-sdk";
import { appendFile } from "node:fs/promises";
// A custom tool: the live list of feature flags from your own service.
const listFlags = tool(
"list_flags",
"List every feature flag that currently exists, with its state.",
{},
async () => {
const res = await fetch(process.env.FLAGS_URL + "/flags");
if (!res.ok) {
return { content: [{ type: "text", text: `Flag service returned ${res.status}. Try again later.` }], isError: true };
}
return { content: [{ type: "text", text: await res.text() }] };
},
{ annotations: { readOnlyHint: true } }
);
const flags = createSdkMcpServer({ name: "flags", version: "1.0.0", tools: [listFlags] });
// A hook: one audit line for every tool call, whatever it was.
const audit: HookCallback = async (input) => {
if (input.hook_event_name === "PostToolUse") {
await appendFile("agent-audit.log", `${new Date().toISOString()} ${input.tool_name}\n`);
}
return {};
};
for await (const message of query({
prompt: "Find flags the code still checks that no longer exist in the flag service. List file:line for each.",
options: {
tools: ["Read", "Grep", "Glob"],
mcpServers: { flags },
allowedTools: ["Read", "Grep", "Glob", "mcp__flags__list_flags"],
permissionMode: "dontAsk",
hooks: { PostToolUse: [{ hooks: [audit] }] },
maxTurns: 25,
maxBudgetUsd: 1,
},
})) {
if (message.type === "result" && message.subtype === "success") console.log(message.result);
}The tool is named mcp__flags__list_flags to Claude: server name, then tool name. When the service fails, the tool returns isError: true with a sentence of its own. According to the SDK’s custom tools guide (opens in a new tab), an exception thrown in a handler does not stop the run either, but Claude then sees only the raw exception message; composing the error yourself tells it what failed and what to try. A hook without a matcher runs for every tool.
Three traps
- Anything approved earlier never reaches
canUseTool. The permissions guide (opens in a new tab) is explicit: a call approved by an allow rule,acceptEditsorbypassPermissionsskips your callback, so a check you put there is silently bypassed for those tools. For a check that must run on every call, use aPreToolUsehook. allowedToolsdoes not limitbypassPermissions.allowedTools: ["Read"]withbypassPermissionsstill approvesBash,WriteandEdit. Block tools withdisallowedToolsor leave them out oftools.- The default
settingSourcesloads whatever is on the machine, including user settings and hooks. For a reproducible job, or one serving several customers, pass[]and configure everything in code.
The SDK or claude -p?
- Use
claude -pwhen the job is a shell step: a prompt in, text or JSON out, permissions fixed in advance. It fits in a CI file and needs no code, and the same run loads your usual Claude Code setup. - Use the SDK when your program must take part while the agent runs: approving calls one by one, rewriting tool input in a hook, providing a tool that lives in your process, reacting to each message as it streams, or keeping sessions per user.
- Use the SDK when the agent is part of a product. Its identity, system prompt, tools and storage are then yours to decide, rather than inherited from a developer’s terminal.
- For languages other than TypeScript and Python, the documented route is the CLI as a subprocess with
-pand--output-format json.
Put the results where people look
An agent that prints a report to a log has done half the job. The other half is putting the finding on the task it concerns. fenbs is an MCP server, so an SDK agent reaches it through mcpServers with a token you issue by hand under Settings, “Connect an AI assistant”, with a name, the scopes it needs and an optional expiry. Pre-approve mcp__fenbs__fenbs_comment and the reviewer in Example 1 can comment its findings on the bug it was checking, recorded in the board’s history as an AI assistant acting for you. Revoking the token stops it at once. The exact configuration is step 7 of how to build your first AI agent.
Related
Claude Code’s own subagents, defined as files: Claude Code subagent examples. The same hooks in the CLI: Claude Code hooks. The fenbs tools an agent can call: MCP docs and assistant tokens and scopes.