Claude API: A Plain Guide to Getting Started
What the Claude API is, what you need before the first call, how a Messages API request and response look, which model to pick, the SDKs, rate limits in plain words, tool use and the MCP connector.
8 min read
The Claude API is Anthropic’s REST API for calling Claude models from your own code. It lives at https://api.anthropic.com, and almost everything goes through one endpoint, POST /v1/messages: you send a model name, a max_tokens cap and a list of messages, and Claude sends back its reply as JSON. To start you need a Claude Console account at platform.claude.com and an API key. Anthropic publishes official SDKs in seven languages, limits usage by tier, and lets a request call your own tools or a remote MCP server. This guide covers each of those in plain words, as Anthropic’s documentation describes them as of September 30, 2026.
Getting and storing the key itself has its own guide: how to get a Claude API key. If what you want is an agent that reads files and runs commands rather than a single request and reply, see the Claude Agent SDK, which wraps this API in Claude Code’s agent loop.
What you need before the first call
- A Claude Console account. The Console is where you create keys, add team members, set up billing and try prompts in the playground.
- Separate billing. Anthropic’s help article on accessing the Claude API (opens in a new tab) is plain about it: a Pro, Max, Team or Enterprise plan does not include the API, and you pay separately to use the API and the Console.
- An API key, created under Settings, API keys. It starts with
sk-ant-and is shown once. - The key in an environment variable named
ANTHROPIC_API_KEY. Every official SDK reads it automatically, so it never needs to appear in your code.
Your first request
curl https://api.anthropic.com/v1/messages \
-H "content-type: application/json" \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-d '{
"model": "claude-opus-5-5",
"max_tokens": 1000,
"messages": [
{"role": "user", "content": "Summarize the difference between a feature and a bug in two sentences."}
]
}'Three headers matter. anthropic-version is required on every request, and 2023-06-01 is the value the docs use. content-type is JSON. The key goes in a header too, and here Anthropic’s own pages differ slightly: the quickstart sends it as x-api-key, while the API overview (opens in a new tab) lists Authorization: Bearer <key> first and calls x-api-key a legacy fallback that is still supported. Both work; the SDKs handle it for you.
The reply is a message object. Its content is a list of blocks, usually one text block; stop_reason says why Claude stopped, such as end_turn when it finished or tool_use when it wants to call a tool; and usage gives the input and output token counts you are billed on. To continue a conversation, you append Claude’s reply and the next user message to messages and send the whole list again.
Which model to call
Anthropic’s models overview (opens in a new tab) lists four current models and says to start with Claude Opus 5.5 for most workloads:
claude-opus-5-5: long-running agentic coding and knowledge work, the suggested default.claude-fable-5-1: demanding reasoning and long-horizon agentic work, or when Opus 5.5 at higher effort still falls short. It is the slowest of the four.claude-sonnet-5-5: described as the best combination of speed and intelligence.claude-haiku-4-5: the fastest. Its alias resolves to the dated IDclaude-haiku-4-5-20251001.
Each model ID is a pinned snapshot, so a model does not change under you until you change the ID. Older models stay available for a while under their own IDs; the deprecations page lists retirement dates. For a side-by-side of the two middle tiers in practice, see Claude Sonnet vs Opus.
The Claude API context window
The context window is everything a request carries: system prompt, tool definitions, the conversation so far and the reply. Fable 5.1, Opus 5.5 and Sonnet 5.5 each take 1M tokens, which the docs put at roughly 555,000 words on the current tokenizer, and can write up to 128K tokens in one synchronous reply. Haiku 4.5 takes 200K tokens and writes up to 64K. The Models API, GET /v1/models, returns each model’s max_input_tokens and max_tokens, so code can check instead of hard-coding them.
SDKs and the command line
The official client SDKs cover Python, TypeScript, C#, Go, Java, PHP and Ruby. They send the headers, retry on errors, stream replies and give you typed objects. Anthropic also ships ant, a command-line tool for shell scripts. The Python version of the request above:
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
message = client.messages.create(
model="claude-opus-5-5",
max_tokens=1000,
messages=[{"role": "user", "content": "Summarize the difference between a feature and a bug in two sentences."}],
)
for block in message.content:
if block.type == "text":
print(block.text)TypeScript is the same shape with npm install @anthropic-ai/sdk and new Anthropic(). Besides Messages, the API has a Message Batches API for large asynchronous jobs at half the price, a token counting endpoint, a Files API and a Skills API. Claude Managed Agents, which runs agent sessions in Anthropic’s own sandboxes, is a separate set of endpoints in beta.
Rate limits, in plain words
Limits belong to your organization, not to a key. Anthropic’s rate limits page (opens in a new tab) describes two kinds, and the Console’s Rate limits page shows your own numbers:
- Usage tiers. Your organization is placed on a tier automatically and moves up as it builds usage history. New organizations may start on a lower Evaluation tier.
- Rate limits per model, counted as requests per minute, input tokens per minute and output tokens per minute. They refill continuously, token-bucket style, rather than resetting on the minute, so a short burst can trip a limit.
- Going over returns a
429error naming the limit, with aretry-afterheader saying how many seconds to wait. A sudden jump in traffic can also hit an acceleration limit, so ramp up gradually. - For most models, input read from the prompt cache does not count toward the input limit, and
max_tokensdoes not count toward the output limit; only tokens actually generated do. - Spend limits. Each tier has a monthly spend cap, and you can set a lower limit of your own. Hitting the tier cap also returns
429, but withoutretry-after, and retrying does not help until the month turns or the cap is raised. - Workspaces can have their own lower limits, which keeps a test project from starving production.
Tool use
Tool use, also called function calling, lets Claude ask your code to do something. You pass tools, each with a name, a description and an input_schema. When Claude wants one, the reply ends with stop_reason: "tool_use" and a tool_use block holding the arguments; your code runs the function and sends the output back in a tool_result block, and Claude carries on. Anthropic’s tool use overview (opens in a new tab) separates these client tools, which run in your application, from server tools such as web search, web fetch and code execution, which run on Anthropic’s side. tool_choice can force a tool call, and strict: true makes Claude’s arguments match your schema exactly.
How this relates to MCP, which standardizes the same idea across apps, is covered in MCP vs function calling.
The MCP connector
The API can also act as the MCP client for you. With the beta header mcp-client-2025-11-20, a request lists remote servers in mcp_servers and enables their tools with an mcp_toolset entry in tools, and Claude calls those tools without you writing any client code. The MCP connector documentation (opens in a new tab) lists its limits: only tool calls are supported from the MCP feature set, the server must be reachable over HTTP, so a local stdio server cannot be connected, and for servers that need OAuth you obtain and refresh the access token yourself and pass it as authorization_token. It is in beta on the Claude API, Claude Platform on AWS and Microsoft Foundry, and not available on Amazon Bedrock or Google Cloud.
Using the API with Claude Code
Claude Code can run on an API key instead of a Claude subscription: set ANTHROPIC_API_KEY and it bills the key’s Console organization. That trade-off, how Claude Code asks you to approve the key, and how to test that a key works are in how to get a Claude API key.
Letting an API call update your board
The MCP connector is a short path from a script to a task board. fenbs is a remote MCP server at https://fenbs.ai/api/mcp, and a token you issue by hand under Settings, “Connect an AI assistant”, with a name, the scopes it needs and an optional expiry, goes straight into authorization_token. A nightly job can then read the board, add a bug it found, or comment on the task it checked:
{
"model": "claude-opus-5-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "List the bugs in Next Up and add a comment to any that mention checkout."}],
"mcp_servers": [
{"type": "url", "url": "https://fenbs.ai/api/mcp", "name": "fenbs", "authorization_token": "YOUR_FENBS_TOKEN"}
],
"tools": [{"type": "mcp_toolset", "mcp_server_name": "fenbs"}]
}The token holds a role on the board like a person does, every change it makes shows in the board’s history as an AI assistant acting for you, and revoking it under Settings stops it at once. fenbs keeps to four lanes (To Do, Next Up, In Progress, Completed) and has no due dates or assignee field, so a job like this reports and files work rather than scheduling it.
Related
Keys, workspaces and rotation: how to get a Claude API key. The agent loop on top of the API: the Claude Agent SDK. The fenbs tools a request can call: MCP docs and assistant tokens and scopes. What MCP is: MCP in the glossary.