Claude API: A Plain Guide to Getting Started

What the Claude API is, what you need before the first call, how a Messages API request and response look, which model to pick, the SDKs, rate limits in plain words, tool use and the MCP connector.

8 min read

The Claude API is Anthropic’s REST API for calling Claude models from your own code. It lives at https://api.anthropic.com, and almost everything goes through one endpoint, POST /v1/messages: you send a model name, a max_tokens cap and a list of messages, and Claude sends back its reply as JSON. To start you need a Claude Console account at platform.claude.com and an API key. Anthropic publishes official SDKs in seven languages, limits usage by tier, and lets a request call your own tools or a remote MCP server. This guide covers each of those in plain words, as Anthropic’s documentation describes them as of September 30, 2026.

Getting and storing the key itself has its own guide: how to get a Claude API key. If what you want is an agent that reads files and runs commands rather than a single request and reply, see the Claude Agent SDK, which wraps this API in Claude Code’s agent loop.

What you need before the first call

  • A Claude Console account. The Console is where you create keys, add team members, set up billing and try prompts in the playground.
  • Separate billing. Anthropic’s help article on accessing the Claude API (opens in a new tab) is plain about it: a Pro, Max, Team or Enterprise plan does not include the API, and you pay separately to use the API and the Console.
  • An API key, created under Settings, API keys. It starts with sk-ant- and is shown once.
  • The key in an environment variable named ANTHROPIC_API_KEY. Every official SDK reads it automatically, so it never needs to appear in your code.

Your first request

Terminal
curl https://api.anthropic.com/v1/messages \
  -H "content-type: application/json" \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "claude-opus-5-5",
    "max_tokens": 1000,
    "messages": [
      {"role": "user", "content": "Summarize the difference between a feature and a bug in two sentences."}
    ]
  }'

Three headers matter. anthropic-version is required on every request, and 2023-06-01 is the value the docs use. content-type is JSON. The key goes in a header too, and here Anthropic’s own pages differ slightly: the quickstart sends it as x-api-key, while the API overview (opens in a new tab) lists Authorization: Bearer <key> first and calls x-api-key a legacy fallback that is still supported. Both work; the SDKs handle it for you.

The reply is a message object. Its content is a list of blocks, usually one text block; stop_reason says why Claude stopped, such as end_turn when it finished or tool_use when it wants to call a tool; and usage gives the input and output token counts you are billed on. To continue a conversation, you append Claude’s reply and the next user message to messages and send the whole list again.

Which model to call

Anthropic’s models overview (opens in a new tab) lists four current models and says to start with Claude Opus 5.5 for most workloads:

  • claude-opus-5-5: long-running agentic coding and knowledge work, the suggested default.
  • claude-fable-5-1: demanding reasoning and long-horizon agentic work, or when Opus 5.5 at higher effort still falls short. It is the slowest of the four.
  • claude-sonnet-5-5: described as the best combination of speed and intelligence.
  • claude-haiku-4-5: the fastest. Its alias resolves to the dated ID claude-haiku-4-5-20251001.

Each model ID is a pinned snapshot, so a model does not change under you until you change the ID. Older models stay available for a while under their own IDs; the deprecations page lists retirement dates. For a side-by-side of the two middle tiers in practice, see Claude Sonnet vs Opus.

The Claude API context window

The context window is everything a request carries: system prompt, tool definitions, the conversation so far and the reply. Fable 5.1, Opus 5.5 and Sonnet 5.5 each take 1M tokens, which the docs put at roughly 555,000 words on the current tokenizer, and can write up to 128K tokens in one synchronous reply. Haiku 4.5 takes 200K tokens and writes up to 64K. The Models API, GET /v1/models, returns each model’s max_input_tokens and max_tokens, so code can check instead of hard-coding them.

SDKs and the command line

The official client SDKs cover Python, TypeScript, C#, Go, Java, PHP and Ruby. They send the headers, retry on errors, stream replies and give you typed objects. Anthropic also ships ant, a command-line tool for shell scripts. The Python version of the request above:

quickstart.py (pip install anthropic)
import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

message = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=1000,
    messages=[{"role": "user", "content": "Summarize the difference between a feature and a bug in two sentences."}],
)

for block in message.content:
    if block.type == "text":
        print(block.text)

TypeScript is the same shape with npm install @anthropic-ai/sdk and new Anthropic(). Besides Messages, the API has a Message Batches API for large asynchronous jobs at half the price, a token counting endpoint, a Files API and a Skills API. Claude Managed Agents, which runs agent sessions in Anthropic’s own sandboxes, is a separate set of endpoints in beta.

Rate limits, in plain words

Limits belong to your organization, not to a key. Anthropic’s rate limits page (opens in a new tab) describes two kinds, and the Console’s Rate limits page shows your own numbers:

  • Usage tiers. Your organization is placed on a tier automatically and moves up as it builds usage history. New organizations may start on a lower Evaluation tier.
  • Rate limits per model, counted as requests per minute, input tokens per minute and output tokens per minute. They refill continuously, token-bucket style, rather than resetting on the minute, so a short burst can trip a limit.
  • Going over returns a 429 error naming the limit, with a retry-after header saying how many seconds to wait. A sudden jump in traffic can also hit an acceleration limit, so ramp up gradually.
  • For most models, input read from the prompt cache does not count toward the input limit, and max_tokens does not count toward the output limit; only tokens actually generated do.
  • Spend limits. Each tier has a monthly spend cap, and you can set a lower limit of your own. Hitting the tier cap also returns 429, but without retry-after, and retrying does not help until the month turns or the cap is raised.
  • Workspaces can have their own lower limits, which keeps a test project from starving production.

Tool use

Tool use, also called function calling, lets Claude ask your code to do something. You pass tools, each with a name, a description and an input_schema. When Claude wants one, the reply ends with stop_reason: "tool_use" and a tool_use block holding the arguments; your code runs the function and sends the output back in a tool_result block, and Claude carries on. Anthropic’s tool use overview (opens in a new tab) separates these client tools, which run in your application, from server tools such as web search, web fetch and code execution, which run on Anthropic’s side. tool_choice can force a tool call, and strict: true makes Claude’s arguments match your schema exactly.

How this relates to MCP, which standardizes the same idea across apps, is covered in MCP vs function calling.

The MCP connector

The API can also act as the MCP client for you. With the beta header mcp-client-2025-11-20, a request lists remote servers in mcp_servers and enables their tools with an mcp_toolset entry in tools, and Claude calls those tools without you writing any client code. The MCP connector documentation (opens in a new tab) lists its limits: only tool calls are supported from the MCP feature set, the server must be reachable over HTTP, so a local stdio server cannot be connected, and for servers that need OAuth you obtain and refresh the access token yourself and pass it as authorization_token. It is in beta on the Claude API, Claude Platform on AWS and Microsoft Foundry, and not available on Amazon Bedrock or Google Cloud.

Using the API with Claude Code

Claude Code can run on an API key instead of a Claude subscription: set ANTHROPIC_API_KEY and it bills the key’s Console organization. That trade-off, how Claude Code asks you to approve the key, and how to test that a key works are in how to get a Claude API key.

Letting an API call update your board

The MCP connector is a short path from a script to a task board. fenbs is a remote MCP server at https://fenbs.ai/api/mcp, and a token you issue by hand under Settings, “Connect an AI assistant”, with a name, the scopes it needs and an optional expiry, goes straight into authorization_token. A nightly job can then read the board, add a bug it found, or comment on the task it checked:

Messages API request body (beta header mcp-client-2025-11-20)
{
  "model": "claude-opus-5-5",
  "max_tokens": 1024,
  "messages": [{"role": "user", "content": "List the bugs in Next Up and add a comment to any that mention checkout."}],
  "mcp_servers": [
    {"type": "url", "url": "https://fenbs.ai/api/mcp", "name": "fenbs", "authorization_token": "YOUR_FENBS_TOKEN"}
  ],
  "tools": [{"type": "mcp_toolset", "mcp_server_name": "fenbs"}]
}

The token holds a role on the board like a person does, every change it makes shows in the board’s history as an AI assistant acting for you, and revoking it under Settings stops it at once. fenbs keeps to four lanes (To Do, Next Up, In Progress, Completed) and has no due dates or assignee field, so a job like this reports and files work rather than scheduling it.

Related

Keys, workspaces and rotation: how to get a Claude API key. The agent loop on top of the API: the Claude Agent SDK. The fenbs tools a request can call: MCP docs and assistant tokens and scopes. What MCP is: MCP in the glossary.

Questions people ask.

Is the Claude API included in a Claude Pro or Max plan?

No. Anthropic says Pro, Max, Team and Enterprise plans do not include the API. You create a Claude Console account at platform.claude.com, set up billing there and pay for API usage separately.

What is the Claude API context window?

As of September 30, 2026, Claude Fable 5.1, Claude Opus 5.5 and Claude Sonnet 5.5 each accept 1M tokens of context and write up to 128K tokens per reply. Claude Haiku 4.5 accepts 200K tokens and writes up to 64K. The Models API reports these limits for each model.

Which languages have an official Claude SDK?

Python, TypeScript, C#, Go, Java, PHP and Ruby, plus the ant command-line tool. Each SDK reads the ANTHROPIC_API_KEY environment variable and handles headers, retries and streaming.

Can the Claude API connect to an MCP server?

Yes, through the MCP connector, which is in beta. A Messages API request names remote servers in mcp_servers and enables their tools. Only tool calls are supported, the server must be reachable over HTTP, and you supply any OAuth token yourself.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.