Writing MCP Tool Descriptions a Model Can Use

A model picks and calls your MCP tools from their names, descriptions, schemas and error messages alone. How to write each one so it picks the right tool, fills the arguments correctly and recovers when a call fails, with before-and-after rewrites.

8 min read

A good MCP tool description tells the model what the tool does, when to call it, when to call something else instead, what comes back and what it changes, in that order, with the most important sentence first. Around it, give the tool a name prefixed with your product, arguments with unambiguous names and enums for fixed values, error results that say what to do next, honest annotations (readOnlyHint, destructiveHint, idempotentHint, openWorldHint), and an output schema where the result has a shape. Then keep the number of tools small, because every definition competes for the same attention.

The basics, with a working server, are in how to build an MCP server. This post goes one level down, into the wording. When a model already misuses a tool and you need to find out why, start with how to debug MCP tools.

What the model actually reads

The tools page of the MCP specification (opens in a new tab) defines a tool as a name, an optional display title, a description, an inputSchema, an optional outputSchema and optional annotations. The title is for people; the name, description and schemas are what the model reasons over.

Then there is what the client does with them. According to the Claude Code MCP documentation (opens in a new tab), Claude Code’s tool search, on by default, loads only tool names and server instructions at the start of a session and fetches full definitions when the model looks for them, and it cuts each description off at 2,048 characters. So the name has to be searchable on its own, and the first sentence of the description has to carry the decision.

Names: product first, then the job

  • Stay inside the specification’s rules: 1 to 128 characters, case-sensitive, only ASCII letters, digits, underscore, hyphen and dot, no spaces.
  • Prefix with your product. Names only have to be unique within one server, and the specification expects clients to hit collisions, two servers each with a search. billing_find_invoices survives next to a CRM’s find_invoices; search does not.
  • Name the job, not the endpoint. billing_find_invoices says what a person would ask for; get_v2_invoice_records says how your API happens to be laid out.
  • Be consistent. Pick one verb for one kind of action (find, create, update) and one order (product_resource_verb or product_verb_resource) across every tool.

Descriptions: when to use it, and when not

Anthropic’s best practices for tool definitions (opens in a new tab) call detailed descriptions “by far the most important factor in tool performance”: what the tool does, when it should be used and when it should not, what each parameter means, and its caveats, in at least three or four sentences. A shape that works:

  1. What it does and when to call it, in one sentence. This is the one that survives truncation.
  2. When not to call it, naming the tool to use instead. Near-duplicates are where models go wrong, so say which sibling is for which case.
  3. What comes back: the fields, the order, the page size, and what to do when there is more.
  4. What it changes, and what it needs: side effects, permissions, anything irreversible.

Write it the way you would brief a new colleague, which is how Anthropic’s engineering guide to writing tools for agents (opens in a new tab) puts it: the query format you take for granted, the meaning of an internal term, how two resources relate. Do not put instructions in a description that you would not want a model to follow on every call; it is read as instructions.

A rewrite, before and after

Before
{
  "name": "get_invoices",
  "description": "Gets invoices.",
  "inputSchema": {
    "type": "object",
    "properties": {
      "customer": { "type": "string" },
      "status": { "type": "string" },
      "from": { "type": "string" }
    }
  }
}
After
{
  "name": "billing_find_invoices",
  "title": "Find invoices",
  "description": "Find a customer's invoices by status and date, newest first. Use it to answer questions about what a customer owes or has paid. To read one invoice's lines, call billing_get_invoice with its number instead. Returns up to 20 invoices (number, date, total, status) and a nextCursor when there are more. Read-only.",
  "inputSchema": {
    "type": "object",
    "properties": {
      "customer_id": { "type": "string", "description": "The customer's id, e.g. CUS-0042. Call crm_find_customer if you only have a name." },
      "status": { "type": "string", "enum": ["draft", "open", "paid", "void"] },
      "issued_after": { "type": "string", "format": "date", "description": "Only invoices issued on or after this date, YYYY-MM-DD." },
      "cursor": { "type": "string", "description": "nextCursor from the previous page." }
    },
    "required": ["customer_id"],
    "additionalProperties": false
  },
  "annotations": { "readOnlyHint": true, "openWorldHint": false }
}

Every change answers a question the model would otherwise guess at: which customer field it wants, which statuses exist, what date format, what to call for the detail, how many come back, and whether calling it is safe.

Argument schemas and enums

  • Name arguments so they cannot be misread: customer_id, not customer, which a model will fill with a name.
  • Use enum for every fixed set. A model can only send a value it has seen, and the list itself documents the choices.
  • State formats and ranges in the schema, not just the prose: format, pattern, minimum, maximum, minLength. Say units and time zones in the description.
  • Mark what is required, give optional arguments sensible defaults, and set additionalProperties: false so a stray argument fails loudly instead of being ignored. The specification recommends it for tools that take no arguments at all.
  • Keep it flat. The 2026-07-28 revision allows any JSON Schema 2020-12 keyword, but the model provider behind a client may accept less; see MCP specification changes.

Errors a model can act on

The specification separates protocol errors (an unknown tool, a malformed request) from tool execution errors, which come back as a normal result with isError: true so the model can read them and correct itself. Input validation failures belong in the second group. An error written for a model says what was wrong, what would be right, and what to call next:

Before and after
// Before
{ "isError": true, "content": [{ "type": "text", "text": "Error 422" }] }

// After
{ "isError": true, "content": [{ "type": "text",
  "text": "No customer CUS-42. Customer ids have four digits, e.g. CUS-0042. Call crm_find_customer with the customer's name to get the id." }] }

No stack traces and no internal codes: they cost tokens and cannot be acted on. If a tool refuses for permission reasons, name the permission that is missing, so the assistant can tell its person why rather than retrying.

Annotations, and why clients read them

The four hints and their defaults, from the specification’s schema (opens in a new tab):

  • readOnlyHint: the tool does not modify its environment. Default false.
  • destructiveHint: it may make destructive updates, rather than only additive ones. Default true; only meaningful when readOnlyHint is false.
  • idempotentHint: calling it again with the same arguments has no further effect. Default false; also only meaningful for tools that write.
  • openWorldHint: it reaches an open world of external entities, such as the web, rather than a closed domain. Default true.

The defaults are the cautious reading: a tool with no annotations is presumed to write, destructively, to the open world. Clients use the hints to decide what to confirm. ChatGPT treats any tool without readOnlyHint as a write action that asks for confirmation, as MCP with OpenAI explains, and OpenAI’s app submission guidelines (opens in a new tab) require accurate values for all three of readOnlyHint, openWorldHint and destructiveHint, including openWorldHint: false for a tool limited to a private account or workspace. The specification is equally clear on the other side: annotations are hints, and clients must treat them as untrusted unless the server is trusted. Set them honestly; never use them as a security control.

Output schemas

When a result has a fixed shape, declare it in outputSchema and return the data in structuredContent. If you declare one, your results must conform to it, and for older clients the specification asks you to repeat the same JSON as text. What goes into it matters as much as the shape: Anthropic’s guide recommends high-signal fields, readable names over opaque identifiers, and pagination, filtering or truncation with sensible defaults, noting that Claude Code limits tool responses to 25,000 tokens by default. For tools that can return a lot, a response_format argument with concise and detailed lets the model ask for ids only when it needs them for the next call.

Keep the tool count low

More tools do not make a better server. Each definition is context the model has to weigh, overlapping tools get confused with each other, and a tool per endpoint pushes the work of chaining calls onto the model. Consolidate around jobs: Anthropic’s example is one schedule_event tool that finds availability and books, in place of separate tools to list users, list events and create one. Keep the list in a stable order, which the specification asks for and which helps clients cache it. What the definitions cost in Claude Code, and the settings that trim them, are in how to reduce Claude Code token usage.

A checklist before you ship

  1. Every name is prefixed, within the allowed characters, and follows one verb pattern.
  2. Every description opens with what and when, names the sibling to use instead, says what comes back and what changes.
  3. Every argument has a clear name, a type, and an enum, format or range where one applies; required fields are marked.
  4. Every refusal is a tool result with isError: true and a next step.
  5. Every tool has honest annotations, and read-only tools say so.
  6. Large results are paginated, and structured ones have an output schema.
  7. You have read the tool list cold, as the model will, and removed anything you could not choose between.

How fenbs writes its tools

fenbs, a task board where AI assistants are members with roles, has around thirty tools on its MCP server, all prefixed fenbs_. A few of their descriptions show these rules in use. fenbs_whoami opens with “Call this first — it tells you what you are allowed to do before you try.” fenbs_create_item explains its duplicate check: if a similar task exists, nothing is made, the likely matches come back with created: false, and the description says what to do with an open match and with a finished one. Priority is an integer from 1 to 10 whose description says 1 is most urgent and gives the bands. Lanes are an enum of keys (backlog, next, doing, done) that the board shows to people as To Do, Next Up, In Progress and Completed. And a refusal is a tool result with isError: true that names the missing permission and the role the caller holds. The full list is in the MCP docs.

Related

Building the server: how to build an MCP server, or in .NET, building an MCP server in C#. Finding out why a tool is misused: how to debug MCP tools. What a malicious description can do: MCP security risks.

Questions people ask.

How long should an MCP tool description be?

Long enough to say what the tool does, when to use it and when not, what it returns and what it changes. Anthropic suggests at least three or four sentences. Put the most important sentence first, because some clients truncate long descriptions; Claude Code cuts them at 2,048 characters by default.

What are MCP tool annotations?

Optional hints about a tool’s behaviour: readOnlyHint, destructiveHint, idempotentHint and openWorldHint, plus a title. Clients use them to decide what to confirm with the user. They are hints, not guarantees, and clients must treat them as untrusted unless the server is trusted.

Should an MCP tool return an error or throw one?

For anything the model could fix, such as bad input, a missing record or a refused permission, return a normal result with isError set to true and a sentence saying what to do next. Reserve protocol errors for unknown tools and malformed requests.

How many tools should an MCP server have?

As few as cover the jobs people actually ask for. Merge tools that are always called together, give each a distinct purpose, and split one only when its description no longer fits in a few sentences.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.