MCP Sampling Explained: When a Server Asks the Model

Sampling lets an MCP server borrow the client’s model for a completion. How it worked, how it works under the 2026-07-28 specification, which clients support it, and why the specification now tells new servers not to use it.

7 min read

MCP sampling is the feature that lets a server ask the client to run a completion on the client’s own model: “summarise this”, “classify that”, with the client choosing the model, showing the request to the user and returning the answer. The server needs no API key of its own. Two things changed in the 2026-07-28 revision of the specification. The request no longer travels as a message the server sends on an open connection; the server returns a result marked input_required and the client retries the original call with the answer. And sampling is now deprecated: new implementations should not adopt it, and existing ones should call a model provider’s API directly instead. Few mainstream clients ever documented support, so for most servers the practical answer is to design around it.

This post covers sampling only. The three things a server offers are in MCP tools, resources and prompts, and everything else that changed in 2026-07-28 is in MCP specification changes.

What sampling was designed for

The sampling page of the specification (opens in a new tab) describes it as a way for servers to request completions “with no server API keys necessary”, while the client keeps control of model access, selection and permissions. A server builds a list of messages, much like a chat request, and the client decides what actually happens to it.

  • Model choice stays with the client. The server can only express preferences: costPriority, speedPriority and intelligencePriority from 0 to 1, plus hints such as a model family name, which the client may map to whatever it has.
  • maxTokens is required, and the client must respect it. The client may modify or ignore systemPrompt, temperature, stopSequences and includeContext without telling the server.
  • Content can be text, images or audio, and the result names the model that answered and why it stopped (endTurn, maxTokens, toolUse and so on).

How it worked before 2026-07-28

In earlier revisions, a client that supported sampling declared the sampling capability during the initialize handshake. While handling a tool call, the server sent its own JSON-RPC request, sampling/createMessage, back to the client over the same connection and waited. That needed a connection that stayed open in both directions: stdio, or a Streamable HTTP session with a stream the server could write to. It is the main reason sampling was awkward to run behind a load balancer.

How it works now: input_required

The 2026-07-28 revision removed server-initiated requests altogether and replaced them with the multi round-trip requests pattern (opens in a new tab). Sampling rides on it, alongside elicitation and roots:

  1. The client calls a tool as usual, declaring sampling in the client capabilities it sends in _meta on every request.
  2. The server decides it needs the model, and instead of a final result returns resultType: "input_required" with an inputRequests map. One entry holds a sampling/createMessage request. It can add an opaque requestState string to remember where it was.
  3. The client shows the request to the user, runs it on its model if approved, and retries the same tools/call under a new JSON-RPC id, with the answer in inputResponses under the same key and requestState echoed back unchanged.
  4. The server reads the answer and finishes with an ordinary result, or asks again.
The server’s reply to the first call (trimmed)
{
  "resultType": "input_required",
  "inputRequests": {
    "summary": {
      "method": "sampling/createMessage",
      "params": {
        "messages": [{ "role": "user",
          "content": { "type": "text", "text": "Summarise in three bullets: ..." } }],
        "maxTokens": 300
      }
    }
  },
  "requestState": "summary-v1"
}

Because each round is a separate request, any server instance can handle the retry, with no shared memory. Two rules keep that safe. A server must not put a sampling request in inputRequests unless the client declared the capability. And it must treat requestState as attacker-controlled input: if the state affects authorisation or business logic, protect it with an HMAC or authenticated encryption, and bind it to the caller, a short expiry and the original request. If the user declines, the client simply does not retry; there is no error to send back.

A working example in C#

The official C# SDK implements the new flow with an exception: throw InputRequiredException on the first call and read InputResponses on the retry. The tool below builds against version 2.2.0 of the SDK and was checked over stdio by hand: with the capability declared, the first call returned input_required, and the retry with a model answer returned it as the tool result. Every sampling type in it compiles with warning MCP9005, the SDK’s marker for the deprecation. The SDK’s sampling guide (opens in a new tab) has the older SampleAsync style, which only works on stateful connections. Setting up the server around it is in building an MCP server in C#.

ReleaseTools.cs (ModelContextProtocol 2.2.0)
[McpServerToolType]
public sealed class ReleaseTools
{
    [McpServerTool(Name = "release_summarise")]
    [Description("Summarise a changelog in three bullet points for a release note.")]
    public static string Summarise(
        McpServer server,
        RequestContext<CallToolRequestParams> context,
        [Description("The raw changelog text.")] string changelog)
    {
        // Second call: the client has retried with the model's answer.
        if (context.Params!.InputResponses?.TryGetValue("summary", out var answer) is true)
        {
            var text = answer.Deserialize(InputResponse.CreateMessageResultJsonTypeInfo)?
                .Content.OfType<TextContentBlock>().FirstOrDefault()?.Text;
            return text ?? "The client returned no text.";
        }

        // Never send a sampling request to a client that did not offer it.
        if (!server.IsMrtrSupported || server.ClientCapabilities?.Sampling is null)
            return "This client does not offer sampling. Paste the changelog into the chat instead.";

        // First call: ask the client's model, and stop here.
        throw new InputRequiredException(
            inputRequests: new Dictionary<string, InputRequest>
            {
                ["summary"] = InputRequest.ForSampling(new CreateMessageRequestParams
                {
                    Messages =
                    [
                        new SamplingMessage
                        {
                            Role = Role.User,
                            Content = [new TextContentBlock { Text = "Summarise in three bullets:\n" + changelog }],
                        },
                    ],
                    MaxTokens = 300,
                }),
            },
            requestState: "summary-v1");
    }
}

The capability check matters. In testing, the SDK returned the sampling request to a client that had not declared sampling at all, so the guard has to be in your code.

Sampling with tools

Since 2025-11-25, a sampling request can carry its own tools list and a toolChoice of auto, required or none, for clients that declare sampling.tools. The model may answer with tool-use blocks; the server runs those tools itself, appends the results and asks again, round after round, until the model answers with text. The specification asks both sides to cap the number of iterations, and suggests forcing toolChoice: none on the last one.

Deprecated: what that means in practice

The 2026-07-28 changelog (opens in a new tab) deprecates Roots, Sampling and Logging together (SEP-2577). The features remain fully functional during the deprecation window, which lasts at least twelve months from the revision’s release, but new implementations should not add them. The suggested migration for sampling is to integrate directly with model provider APIs. The includeContext values thisServer and allServers are deprecated too, and will go no later than sampling itself. If you already ship a sampling feature, it keeps working with clients that support it; if you are designing one, do not start.

Which clients support it

Support was thin even before the deprecation. From each vendor’s own documentation on 28 September 2026:

  • Claude (web, desktop and mobile): not supported. Anthropic’s guide to building connectors (opens in a new tab) lists sampling among the MCP features Claude does not yet support, and says not to build features that depend on it.
  • Claude Code: not documented. Its MCP pages cover tools, resources, prompts and elicitation dialogs, but not sampling.
  • VS Code: its 1.101 release notes (opens in a new tab) (May 2025) added experimental sampling, with a confirmation the first time a server asks and a per-server choice of which models it may use. Its current MCP pages do not mention it.
  • Cursor: its MCP page lists tools, prompts, resources, roots and elicitation as supported. Sampling is not listed.
  • ChatGPT: not documented.
  • MCP Inspector: routes an embedded sampling request to its Sampling panel, so you can answer one by hand while testing.

The human in the loop

The specification says there should always be a human in the loop able to deny a sampling request, and that applications should make requests easy to review, let the user see and edit the prompt before it is sent, and show the generated response before it goes back to the server. That is stricter than it looks. The server wrote the prompt, the user’s model runs it, and the user’s account pays for it, so the person approving should see all three.

Security

  • The prompt is the server’s text running on your model. A malicious or compromised server can use it for prompt injection or to spend your tokens. Clients should require approval and rate-limit requests.
  • Context can leak. includeContext once asked the client to add context from this server or all servers; a client may refuse when that would share sensitive information, and the values are now deprecated.
  • Returned text is untrusted in both directions. Both sides should validate message content, and a server should not treat a model’s answer as an instruction.
  • State passes through the client. Protect requestState as described above, or a client can replay or alter it.
  • Tool loops need limits. With tools in sampling, cap the rounds on both sides.

The broader list of what goes wrong with MCP servers is in MCP security risks.

Use cases, and what to do instead

The classic examples were summarising a large document before returning it, classifying an item, turning a question into a query, and letting a server run a small agent loop without owning a model. Each has an alternative that works in every client today:

  • Let the calling model do it. The assistant that called your tool is already a model; return the data, filtered and paginated, and let it summarise or classify in its next step.
  • Call a provider yourself. If the server genuinely needs its own model call, use your own key and a provider API, which is the migration the specification names.
  • Ask the person, not the model. If the server needs a decision rather than a completion, elicitation uses the same input_required flow to ask the user.

Where fenbs stands

fenbs, a task board where AI assistants are members with roles, runs a tools-only MCP server at https://fenbs.ai/api/mcp: it declares the tools capability and nothing else, so it never asks your model for anything. Every task an assistant creates, moves or comments on is an explicit tool call it chose to make, and each one is recorded in History under the assistant’s name. The tool list is in the MCP docs.

Related

What servers offer: MCP tools, resources and prompts. The rest of 2026-07-28: MCP specification changes. Writing the server in .NET: building an MCP server in C#. Testing an input request by hand: the MCP Inspector.

Questions people ask.

What is sampling in MCP?

It is a client feature that lets an MCP server ask the client to run a completion on the client’s model. The server sends messages and preferences; the client picks the model, shows the request to the user and returns the answer. The server needs no API key of its own.

Is MCP sampling deprecated?

Yes. The 2026-07-28 revision deprecated Roots, Sampling and Logging. Sampling still works during a deprecation window of at least twelve months, but new implementations should not adopt it, and existing ones should move to calling a model provider’s API directly.

How does MCP sampling work in the 2026-07-28 specification?

The server answers a tool call with a result whose resultType is input_required and whose inputRequests include a sampling/createMessage request. The client runs it, with the user’s approval, and retries the original tool call with the answer in inputResponses and any requestState echoed back.

Does Claude support MCP sampling?

Not in the Claude apps: Anthropic’s connector guide lists sampling as unsupported. Claude Code’s MCP documentation does not mention it. Design servers so they work without it.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.