MCP Sampling Explained: When a Server Asks the Model
Sampling lets an MCP server borrow the client’s model for a completion. How it worked, how it works under the 2026-07-28 specification, which clients support it, and why the specification now tells new servers not to use it.
7 min read
MCP sampling is the feature that lets a server ask the client to run a completion on the client’s own model: “summarise this”, “classify that”, with the client choosing the model, showing the request to the user and returning the answer. The server needs no API key of its own. Two things changed in the 2026-07-28 revision of the specification. The request no longer travels as a message the server sends on an open connection; the server returns a result marked input_required and the client retries the original call with the answer. And sampling is now deprecated: new implementations should not adopt it, and existing ones should call a model provider’s API directly instead. Few mainstream clients ever documented support, so for most servers the practical answer is to design around it.
This post covers sampling only. The three things a server offers are in MCP tools, resources and prompts, and everything else that changed in 2026-07-28 is in MCP specification changes.
What sampling was designed for
The sampling page of the specification (opens in a new tab) describes it as a way for servers to request completions “with no server API keys necessary”, while the client keeps control of model access, selection and permissions. A server builds a list of messages, much like a chat request, and the client decides what actually happens to it.
- Model choice stays with the client. The server can only express preferences:
costPriority,speedPriorityandintelligencePriorityfrom 0 to 1, plushintssuch as a model family name, which the client may map to whatever it has. maxTokensis required, and the client must respect it. The client may modify or ignoresystemPrompt,temperature,stopSequencesandincludeContextwithout telling the server.- Content can be text, images or audio, and the result names the model that answered and why it stopped (
endTurn,maxTokens,toolUseand so on).
How it worked before 2026-07-28
In earlier revisions, a client that supported sampling declared the sampling capability during the initialize handshake. While handling a tool call, the server sent its own JSON-RPC request, sampling/createMessage, back to the client over the same connection and waited. That needed a connection that stayed open in both directions: stdio, or a Streamable HTTP session with a stream the server could write to. It is the main reason sampling was awkward to run behind a load balancer.
How it works now: input_required
The 2026-07-28 revision removed server-initiated requests altogether and replaced them with the multi round-trip requests pattern (opens in a new tab). Sampling rides on it, alongside elicitation and roots:
- The client calls a tool as usual, declaring
samplingin the client capabilities it sends in_metaon every request. - The server decides it needs the model, and instead of a final result returns
resultType: "input_required"with aninputRequestsmap. One entry holds asampling/createMessagerequest. It can add an opaquerequestStatestring to remember where it was. - The client shows the request to the user, runs it on its model if approved, and retries the same
tools/callunder a new JSON-RPC id, with the answer ininputResponsesunder the same key andrequestStateechoed back unchanged. - The server reads the answer and finishes with an ordinary result, or asks again.
{
"resultType": "input_required",
"inputRequests": {
"summary": {
"method": "sampling/createMessage",
"params": {
"messages": [{ "role": "user",
"content": { "type": "text", "text": "Summarise in three bullets: ..." } }],
"maxTokens": 300
}
}
},
"requestState": "summary-v1"
}Because each round is a separate request, any server instance can handle the retry, with no shared memory. Two rules keep that safe. A server must not put a sampling request in inputRequests unless the client declared the capability. And it must treat requestState as attacker-controlled input: if the state affects authorisation or business logic, protect it with an HMAC or authenticated encryption, and bind it to the caller, a short expiry and the original request. If the user declines, the client simply does not retry; there is no error to send back.
A working example in C#
The official C# SDK implements the new flow with an exception: throw InputRequiredException on the first call and read InputResponses on the retry. The tool below builds against version 2.2.0 of the SDK and was checked over stdio by hand: with the capability declared, the first call returned input_required, and the retry with a model answer returned it as the tool result. Every sampling type in it compiles with warning MCP9005, the SDK’s marker for the deprecation. The SDK’s sampling guide (opens in a new tab) has the older SampleAsync style, which only works on stateful connections. Setting up the server around it is in building an MCP server in C#.
[McpServerToolType]
public sealed class ReleaseTools
{
[McpServerTool(Name = "release_summarise")]
[Description("Summarise a changelog in three bullet points for a release note.")]
public static string Summarise(
McpServer server,
RequestContext<CallToolRequestParams> context,
[Description("The raw changelog text.")] string changelog)
{
// Second call: the client has retried with the model's answer.
if (context.Params!.InputResponses?.TryGetValue("summary", out var answer) is true)
{
var text = answer.Deserialize(InputResponse.CreateMessageResultJsonTypeInfo)?
.Content.OfType<TextContentBlock>().FirstOrDefault()?.Text;
return text ?? "The client returned no text.";
}
// Never send a sampling request to a client that did not offer it.
if (!server.IsMrtrSupported || server.ClientCapabilities?.Sampling is null)
return "This client does not offer sampling. Paste the changelog into the chat instead.";
// First call: ask the client's model, and stop here.
throw new InputRequiredException(
inputRequests: new Dictionary<string, InputRequest>
{
["summary"] = InputRequest.ForSampling(new CreateMessageRequestParams
{
Messages =
[
new SamplingMessage
{
Role = Role.User,
Content = [new TextContentBlock { Text = "Summarise in three bullets:\n" + changelog }],
},
],
MaxTokens = 300,
}),
},
requestState: "summary-v1");
}
}The capability check matters. In testing, the SDK returned the sampling request to a client that had not declared sampling at all, so the guard has to be in your code.
Sampling with tools
Since 2025-11-25, a sampling request can carry its own tools list and a toolChoice of auto, required or none, for clients that declare sampling.tools. The model may answer with tool-use blocks; the server runs those tools itself, appends the results and asks again, round after round, until the model answers with text. The specification asks both sides to cap the number of iterations, and suggests forcing toolChoice: none on the last one.
Deprecated: what that means in practice
The 2026-07-28 changelog (opens in a new tab) deprecates Roots, Sampling and Logging together (SEP-2577). The features remain fully functional during the deprecation window, which lasts at least twelve months from the revision’s release, but new implementations should not add them. The suggested migration for sampling is to integrate directly with model provider APIs. The includeContext values thisServer and allServers are deprecated too, and will go no later than sampling itself. If you already ship a sampling feature, it keeps working with clients that support it; if you are designing one, do not start.
Which clients support it
Support was thin even before the deprecation. From each vendor’s own documentation on 28 September 2026:
- Claude (web, desktop and mobile): not supported. Anthropic’s guide to building connectors (opens in a new tab) lists sampling among the MCP features Claude does not yet support, and says not to build features that depend on it.
- Claude Code: not documented. Its MCP pages cover tools, resources, prompts and elicitation dialogs, but not sampling.
- VS Code: its 1.101 release notes (opens in a new tab) (May 2025) added experimental sampling, with a confirmation the first time a server asks and a per-server choice of which models it may use. Its current MCP pages do not mention it.
- Cursor: its MCP page lists tools, prompts, resources, roots and elicitation as supported. Sampling is not listed.
- ChatGPT: not documented.
- MCP Inspector: routes an embedded sampling request to its Sampling panel, so you can answer one by hand while testing.
The human in the loop
The specification says there should always be a human in the loop able to deny a sampling request, and that applications should make requests easy to review, let the user see and edit the prompt before it is sent, and show the generated response before it goes back to the server. That is stricter than it looks. The server wrote the prompt, the user’s model runs it, and the user’s account pays for it, so the person approving should see all three.
Security
- The prompt is the server’s text running on your model. A malicious or compromised server can use it for prompt injection or to spend your tokens. Clients should require approval and rate-limit requests.
- Context can leak.
includeContextonce asked the client to add context from this server or all servers; a client may refuse when that would share sensitive information, and the values are now deprecated. - Returned text is untrusted in both directions. Both sides should validate message content, and a server should not treat a model’s answer as an instruction.
- State passes through the client. Protect
requestStateas described above, or a client can replay or alter it. - Tool loops need limits. With tools in sampling, cap the rounds on both sides.
The broader list of what goes wrong with MCP servers is in MCP security risks.
Use cases, and what to do instead
The classic examples were summarising a large document before returning it, classifying an item, turning a question into a query, and letting a server run a small agent loop without owning a model. Each has an alternative that works in every client today:
- Let the calling model do it. The assistant that called your tool is already a model; return the data, filtered and paginated, and let it summarise or classify in its next step.
- Call a provider yourself. If the server genuinely needs its own model call, use your own key and a provider API, which is the migration the specification names.
- Ask the person, not the model. If the server needs a decision rather than a completion, elicitation uses the same
input_requiredflow to ask the user.
Where fenbs stands
fenbs, a task board where AI assistants are members with roles, runs a tools-only MCP server at https://fenbs.ai/api/mcp: it declares the tools capability and nothing else, so it never asks your model for anything. Every task an assistant creates, moves or comments on is an explicit tool call it chose to make, and each one is recorded in History under the assistant’s name. The tool list is in the MCP docs.
Related
What servers offer: MCP tools, resources and prompts. The rest of 2026-07-28: MCP specification changes. Writing the server in .NET: building an MCP server in C#. Testing an input request by hand: the MCP Inspector.