MCP Client, Host and Server: What Each One Does

The host is the AI app you use, a client is the connection it opens for each server, and a server offers the tools. What each part is responsible for under the current specification, examples of each, and why the split decides who can stop what.

7 min read

In MCP, the host is the AI application you use, such as Claude Desktop, Claude Code or VS Code. It holds the conversation and the model, and it decides what happens. A client is a component the host creates for each server it connects to: one client, one server, and the client carries the messages between them. A server is the program that offers something to use, as tools, resources and prompts: a file system, a database, a task board. So when you add three servers to an editor, you get one host, three clients and three servers. The distinction sounds academic until you ask who can see your conversation and who can refuse a request; then it is the whole answer.

This explainer follows the architecture page of the MCP specification (opens in a new tab), revision 2026-07-28. If MCP itself is new, start with what MCP is or MCP for beginners; for what a server offers, see tools, resources and prompts.

The picture

One host, one client per server
+------------------ Host (e.g. VS Code) ------------------+
|  conversation, model, your approvals, security policy    |
|                                                          |
|   Client A          Client B            Client C         |
+------|-----------------|-------------------|-------------+
       | stdio           | stdio             | Streamable HTTP
       v                 v                   v
  Server A          Server B            Server C
  files & git       local database      remote service
  (your machine)    (your machine)      (the internet)

The specification’s own diagram has the same shape: clients live inside the host’s process, local servers run on your machine, and remote servers sit on the internet. Each line is a separate connection. Server A has no line to Server C, and neither has a line to the conversation.

The host: the app that coordinates

The host is the container and coordinator. The specification gives it six jobs:

  • Create and manage the client instances, one per server.
  • Control which connections are allowed and when they start and stop.
  • Enforce security policies and consent requirements.
  • Handle the user’s authorisation decisions.
  • Coordinate the AI model, including any request a server makes for it.
  • Gather context from all the clients into one conversation.

Examples: Claude Desktop, Claude Code, Cursor, VS Code with GitHub Copilot, ChatGPT, Gemini CLI. The MCP site’s architecture overview (opens in a new tab) uses VS Code as its worked example of a host. When people ask “which apps support MCP?”, they are asking which hosts exist; MCP clients compared lists them.

The client: one connection, one server

An MCP client is not an app you install. It is the part of the host that talks to exactly one server. In the current revision, each client:

  • Communicates with exactly one server.
  • Attaches the protocol version and its capabilities to every request.
  • Routes protocol messages in both directions.
  • Manages subscriptions and notifications, such as a server saying its tool list changed.
  • Maintains the security boundary between its server and every other one.

The architecture overview’s example makes it concrete: when VS Code connects to the Sentry MCP server, its runtime creates a client object for that connection; when it then connects to a local file system server, it creates another. A host can even hold two clients for the same remote server, which is how one remote server serves many connections at once, while a local stdio server typically serves a single client.

So what does “MCP client” mean in everyday use? Two things. Strictly, the connection component above. Loosely, the host app itself, as in “Cursor is an MCP client”. Both are common, and neither is wrong in context, but when you read the specification or an SDK, “client” means the connection. If you are writing one, building an MCP client in Python shows it in code.

The server: focused capabilities

A server exposes resources, tools and prompts, the protocol’s primitives, and does one focused job. It can be a local process the host starts over stdio, or a remote service reached over Streamable HTTP; MCP transports explains the two. Examples: the reference file system server, a Postgres server on your laptop, Sentry’s hosted server, GitHub’s, and a task board such as fenbs at https://fenbs.ai/api/mcp.

A server can also ask for something back. It can ask the user for more information, which the protocol calls elicitation. In the current revision it does that by returning a result marked as needing input, and the client sends the original request again with the answer. Sampling, a server asking the host’s model for a completion, still exists but is deprecated as of 2026-07-28.

One request, end to end

  1. You ask the host: “What is in progress on the board?”
  2. The host has already asked each client for its server’s tool list, and shows the model the combined list.
  3. The model picks a tool. The host checks your approval settings and, if needed, asks you.
  4. The host hands the call to the client for that server, and the client sends tools/call with the protocol version and its capabilities in _meta.
  5. The server checks the caller may do this, runs the tool and returns the result.
  6. The client passes the result back; the host adds it to the conversation and the model answers you.

Notice what the server saw: one tool call with its arguments. Not your question, not the rest of the conversation, not the other servers’ results.

What changed with the 2026-07-28 revision

The roles did not change; how the client and server talk did. In the 2025-11-25 architecture (opens in a new tab), a client established one stateful session per server, and capabilities were exchanged once, during initialisation, then held for the session. Server-initiated requests, such as sampling, went from server to client over that session.

Now MCP is stateless. Every request is self-contained and carries its own protocol version and client capabilities; a client may call server/discover first to learn the server’s versions and capabilities; a server that needs input answers with a result asking for it instead of sending its own request; and change notifications arrive on a stream the client opens with subscriptions/listen. For a host, that means a client can be thin and short-lived. For a server, it means no session to lean on: anything that spans calls has to be passed back explicitly. The full list is in MCP specification changes.

Why the split matters for security

Each part can protect you from a different thing, and none can do another’s job. The design principles say servers should not be able to read the whole conversation, nor see into other servers: the full history stays with the host, each server gets only the context it needs, and cross-server interactions are controlled by the host.

  • The host is where consent lives. The specification’s security principles (opens in a new tab) say hosts must get your explicit consent before invoking any tool or exposing your data to a server, and that tool descriptions and annotations should be treated as untrusted unless the server is trusted. Approval prompts, allowlists and auto-approve settings are host features; how one host handles them is in VS Code MCP security.
  • The client is the wall between servers. Because each client talks to one server, a malicious server cannot call another server’s tools directly. It can still try through the model, by planting instructions in a tool result, which is why the host’s review of results matters; see MCP security risks.
  • The server is the last check. The host decides whether a call is made; only the server decides what that call may do. It has to authenticate the caller and check permission on every request, whatever the host allowed, because a different host with different settings can send the same call.

The specification is candid that the protocol cannot enforce any of this by itself; implementers are expected to build the consent flows and access controls. That is the practical reason to know which part is which: when you choose a setup, you are choosing who holds each of those three checks. MCP security best practices turns it into a checklist.

Where fenbs sits

fenbs, a task board where people and AI assistants are members with roles, is a server. Your host, whether Claude Code, Cursor or VS Code, creates one client for it when you add https://fenbs.ai/api/mcp and sign in. The host decides when a tool is called; fenbs decides what the call may do. Each call is checked against the scopes you ticked when you approved the connection and against your role on the board, and a refusal comes back as a tool result naming the permission that was missing, so the assistant can tell you instead of retrying. Every change is recorded in History with the assistant’s name and the person it acted for, and revoking the connection in Settings stops it without signing you out.

Related

Connect a host to fenbs: MCP docs. How scopes and roles combine: assistant tokens and scopes. Nine real servers to try: MCP examples. How MCP compares with calling an API directly: MCP vs API.

Questions people ask.

What is an MCP client?

Strictly, a component inside an AI application that holds the connection to one MCP server, sending requests with the protocol version and capabilities and passing results back. Loosely, people also call the whole application, such as Cursor or Claude Desktop, an MCP client.

Is Claude Desktop an MCP host or an MCP client?

In the specification’s terms it is a host. It holds the conversation and the model and creates one client for each server you add. In everyday speech it is often called an MCP client, meaning an app that can use MCP servers.

Can one MCP server have many clients?

Yes. Each client connects to exactly one server, but a server can serve many clients. A remote server over Streamable HTTP typically serves many at once, while a local stdio server started by a host usually serves just that one client.

Does the client or the server run the tool?

The server runs it. The model chooses a tool, the host approves the call, the client sends it, and the server executes the tool and returns the result, which the client passes back to the host for the model to read.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.