Playwright MCP: Letting an AI Agent Drive the Browser for Tests
Playwright MCP is Microsoft’s MCP server that lets a coding agent open pages, click, type and read what happened, using the page’s accessibility tree. How to install it in each client, the options that matter, how it compares with the Playwright CLI and test agents, and how to keep a browser full of your sessions safe.
7 min read
Playwright MCP is Microsoft’s Model Context Protocol server for browser automation. You add it to a client such as Claude Code, VS Code, Cursor or Codex, and the agent gets tools to open a page, click, type, fill forms and read what the page now says. It reads the page as a structured accessibility snapshot rather than a picture, so it needs no vision model. It is most useful for reproducing a bug in a real browser, checking a change end to end, and drafting Playwright tests from a flow the agent has actually walked through. It also runs a real browser that can hold your logins, which is the part to set up carefully.
What the server is
The project lives in Microsoft’s playwright-mcp repository (opens in a new tab) and is published to npm as @playwright/mcp. It needs Node.js 18 or newer. It runs on your machine as a local process: your client starts it, and it starts a browser. By default the browser is headed, so you can watch it work.
The core tools are named plainly: browser_navigate, browser_snapshot, browser_click, browser_type, browser_fill_form, browser_select_option, browser_press_key, browser_wait_for, browser_tabs, browser_console_messages and browser_network_requests, among others. Further groups of tools, for network mocking, storage, DevTools recording, PDF output and test assertions, are switched on only when you ask for them.
Install it in your client
Every client takes the same standard entry: a server called playwright that runs npx @playwright/mcp@latest. Playwright’s own getting started page for MCP (opens in a new tab) gives the one-line forms for the clients that have a command for it.
# Claude Code
claude mcp add playwright npx @playwright/mcp@latest
# VS Code
code --add-mcp '{"name":"playwright","command":"npx","args":["@playwright/mcp@latest"]}'
# Codex CLI
codex mcp add playwright npx "@playwright/mcp@latest"- Cursor: Cursor Settings, then MCP, then Add new MCP Server, with the command
npx @playwright/mcp@latest. The README also has a one-click install link. - Claude Desktop and most other clients: paste the standard JSON entry into the client’s MCP configuration file. Where that file lives, and how to fix a server that will not start, is covered in the Claude Desktop MCP config guide.
- GitHub Copilot CLI: Playwright is one of its built-in servers, so there is nothing to add; see Copilot CLI and MCP.
To pass options in Claude Code, put them after a double dash so Claude Code does not try to read them as its own, as Claude Code’s MCP documentation (opens in a new tab) explains: claude mcp add playwright -- npx @playwright/mcp@latest --isolated --headless. If the server shows as failed, Claude Code MCP not working walks through the usual causes.
Snapshots, screenshots and vision mode
By default the agent sees the page through browser_snapshot: a text outline of the elements, their roles and their text, each with a reference such as ref=e5. It clicks and types by reference, so there is no guessing at coordinates, and the same snapshot works whether the button is blue or green. Microsoft’s tool description says this is better than a screenshot for acting on the page.
Screenshots still exist. browser_take_screenshot captures the page as an image for the agent or for you to look at, but the agent cannot act on it. If you need clicks by position, for a canvas, a map or a drawing tool with no accessible elements, start the server with --caps vision. That adds coordinate tools such as browser_mouse_click_xy and browser_mouse_drag_xy. Leave it off otherwise: images cost more context than text, and coordinates break when the layout moves.
The options that matter
The configuration reference (opens in a new tab) lists dozens of flags. Five decide how safe and how repeatable a session is.
--headless: run without a visible window. Headed is the default, which is what you want while you are learning what the agent does; headless suits CI.--isolated: keep the browser profile in memory and throw it away when the browser closes. Without it, the server uses a persistent profile, so cookies and logins carry over from one session to the next.--storage-stateand--user-data-dir: load a known login into an isolated session from a file, or point at a profile directory of your choosing. A dedicated test account’s storage state is far safer than your own profile.--allowed-originsand--blocked-origins: semicolon-separated lists of origins the browser may or may not request. The blocklist is checked first.--caps: turn on extra tool groups, such asvision,pdf,devtoolsortesting. Each group adds tools to the list the model has to read, so add only what the job needs.
Read the origin lists for what they are. Playwright says plainly that they and the file-access guardrail are convenience defences to catch unintended access, not a security boundary: they do not affect redirects and can be worked around deliberately.
Playwright MCP, the Playwright CLI and test agents
Microsoft now offers three ways for an agent to use Playwright, and its own guidance says which suits what. The Playwright CLI page for coding agents (opens in a new tab) recommends playwright-cli, installed with npm install -g @playwright/cli@latest and taught to the agent with playwright-cli install --skills, for coding agents such as Claude Code and GitHub Copilot. The reason is tokens: short commands avoid loading large tool schemas and full accessibility trees into the model’s context.
- Playwright MCP: best, in Microsoft’s words, for loops that benefit from persistent state and reasoning over page structure, such as exploratory automation, self-healing tests and long-running autonomous work.
- Playwright CLI with skills: best for a coding agent that also has a large codebase and test suite to hold in context, and needs the browser in short bursts.
- Playwright test agents (opens in a new tab): three agent definitions, a planner that explores the app and writes a Markdown test plan, a generator that turns the plan into test files, and a healer that runs the suite and repairs failing tests.
npx playwright init-agents --loop=claudesets them up for Claude Code;vscode,codexandopencodeare the other loops.
The general version of this choice, a protocol server or a command-line tool, is the subject of MCP vs CLI. For Playwright the practical answer is often both: MCP while you explore a flow together, and committed test files, run by npx playwright test, for everything after.
Using it to write tests
The agent can walk a flow, but the flow is only worth keeping once it is a test file in your repository that runs without the agent. A short loop gets you there.
- Start the server with
--isolated, a test account’s--storage-state, and--caps testingfor the assertion tools, such asbrowser_verify_text_visibleandbrowser_generate_locator. - Ask for one flow with a clear end: “Sign up with a new address, confirm the welcome banner appears, and stop.”
- Ask it to write the Playwright test from what it did, using the locators it generated rather than CSS selectors it guessed. Code is generated in TypeScript by default;
--codegenswitches to Python, Java or C#. - Run the test yourself with
npx playwright test. A test the agent wrote and never ran is a draft. - Ask it to break the test on purpose once, by changing the expected text, and confirm it fails. A test that cannot fail checks nothing.
Security: it is a browser with your sessions
The default persistent profile keeps every login the agent makes, and the --extension mode connects to your running Chrome or Edge with your real tabs and sessions. Either way, the agent can reach whatever that browser is signed in to. Microsoft states that Playwright MCP is not a security boundary.
- Pages are input. Text on a page the agent reads can carry instructions meant for it, which is indirect prompt injection. Point it at your own app and sites you trust.
- Use
--isolatedwith a test account, never your everyday profile, and keep production admin accounts out of it. - The core tools include
browser_run_code_unsafe, which runs arbitrary JavaScript in the server process. Playwright calls it equivalent to remote code execution; keep your client’s approval prompt on for it, or leave it out of the tools you allow. - Keep the approval prompt on for anything that submits forms against a real system. The wider list of what can go wrong with a server like this is in MCP security risks.
If what you want is Claude working in your own everyday browser, with site-by-site permissions, that is a different product: see Claude in Chrome.
Where the findings go
A browser session finds things: a console error on the checkout page, a button with no label, a form that accepts an empty email. They are worth keeping only if they land somewhere a person will see them. With fenbs connected as a second MCP server, the agent can file each one with fenbs_create_item as a bug, with the steps and the console output in the note, and you triage them in To Do later. Its changes are recorded in the board’s history as “Claude via” you, so you can tell what the agent filed from what you did.
Related
Connect a board next to Playwright: Claude Code integration and the MCP setup guide. A ready board for what the browser finds: the bug tracker template. More real server workflows: MCP examples.