Firecrawl MCP: Scraping and Crawling From an Agent
The Firecrawl MCP server gives an AI agent tools to search the web, scrape a page, map a site and crawl it, returning clean Markdown or JSON. What each tool does, how to connect it hosted or locally, where the API key goes, what it is good for, the risks from the pages it reads, and when Playwright MCP fits better.
7 min read
Firecrawl MCP is the MCP server for Firecrawl, a web data service. It gives an AI agent tools to search the web, scrape a known URL into clean Markdown or JSON, list the URLs on a site, crawl a section of a site, and parse documents such as PDFs. The pages are fetched by Firecrawl, not by your browser, and come back as text the model can use without wading through HTML. You can connect to Firecrawl’s hosted server with a browser sign-in or an API key, or run the open-source server locally with npx. Two things need thought before you point it at the web: every page it reads is untrusted input, and a crawl is still subject to the site’s rules.
What the tools do
The firecrawl-mcp-server repository (opens in a new tab) is MIT-licensed and publishes the npm package firecrawl-mcp. The core tools:
firecrawl_scrape: one known URL, returned as Markdown or as JSON that matches a schema you give it. Firecrawl recommends JSON for most jobs because it keeps responses small.firecrawl_map: lists the URLs on a site without fetching their content. Use it to decide what to scrape.firecrawl_search: a web search from a query rather than a URL, optionally fetching the content of the results in the same call.firecrawl_crawlandfirecrawl_check_crawl_status: many pages under a site, bounded bylimit,maxDiscoveryDepthand include or exclude paths. The crawl tool waits for the job to finish before it returns.firecrawl_parse: PDFs, Word files, spreadsheets and HTML files into Markdown or JSON.firecrawl_agentandfirecrawl_agent_status: an asynchronous research job that searches and reads across several sites and returns structured data.firecrawl_interact: clicks, typing and navigation on a page before reading it.
There are more: scheduled page monitors with diffs, a developer search over issues and docs, a research index of scientific papers, and a credit usage check. With the full profile the server lists 26 tools, which is a lot of context for a model to read; the next section shows how to get fewer.
Extract has been retired
Older guides describe a firecrawl_extract tool. The Firecrawl MCP tools page (opens in a new tab) says the former Extract tool is deprecated and not part of the current tool surface. In the server code it still exists as a name, but it returns a deprecation error that points elsewhere. For structured data from a known page, call firecrawl_scrape with the JSON format and a prompt and schema; when the pages are not known, use firecrawl_agent.
Setup: hosted or local
Firecrawl’s MCP setup guide (opens in a new tab) offers three hosted connections, all over Streamable HTTP:
- Keyless, at
https://mcp.firecrawl.dev/v2/mcpwith no credential. Rate-limited, and onlyfirecrawl_search,firecrawl_scrapeandfirecrawl_parseare offered. Useful for trying it, and for keeping the tool list short. - Sign-in, at
https://mcp.firecrawl.dev/v2/mcp-oauth. Your client opens a browser, you sign in to Firecrawl and approve a team. Connections can be reviewed and revoked in Firecrawl’s MCP settings. - API key, at the same
/v2/mcpaddress with anAuthorization: Bearerheader. For unattended jobs with no browser.
claude mcp add --transport http firecrawl https://mcp.firecrawl.dev/v2/mcp-oauth # then, inside Claude Code, sign in: /mcp
In Claude on the web and in ChatGPT, Firecrawl points to its native connector in each app’s directory instead. For a local server, which you need when a client can only start a process over stdio or when you run your own Firecrawl instance, use npx. Firecrawl’s local guide asks for Node.js 22 or newer.
claude mcp add firecrawl --transport stdio \ --env FIRECRAWL_API_KEY=fc-YOUR-API-KEY -- npx -y firecrawl-mcp # or run it as a local HTTP server at http://localhost:3000/mcp env HTTP_STREAMABLE_SERVER=true FIRECRAWL_API_KEY=fc-YOUR-API-KEY npx -y firecrawl-mcp
Set FIRECRAWL_API_URL to send requests to a self-hosted Firecrawl instead of the cloud service; the key is then optional if your instance does not require one.
Firecrawl MCP in n8n
In n8n, connect Firecrawl to an AI Agent with the MCP Client Tool node. The node’s documentation (opens in a new tab) lists bearer, header and OAuth2 authentication and a Tools to Include setting, so you can pass the API key as a bearer credential and expose only scrape and map rather than all 26 tools. Firecrawl’s local guide names n8n as the kind of client that connects to the local HTTP server. How the agent itself is built, with its trigger and review steps, is in n8n AI agents with MCP.
The API key
The key is a credential that spends your team’s Firecrawl credits, so handle it like a password. Firecrawl’s README is blunt about the first two points, and the third follows from them:
- Never put the key in the server URL. URLs end up in logs, shell history and committed config files.
- Never paste the key into an agent chat. Put it in the client’s secret or header setting, or an environment variable.
- Prefer sign-in where a person is present. The client stores and refreshes the token, and you can revoke the connection from Firecrawl without rotating a key.
Give each unattended job its own key, so you can revoke one without breaking the others, and keep keys out of any .mcp.json you commit.
What people use it for
- Reading current documentation. “Scrape the migration guide for version 5 of this library and list the breaking changes that affect our code.”
- Research with sources. “Search for the three most recent FTC statements on AI disclosures and summarize each with its URL.”
- Structured data from known pages. “Scrape these ten product pages as JSON with name, list price and availability.”
- A site inventory. “Map our marketing site and list every URL under /blog that has no meta description.”
- Watching for change. A monitor on a vendor’s status or pricing terms page, with diffs, instead of someone checking by hand.
Keep crawls small. Firecrawl warns that crawl responses can be very large and exceed token limits, and that setting limit or maxDiscoveryDepth too high is the common mistake. Map first, then scrape the pages you actually need; it is cheaper in credits and in context.
The risks
Pages are untrusted input
Every page Firecrawl returns lands in the model’s context, and a page can carry text written for the model: hidden instructions, fake “notes from the user”, requests to send data somewhere. That is indirect prompt injection, and a web scraping tool is its most direct route in. Clean Markdown does not help here; it removes the HTML, not the words.
- Do not run Firecrawl in the same session as tools that send, publish, deploy or write to production, unless every one of those calls needs your approval.
- Treat what comes back as a quotation, not an instruction. Ask for summaries with source URLs, and check the ones you act on.
- Be careful with
firecrawl_interactandfirecrawl_agent: they act on pages and follow links on their own, so they are the tools an injected instruction can steer furthest.
Site terms and robots.txt
Firecrawl’s crawl documentation (opens in a new tab) says robots.txt is respected unless ignoreRobotsTxt is enabled, and that option is limited to Enterprise plans. Sites can address Firecrawl in robots.txt with the token FirecrawlAgent. Respecting robots.txt is not the whole answer, though. The robots exclusion standard, RFC 9309 (opens in a new tab), says its rules “are not a form of access authorization”, and a site’s terms of use are a separate matter again. Before you scrape a site at scale or reuse what you collect, read its terms, keep request rates modest, and do not collect personal data you have no reason to hold. If the use is commercial or the site is a competitor’s, ask counsel; this post is not legal advice.
When Playwright MCP fits better
Firecrawl is built for reading the public web at scale and handing back clean text. Playwright MCP drives a real browser on your own machine. Choose Playwright when:
- The page is your own app, on
localhostor a staging network, that a hosted service cannot reach. - You need a logged-in session with a test account, kept in a browser you control.
- The goal is testing: clicking through a flow step by step, checking what the page shows, and turning it into a Playwright test file.
- You want to watch the browser as it works.
Choose Firecrawl when the job is reading many public pages, turning them into Markdown or JSON, searching the web, or parsing documents, and you would rather not run a browser at all. The two sit side by side without trouble.
Turning findings into work
A crawl of your own site finds things to fix: broken links, pages with no meta description, a pricing page that disagrees with the docs. With fenbs connected as a second MCP server at https://fenbs.ai/api/mcp, the agent can file each one as a bug or an enhancement, with the URL and what it found in the note, and you sort them in To Do. Every task has a plan and a test status, and History records each change under the assistant’s name. If a person decides which sites an agent may crawl, record that as a rule on the Decisions and rules page, which every connected assistant reads before it starts. fenbs does not fetch pages or check robots.txt itself; that stays with Firecrawl.
Related
The browser alternative: Playwright MCP. Why large tool results cost so much context: MCP token usage. The general checklist: MCP security best practices. Connecting fenbs: the MCP docs.