Grafana MCP Server: Dashboards and Alerts for an Agent
Grafana Labs maintains mcp-grafana, an open source MCP server for dashboards, datasources, Prometheus and Loki queries, alerting, incidents and OnCall, and runs a hosted Cloud MCP server for Grafana Cloud. What the tools cover, how to scope the service account token, how to run it read-only, and how to set it up in Claude Code and Cursor.
7 min read
The Grafana MCP server, mcp-grafana, is Grafana Labs’ open source MCP server. It lets an AI agent search and read dashboards, list datasources, run PromQL and LogQL queries, read alert rules, and work with Grafana Incident, OnCall and Sift investigations, against your own Grafana or a Grafana Cloud stack. It signs in with a Grafana service account token, so the agent can do exactly what that service account can do. For Grafana Cloud there is also a hosted server at https://mcp.grafana.com/mcp with a browser sign-in and nothing to run. For most agent work the right setup is narrow: a service account with read permissions, the server started with --disable-write, and only the tool categories the job needs.
Two ways to run it
- Open source, self-run: the mcp-grafana repository (opens in a new tab) is Apache 2.0 and works with self-hosted Grafana 9.0 or later and with Grafana Cloud. You start it with
uvx mcp-grafana, a Docker image, a binary or a Helm chart. It runs over stdio by default, or as an SSE or Streamable HTTP server. - Grafana Cloud MCP: a server hosted by Grafana Labs, for Grafana Cloud stacks only. You point your client at one URL and authorize in the browser; there is no token to create and nothing to deploy.
What the tools cover
The README lists the tools by category. The ones an agent uses most:
- Dashboards:
search_dashboards,get_dashboard_summaryfor a compact overview,get_dashboard_propertyto pull one part with a JSONPath, andget_dashboard_panel_queriesfor each panel’s query and datasource.update_dashboardwrites. The README warns that full dashboard JSON can use a lot of context and recommends the summary and property tools instead. - Datasources:
list_datasourcesandget_datasource. - Prometheus:
query_prometheusfor instant and range queries, metric and label discovery, andquery_prometheus_histogramfor percentiles. - Loki:
query_loki_logsfor log and metric queries in LogQL, label discovery,query_loki_statsandquery_loki_patternsfor common log shapes. - Alerting:
alerting_manage_ruleslists rules and their state and can create, update and delete them;alerting_manage_routingshows notification policies and contact points;alerting_manage_silenceshandles silences. - Incidents and OnCall: search, create and update incidents in Grafana Incident; read on-call schedules, shifts and who is on call now; list and read alert groups, with
update_alert_groupas the write. - Sift: read investigations, and
find_error_pattern_logsandfind_slow_requeststo start one. - Navigation:
generate_deeplinkbuilds real links to dashboards, panels and Explore instead of letting the model guess URLs.
Several categories are off by default and turned on by name with --enabled-tools: SQL datasources, InfluxDB, CloudWatch, Elasticsearch, Graphite, admin tools, running panel queries, and the Grafana Assistant. Others are on and can be turned off with a --disable- flag, such as --disable-oncall. Fewer tools means less context spent on every turn; MCP token usage explains the cost.
The service account token and least privilege
The open source server authenticates with a token from a Grafana service account, set in GRAFANA_SERVICE_ACCOUNT_TOKEN or read from a file named in GRAFANA_SERVICE_ACCOUNT_TOKEN_FILE, which is re-read on every request so a rotated token is picked up without a restart. The older GRAFANA_API_KEY variable is deprecated. Grafana’s service accounts documentation (opens in a new tab) covers creating the account and its token.
The README lists the RBAC permission and scope each tool needs, and it is candid about the shortcut: assigning the built-in Editor role “grants broad read/write access”, which it recommends only when convenience matters more than least privilege. For an agent, take the longer route:
- Start from read permissions:
dashboards:read,datasources:read, anddatasources:queryfor the tools that run queries. - Scope queries to the datasources the job needs, such as
datasources:uid:prometheus-prodanddatasources:uid:loki-prod, rather thandatasources:*. - Incident and Sift tools use basic roles instead: Viewer for reading, Editor for creating or changing.
- Give each agent or job its own service account and token, with an expiry, so you can revoke one without affecting the rest.
- Log lines are written by whoever can reach your services. Treat them as untrusted input to the model; indirect prompt injection explains why.
Read-only use
Permissions limit what the token can do; the server’s own flags limit which tools the agent even sees. Start the server with --disable-write and the write tools are not registered at all: updating dashboards, creating folders, creating and updating incidents, managing alert rules and silences, updating alert groups, annotations, snapshots, and the Sift tools that start investigations. Queries in PromQL, LogQL, TraceQL and the other languages that cannot express a write stay available.
Two refinements are worth knowing. Read-only mode also removes the raw SQL and InfluxDB query tools, because the README notes that query_sql will run a DROP TABLE if the datasource credentials allow it; --enable-query brings them back when those credentials are read-only. And --disable-query goes further, removing every tool that runs a query while keeping discovery tools, for an agent that should see what exists without reading the data. For Loki, a --loki-guardrail-mode setting, off by default, guards against broad queries that scan huge amounts of data.
Setup in Claude Code
The self-run server needs uv installed; Grafana’s open source MCP server docs (opens in a new tab) also cover Docker, binary and Helm installs. Add it as a stdio server with the URL and token in its environment, and read-only flags after the command:
claude mcp add grafana \
--env GRAFANA_URL=https://myinstance.grafana.net \
--env 'GRAFANA_SERVICE_ACCOUNT_TOKEN=${GRAFANA_SERVICE_ACCOUNT_TOKEN}' \
-- uvx mcp-grafana --disable-writeFor a local Grafana, use http://localhost:3000 as the URL. For the hosted Cloud server, add the URL over HTTP and name your stack in a header; the agent prompts you to authorize in the browser the first time it needs Grafana data.
claude mcp add grafana https://mcp.grafana.com/mcp --transport http \ --header "X-Grafana-URL: https://<your-stack>.grafana.net"
Setup in Cursor
In .cursor/mcp.json, the self-run server goes in as a command. The hosted server goes in as a url with the same X-Grafana-URL header.
{
"mcpServers": {
"grafana": {
"command": "uvx",
"args": ["mcp-grafana", "--disable-write"],
"env": {
"GRAFANA_URL": "https://myinstance.grafana.net",
"GRAFANA_SERVICE_ACCOUNT_TOKEN": "<your service account token>"
}
}
}
}That file holds a live token, so keep it out of the repository. If your client lists the server but it will not connect, Claude Code MCP not working covers the usual causes.
Grafana Cloud MCP, the hosted option
The Grafana Cloud MCP documentation (opens in a new tab) describes a server that works only with hosted Grafana Cloud, over Streamable HTTP only; SSE is not supported. The details that matter for permissions:
- Sign-in is OAuth 2.1. The token lasts an hour and refreshes automatically for 30 days, after which you authorize again.
- Users need the Assistant Cloud MCP User role or the matching permission; Editor and above have it by default.
- At authorization you choose a level: read, query (which adds raw SQL), or write (which adds creating and changing dashboards, alerts and incidents). Choose read unless the session is for changing monitoring.
- For several stacks, put the stack in the path, such as
https://mcp.grafana.com/mcp/<prod-stack>.grafana.net, and add one entry per stack. - Grafana notes that each user who connects through MCP counts as an active user for billing.
How the browser sign-in works in general is in how MCP sign-in works. Choose the hosted server when you are on Grafana Cloud and want nothing to run; choose the open source server for self-hosted Grafana, or when you want flags such as --disable-write and --disable-query under your own control.
From a firing alert to a task
An agent with Grafana can do the first ten minutes of an investigation: list the firing rules, query the error rate, pull the log patterns, and check who is on call. With fenbs connected as a second MCP server at https://fenbs.ai/api/mcp, it can then search the board and file a bug in To Do with the alert, the queries it ran and a generate_deeplink link in the note. fenbs_create_item accepts a key from an automated source, so an alert that fires again does not file a second task. When the fix ships, the assistant records the test status and test notes from what the metrics show and moves the task to Completed; History keeps each step under its name. fenbs does not receive alerts or watch Grafana itself; the agent carries them across.
The same flow with Datadog, including its own read and write permissions, is in Datadog MCP server. Watching what your agents do, rather than your services, is AI agent observability.
Related
When the alert is an outage: AI agent incident response and the runbook template. Errors rather than metrics: Sentry MCP. Connecting fenbs: Claude Code, Cursor and the MCP docs.