GitHub Copilot Agent vs Cursor Agent: Handling Multi-Step Tasks
Both agents plan, edit, run commands and keep going until a task is done. The differences are in the structure around that loop: how each plans, asks before acting, lets you undo, hands work to the cloud and reads your rules.
8 min read
On a multi-step task, GitHub Copilot’s agent in VS Code and Cursor’s agent work the same way: plan, edit files, run commands, read the result, try again. Where they differ is the structure around that loop. Copilot is built around VS Code and GitHub, with a cloud agent that runs in GitHub Actions and answers with a pull request. Cursor is a whole editor built around its agent, with run modes that decide what needs your approval and cloud agents you can start from the web, Slack or a pull request comment. Neither is better in general. Choose by where the work should happen, how you want to approve each step, and where your team reviews the result.
What counts as a multi-step task
Renaming a function is one step. Adding a field to a form, the API that saves it, the migration and the tests is several, and each depends on the one before. That is where agents differ: how they agree an approach before editing, how often they stop to ask, what you can roll back when step four goes wrong, and whether the work can carry on while you do something else. The general loop an agent follows is in the steps an AI agent takes to complete a task; this post compares the two tools on those points only.
Planning before the first edit
Both have a planning mode that reads the code, asks questions and proposes an approach without changing anything.
- Copilot in VS Code: pick Plan in the agent picker or type
/planfollowed by the outcome, the constraints and how you will check it. VS Code’s page on planning with agents (opens in a new tab) describes Implement Plan, which approves the plan and starts the work in the same session, and saving a separate copy of the plan to keep it with the project. - Cursor: press Shift+Tab in the chat input to reach Plan Mode, or pick it from the mode menu; Cursor also suggests it for complex tasks. According to Cursor’s Plan Mode documentation (opens in a new tab), the agent asks clarifying questions, researches the codebase and writes a plan you can edit in chat or as Markdown. Plans are saved in your home directory by default, and Save to workspace moves one into the project.
The practical difference is where the plan lives afterwards. Cursor writes it as a file you can keep in the repository; VS Code keeps it with the session unless you save a copy. Either way, a plan that only exists beside one chat is hard for a colleague to find. When to plan at all is covered in Copilot agent vs ask vs plan.
Tool use and what needs your approval
A multi-step task means many tool calls: searches, edits, terminal commands, MCP tools. How often the agent stops to ask shapes the whole session.
- Copilot in VS Code has three permission levels in the chat input. Manual, the default, follows your tool, URL and terminal approval settings; Assisted, marked experimental, uses a model to judge each call; Allow all runs every call without asking. When it asks, you can approve once, for the session, for the workspace or always. Autopilot is a separate mode that approves everything, retries on errors and answers its own blocking questions.
- Cursor sets this with run modes. Cursor’s run modes page (opens in a new tab) lists Auto-review, where allowlisted calls run at once and a classifier reviews riskier ones; Allowlist, where only what you listed runs without approval; and Run Everything, which runs every call with no review. MCP tools ask for approval by default and follow the same modes.
- Both can sandbox terminal commands, restricting file and network access, as a layer on top of whichever approval setting is on.
Checkpoints and undo
Step four going wrong is normal. What matters is how far back you can go, and what going back does not cover.
- Copilot in VS Code: edits in a standard session wait for Keep or Undo, file by file or change by change. VS Code’s page on reviewing agent edits (opens in a new tab) says a checkpoint is taken before each request and restores files and chat history, but does not reverse terminal commands, network requests, deployments or changes tools made to outside services.
- Cursor: checkpoints are created automatically before significant changes and restore every modified file to that state. They are stored locally and kept separate from Git. While the agent works, you can queue a follow-up message, which it takes after the current step.
The limit is the same in both. A checkpoint rewinds files. A row written to a database, a message sent, or a task moved on a board through MCP stays done. Keep approvals strictest for those.
Reviewing what changed
- Copilot: review the diff in the editor before you keep it. In the Agents window you can select a range in a changed file and add feedback for the agent. On GitHub, Copilot code review can read the pull request.
- Cursor: the diff view shows changes as they happen, with Stop if the agent heads the wrong way. When it finishes, Review then Find Issues runs a separate review pass. Bugbot reviews pull requests in your source control provider.
An agent’s own review is a second opinion, not a check. Run the tests yourself; verifying AI-generated work has the routine.
Background and cloud agents
Long tasks do not need your editor open. Both tools can run the work somewhere else and hand back a branch.
- Copilot cloud agent: GitHub’s page about the cloud agent (opens in a new tab) says it works in its own ephemeral environment powered by GitHub Actions, where it can run tests and linters. You start it from the agents panel on GitHub.com, an issue, VS Code, or
@copilotin a pull request comment, and it works on a branch you review as a pull request. In VS Code, the Cloud target runs a session this way; local sessions in their own git worktree run with Allow all. - Cursor cloud agents: Cursor’s cloud agents documentation (opens in a new tab) says they run in isolated virtual machines with full development environments. You start them from Cursor on the web, the desktop app, the iOS app,
@cursoron a GitHub issue or pull request, Slack, Linear or the API. They clone the repository from GitHub, GitLab, Azure DevOps or Bitbucket, work on a separate branch and push it. The Agents window runs several agents in parallel, each in its own worktree, and can move an agent between cloud and local.
The environment matters as much as the agent. A cloud agent that cannot build your project cannot test its own change, so both need setup: Copilot through its Actions environment, Cursor through an environment it sets up, a saved snapshot or a Dockerfile.
Instructions files
- Copilot reads
.github/copilot-instructions.mdfor the repository,.instructions.mdfiles with anapplyTopattern for matching files, andAGENTS.md. Which surface reads which is in does GitHub Copilot support AGENTS.md? - Cursor reads project rules as
.mdcfiles in.cursor/rules, plusAGENTS.md, and aCLAUDE.mdat the project root, which it always applies. Rule types are in Cursor rules for AI projects.
A multi-step task leans on these files more than a one-line edit does, because the agent makes many small choices without asking. If you use both tools, one AGENTS.md reaches both. The full file-by-file map is in AI context files compared.
MCP support
Both reach outside the codebase over MCP, and both support local and remote servers. VS Code keeps servers in .vscode/mcp.json for the workspace or in your user profile, and also reads a portable .mcp.json at the project root. Cursor keeps them in .cursor/mcp.json in the project or ~/.cursor/mcp.json for every project, over stdio, SSE or Streamable HTTP. The same remote server works in both:
// VS Code: .vscode/mcp.json
{ "servers": { "fenbs": { "type": "http", "url": "https://fenbs.ai/api/mcp" } } }
// Cursor: .cursor/mcp.json
{ "mcpServers": { "fenbs": { "url": "https://fenbs.ai/api/mcp" } } }Which fits which situation
- Your team already works in VS Code and reviews on GitHub: Copilot. The agent, the pull request and code review sit in tools you already use, and the cloud agent picks up issues where they are filed.
- You want one editor built around the agent, with plans saved as files: Cursor, with Plan Mode and the Agents window.
- You want to start long tasks from chat or a tracker: Cursor’s cloud agents start from Slack and Linear as well as GitHub; Copilot’s start from GitHub issues, pull request comments and VS Code.
- Your code is not only on GitHub: Cursor’s cloud agents also clone from GitLab, Azure DevOps and Bitbucket. Copilot’s cloud agent is built on GitHub.
- You want fine control over each command: both offer it. Copilot through approval scopes, Cursor through an allowlist. Pick the one your team will actually keep narrow.
Keeping a long task visible outside the agent
Neither tool’s session list is where your colleagues look. A plan in a chat, a cloud run on one person’s account and a checkpoint on one laptop each record part of the task and none of them the whole. Put the task on a board both agents can reach. On fenbs, each task has a note for the problem and a plan for how it will be done, lanes To Do, Next Up, In Progress and Completed, and a testing status. Connect either editor over MCP and add a line to your rules: write the approved plan into the task with fenbs_update_item, move it to In Progress when work starts, and comment what changed and how it was checked when it stops. Every change is recorded with the assistant’s name and yours, whichever editor made it.
Related
Set-up pages: Cursor and GitHub Copilot. The same comparison for a terminal agent is Claude Code vs GitHub Copilot, and the habits for Copilot sessions are in GitHub Copilot agent mode best practices.