AI Pair Programming: How to Work With a Coding Assistant

Treat the assistant as a pair, not a vending machine: decide who drives for each step, move in small steps, write the tests first, review every diff as the navigator would, and know when to hand it the keyboard.

7 min read

AI pair programming means working with a coding assistant the way two developers pair: one drives and writes the code, the other navigates, reviews and thinks ahead, and the roles switch often. With an assistant, you choose the roles per step. Let it drive when the task is clear and you can check the result with a test; drive yourself when the problem is new, the code is sensitive, or you want to learn it. Either way, keep the steps small, write or agree the tests before the code, read every diff before you keep it, and stop to correct it as soon as it drifts. The habits below come from human pairing and from what Anthropic, GitHub and Cursor publish about their own tools.

What AI pair programming is, and what it is not

Pairing is a conversation at the speed of the code: you are both looking at the same file, and each change is small enough to discuss. That sits between two other ways of using AI. Autocomplete is one keystroke at a time with no conversation; the kinds are compared in AI code generators. Delegation is handing an agent a whole task and coming back to a branch; the habits for that are in agentic coding best practices. Most days use all three. This page is about the middle one, when you and the assistant work the same problem together.

Driver and navigator, with an assistant

Birgitta Böckeler and Nina Siessegger’s guide on pair programming (opens in a new tab) defines the roles. The driver is “focussed on completing the tiny goal at hand, ignoring larger issues for the moment.” The navigator “reviews the code on-the-go, gives directions and shares thoughts.” Both roles map onto a coding assistant:

  • The assistant drives, you navigate. It edits in agent mode while you watch the diff, ask why, and stop it when it heads somewhere you did not mean. This is the common mode, and the one where your attention matters most.
  • You drive, the assistant navigates. You write the code and ask it to review as you go, explain an unfamiliar library, or list the edge cases you have missed. A read-only mode suits this: Cursor’s Ask mode (opens in a new tab) answers questions and explores code “without making any edits”, and Copilot and Claude Code have equivalents.
  • Ping-pong. One side writes a failing test, the other makes it pass, then you swap. This works well with an assistant and is covered below.

When to let it drive

  • Let it drive: the change follows a pattern already in the codebase, a test or build can check the result, and you could describe the diff in a sentence or two. Renames, a new endpoint like the last five, wiring a form to an API, writing tests from expected inputs and outputs.
  • Drive yourself: the design is not settled, the code handles money, security or personal data, the bug only shows up in production, or the point of the task is that you learn the code.
  • Navigate first, then let it drive: anything spanning several files. Ask it to read and plan before it edits, agree the plan, then hand over the keyboard.

GitHub’s best practices for Copilot (opens in a new tab) list what the assistant is good at, including “Writing tests and repetitive code” and explaining code, and say it is not designed to “Replace your expertise and skills.” That is the line between the two lists above.

A pairing session, step by step

  1. State the goal and the constraint in two sentences: what should change, and what must not.
  2. Ask it to explain the code you are about to touch. If its explanation is wrong, nothing it writes next will be right.
  3. Agree a plan for anything larger than one file, in plan mode if your tool has one.
  4. Write the tests first, or have it write them and read them yourself.
  5. Take the smallest next step, run the tests, and read the diff.
  6. Keep it or undo it. Commit when a step is green, so every step has a way back.
  7. Switch roles when you are stuck or bored, the same as with a person.

Tests first: ping-pong with an assistant

Tests are the most useful thing you can give a pair that types faster than you read. Cursor’s guide to coding with agents (opens in a new tab) sets out the loop: “Ask the agent to write tests based on expected input/output pairs. Tell the agent to run the tests and confirm they fail. Commit the tests when you’re satisfied with them. Ask the agent to write code that passes the tests, instructing it not to modify the tests.”

Prompts for one round of ping-pong
1. Write tests for calculateSalesTax(cart, zipCode) from these cases,
   using the rates in test/fixtures/tax-rates.json:
   - a ZIP code rated 8.25%, one taxable item at 100.00 -> tax 8.25
   - an Oregon ZIP code, any cart -> tax 0.00 (no state sales tax)
   - a taxable and a tax-exempt item -> tax on the taxable one only
   - empty cart -> tax 0.00
   This is test-driven development: do not write the function yet.
   Run the tests and show me that they fail.

2. (You read the tests, fix any case that is wrong, and commit them.)

3. Now write calculateSalesTax so these tests pass.
   Do not change the tests. Run them and paste the output.

The numbers in the cases come from you, not the assistant. A test it derived from the code it just wrote will agree with that code, right or wrong.

Small steps, and correct it early

A human pair notices drift within a minute. Do the same. Anthropic’s best practices for Claude Code (opens in a new tab) say to “Correct Claude as soon as you notice it going off track”: press Esc to stop it mid-action with the context kept, and rewind to an earlier checkpoint if a step went wrong. If you have corrected it more than twice on the same point, the session is full of failed attempts; clear it and start again with a better prompt that includes what you learned. Cursor gives the same advice in other words: start a new conversation when you move to a different task, when the agent keeps making the same mistake, or when you finish one logical unit of work.

Review like a navigator

  • Read the diff before the summary. The summary says what it meant to do; the diff says what it did.
  • Check the file list against the task. A change to a file you did not mention needs a reason.
  • Read the tests before the code, and look for any test it changed.
  • Ask “why this way?” about anything you would not have written. If the answer does not convince you, it is not done.
  • Run it yourself. The assistant’s report of a passing test is a claim until you have seen the output.

GitHub’s own wording is “Understand suggested code before you implement it.” The full routine, including what assistants most often get wrong, is in reviewing AI-generated code.

Keep learning while it types

The risk of a pair that never tires is that you stop thinking. A few habits keep you sharp: type the tricky part yourself now and then, ask it to explain a choice before you accept it, and when you are new to a codebase, spend the first session asking questions rather than requesting changes. Anthropic suggests asking the questions “you’d ask a senior engineer”, such as how logging works or why one function is called instead of another.

Where the session’s decisions go

A pairing session ends with code and a chat, and the chat is the part that disappears: why you chose one approach, what you tried first, what is left. On fenbs, put the task first and keep it current. The note says what the problem is, the plan says how you agreed to solve it, and the test status and test notes say what was checked when you finished. A connected assistant reads the task with fenbs_get_item, rewrites the plan with fenbs_update_item when the approach changes, and files anything it notices along the way with fenbs_create_item rather than fixing it on the side. When you settle on something that should always hold, such as “tests come before code in this repository”, record it on the Decisions and rules page; every connected AI assistant reads the rules first, and the decider is always a person.

Related

Connecting your assistant: Claude Code, Cursor and GitHub Copilot. Writing the task before the session: how to write a task for an AI agent. Checking the output: verifying AI-generated work.

Questions people ask.

What is AI pair programming?

It is working with a coding assistant the way two developers pair: one drives and writes the code, the other navigates and reviews, and the roles switch. You decide for each step whether the assistant drives in agent mode or you drive and it reviews and explains.

When should I let the AI drive?

When the change follows an existing pattern, a test or build can check it, and you could describe the diff in a sentence or two. Drive yourself when the design is unsettled, the code is sensitive, or you want to learn the code.

Should the AI write the tests or the code?

Either, but not both from the same assumptions. A good loop is that you supply the expected inputs and outputs, the assistant writes failing tests, you read and commit them, and then it writes code that passes them without changing the tests.

Does AI pair programming replace pairing with a person?

No. An assistant is a fast driver and a tireless reviewer, but it does not share your team’s context or carry knowledge to the next person. Human pairing still spreads knowledge across the team in a way an assistant cannot.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.