AI Coding Agent Benchmarks: What SWE-bench and Terminal-Bench Measure

How SWE-bench and its variants, Terminal-Bench, Aider’s polyglot benchmark and LiveCodeBench are built and scored, how to read a leaderboard without being misled by the scaffold or contamination, and how to run a small eval on your own repository.

8 min read

AI coding agent benchmarks give an agent a real task in a sandbox and check the result with tests it cannot see. SWE-bench asks it to fix a GitHub issue in a real repository and runs the tests from the pull request that originally fixed it. Terminal-Bench asks it to finish a job in a terminal, such as building, configuring or debugging something, and runs a verification script. The score is the share of tasks resolved. What a leaderboard ranks is a model and the agent software around it together, on public tasks that may have leaked into training, so use the numbers to shortlist and your own tasks to decide.

No scores appear here, because they change every few weeks. The leaderboards are linked instead. Choosing between the agents themselves is covered in best AI coding agents.

SWE-bench: how it is built and scored

The original SWE-bench paper (opens in a new tab), published at ICLR 2024, collected 2,294 problems from real GitHub issues and the pull requests that fixed them, across 12 popular Python repositories. Each task gives the agent the repository at the commit before the fix, plus the issue title and body. The agent edits the code. The harness then applies the tests that came with the real fix, inside a Docker container.

A task counts as resolved when two sets of tests agree. The dataset lists them as FAIL_TO_PASS, the tests the original fix made pass, and PASS_TO_PASS, tests that passed before and must still pass after. So the agent has to fix the reported behaviour without breaking what worked, and it never sees those tests while it works. The headline number is the percentage of tasks resolved.

The variants

  • SWE-bench Verified: 500 tasks screened by human annotators, in a collaboration with OpenAI, to remove issues that were unclear, unsolvable from the information given, or had tests that rejected correct fixes.
  • SWE-bench Lite: 300 test tasks and 23 development tasks, chosen to be self-contained, with no images, external links or references to other issues, and fixes confined to one file.
  • SWE-bench Multimodal: issues that include screenshots, mockups or diagrams, mostly in JavaScript projects. The original release had 517 tasks; the current version keeps 480 that evaluate reproducibly.
  • SWE-bench Multilingual: 300 tasks from 42 repositories in nine languages, including C, C++, Go, Java, JavaScript, TypeScript, PHP, Ruby and Rust.
  • SWE-bench Pro: a separate benchmark from Scale AI with 1,865 longer, multi-file problems from 41 repositories. Only some are public; the others are held out or come from commercial codebases, so that models cannot have trained on them. The SWE-bench Pro paper (opens in a new tab) describes the split.

The leaderboards for Full, Verified, Lite, Multimodal and Multilingual are on swebench.com (opens in a new tab). One of them, Bash Only, runs every model through the same minimal agent, mini-SWE-agent, so the model is the only thing that changes; the site notes that results from its version 1 and version 2 are not directly comparable.

Terminal-Bench: whole jobs in a terminal

SWE-bench is about patches. Terminal-Bench is about getting something done at a command line: compiling a project, training a small model, recovering data, configuring a server. Each task is a written instruction, a container that sets up the environment, a reference solution written by a person, and tests that check the final state. The Terminal-Bench 2.0 paper describes 89 such tasks. The benchmark is now versioned like software: 3.0 arrived in July 2026 with harder tasks and separate containers for the agent and the verifier, and 4.0 in August 2026 fixed tasks, removed ones every current model solves, and gave all tasks a flat eight-hour limit to cut timeouts.

The leaderboard on tbench.ai (opens in a new tab) ranks agent and model pairs by resolution rate, with confidence intervals, and shows cost and tokens alongside. You can run it yourself with the Harbor harness. Some current tasks need a GPU, which the site says to provide through a sandbox provider.

Running Terminal-Bench (from tbench.ai)
uv tool install 'harbor[modal]'
harbor run -d terminal-bench/terminal-bench@4.0.0 \
  -e modal -a claude-code -m anthropic/claude-sonnet-5 -k 5

Two more worth knowing

  • Aider’s polyglot benchmark: 225 challenging Exercism exercises in C++, Go, Java, JavaScript, Python and Rust. Results show a pass rate after a first attempt and after a second, and the Aider leaderboard (opens in a new tab) also reports how often it used the edit format correctly. It measures editing skill inside one agent, Aider, rather than comparing agents.
  • LiveCodeBench: contest problems from LeetCode, AtCoder and Codeforces, each tagged with its release date, so a model can be scored only on problems published after its training cutoff. It tests generation, self-repair, predicting test output and executing code, scored as pass@1.

How to read an AI coding agent leaderboard

Scaffold versus model

A row on a leaderboard is a model inside an agent: the prompts, the tools, the retry loop, the time and token budget. Anthropic’s write-up of its own SWE-bench agent (opens in a new tab) makes the point directly: the benchmark evaluates an entire agent system, and results can vary significantly with the scaffolding even when the model is the same. Compare models on a fixed-scaffold board such as Bash Only, and compare agents only when the model is held constant.

Contamination

Public benchmark tasks come from public repositories, which are exactly what models train on. In February 2026 OpenAI said it would stop reporting SWE-bench Verified, after an audit found many hard tasks had tests that rejected correct fixes and that the frontier models it tested had seen some of the problems and solutions in training. Newer designs push back in different ways: SWE-bench Pro holds tasks out, LiveCodeBench dates every problem, and Terminal-Bench retires and rewrites tasks between versions. When a score jumps, ask whether the model got better or the test got familiar.

pass@1, trials and cost

  • pass@1 is the share of tasks solved on a single attempt. pass@k, from the paper that introduced HumanEval, counts a task solved if any of k attempts passes, which rises quickly with k and flatters an agent you will run once.
  • Look for the number of trials and a confidence interval. A gap of a point or two between two agents run once each is often noise.
  • Read the cost and token columns. An agent that resolves slightly more tasks at several times the tokens may be the wrong choice for everyday work.
  • Check the date and version. Scores on SWE-bench Verified, Lite and Pro, or on Terminal-Bench 2.0 and 4.0, are different measurements and do not compare.

What benchmarks miss for real teams

  • Your codebase. Benchmarks use popular open-source projects. Your private conventions, internal libraries and odd build are not in them.
  • Vague tasks. Benchmark issues are screened to be solvable. Real tickets are often under-specified, and asking a good question is part of the job.
  • Everything beyond a green test. Readability, whether the change is the smallest one that works, and whether a reviewer can follow it are not scored.
  • Safety of the route. A benchmark grades the final state. It does not mark down an agent that deleted a directory or ran something risky on the way.
  • Working with people. Handing back partial work, saying what it could not test, and following your review process are invisible to a pass rate.

Run a small eval on your own repository

The SWE-bench recipe works on your own history, and twenty tasks tell you more than any leaderboard about how an agent handles your code. The general method for agent evals, with graders and the online side, is in AI agent evaluation. For coding agents specifically:

  1. Pick 20 to 30 merged pull requests that fixed a real bug or added a small feature, each with tests that failed before and pass after. Mix easy and hard, and include a few from your messiest area.
  2. For each, record the parent commit, the task description as it was written at the time, and the test command. Keep the pull request’s tests out of the agent’s sight.
  3. Run each agent or configuration you are comparing from the parent commit, with the same instructions and the same time limit, at least three times per task.
  4. Grade with the hidden tests plus your existing suite, then read a sample of the diffs yourself for size and readability.
  5. Record resolved rate, variance across runs, time and tokens. Re-run the set whenever you change the model, the instructions or the tools.
One task, the shape of the loop
git worktree add ../eval-017 <parent-commit>
cd ../eval-017
# give the agent the task text only; tests from the real PR stay outside
<run the agent with a fixed time limit>
git apply ../evals/017-tests.patch
<your test command> > ../evals/017-run1.log
git diff --stat >> ../evals/017-run1.log

Failures become tasks

The useful output of an eval is the list of what went wrong. On a fenbs board, each repeated failure pattern, such as an agent that edits generated files or skips your migration step, becomes a task: an enhancement to the agent’s instructions, or a bug if it exposed a real defect in your code. The note carries the eval task ID and a log excerpt. When you change the instructions and re-run, the task’s test status records the result, Tested or Partly tested with notes on which runs passed. Sizing each fix from XS to XL shows which are quick and which need splitting. fenbs does not run evals; it keeps the list of what to fix and whether the fix held.

Related

The broader method: AI agent evaluation. Choosing an agent: best AI coding agents and Claude Code vs Codex. Why agents fail on real work: why AI agents fail.

Questions people ask.

What is SWE-bench?

A benchmark built from real GitHub issues and the pull requests that fixed them. An agent gets the repository before the fix and the issue text, edits the code, and is scored by whether the tests from the real fix now pass while existing tests still pass. The score is the percentage of tasks resolved.

What is the difference between SWE-bench and Terminal-Bench?

SWE-bench measures whether an agent can patch a repository to fix an issue. Terminal-Bench measures whether it can complete a whole job at a command line, such as building or configuring something, checked by tests on the final state of a container.

Why do AI coding agent leaderboard scores differ for the same model?

Because each entry is a model inside a particular agent, with its own prompts, tools, retries and budget, and scores also vary between runs. Compare models on a fixed-scaffold leaderboard and compare agents with the model held constant.

Can I trust SWE-bench Verified scores?

Treat them with care. OpenAI stopped reporting them in February 2026, citing flawed tests on hard tasks and evidence that models had seen some problems in training. Use them alongside newer or held-out benchmarks and a small eval on your own code.

Start with one thing.

There is nothing to set up first. Write one line and you’ve started.