Personal Agent Bench

SWE-bench, and the distance between a resolved issue and a finished errand

SWE-bench takes real bug reports from real Python repositories, hands an agent the repository at the commit before the fix, and checks whether the tests pass afterwards. It is the most cited agent benchmark in existence and the one most often quoted at you as evidence that agents work. It earns a page here for a reason that has nothing to do with its scores: it is the clearest example in this section of a benchmark that got its methodology right and its headline number misread anyway.

  • The grading is not an opinion. A test suite either passes or it does not, which is why this benchmark became the standard.
  • The default board holds the harness constant across every entry, so what you are comparing really is models. Very few leaderboards can say that.
  • Two models tie at the same resolution rate while one costs five times less, and a newer model from the same vendor scores lower than its predecessor. Neither is visible from a rank.
  • Nothing it measures involves your accounts, your history or an action that cannot be undone, which is where a personal agent lives.

What it actually measures

Each instance is a GitHub issue paired with the repository state at the moment before someone fixed it. The agent reads the issue, edits the code, and the benchmark runs the project's own tests, including the tests that were added with the real fix. Resolved means those tests pass.

That design is the reason this benchmark mattered. There is no judge model, no rubric, no partial credit argument. The repository's maintainers already wrote down what correct means, in executable form, before anyone thought of pointing an agent at it.

The family has grown into five sets. The original SWE-bench holds 2,294 instances drawn from 12 Python repositories. Verified is a human-filtered subset of 500 and is what the site shows by default. Lite is a 300-instance subset for cheaper evaluation, Multilingual is 300 tasks across 42 repositories and 9 programming languages, and Multimodal is 480 issues described with visual elements. A number quoted without naming which of those it came from is close to meaningless, and most quoted numbers do not name one.

The leaderboard

#SubmissionHarnessOverallAvg. costDated
1Claude 4.5 Opus (high)mini-SWE-agent 2.0.076.80%$0.752026-02-17
2Gemini 3 Flash (high)mini-SWE-agent 2.0.075.80%$0.362026-02-17
3MiniMax M2.5 (high)mini-SWE-agent 2.0.075.80%$0.072026-02-17
4Claude 4.6 Opusmini-SWE-agent 2.0.075.60%$0.552026-02-17
5Claude 4.5 Opus (medium)mini-SWE-agent 1.16.074.40%$0.722025-11-24
As of 2026-10-08source
Percentage of the 500 instances in SWE-bench Verified resolved, meaning the repository's own tests pass after the edit. Every entry runs in the same mini-SWE-agent environment in bash-only mode, which is what makes the column a comparison between models rather than between scaffolds. Read off the rendered board on 2026-10-08; the harness version is printed beside each row because entries from different versions are not strictly comparable. No confidence intervals are published, so gaps of a point or two cannot be separated from sampling noise.

Holding the harness still is the real contribution

The default Verified view carries a line that is easy to scroll past: every model in the same mini-SWE-agent environment, bash only.

That sentence is doing more work than the scores above it. An agent result is always a product of two things, the model and the scaffolding wrapped around it, and most published comparisons vary both at once. A vendor reports a number from its own harness, a competitor reports one from theirs, and the difference between them is attributed to the models. It usually is not.

Fixing the scaffolding across every row is the only way to make the column mean what readers assume it means. The board goes further and shows the harness version beside each entry, which is an admission that even the fixed scaffolding moves over time. Rows from version 1.16 and version 2.0 of that harness are not strictly comparable, and the board says so by printing the version rather than hiding it.

Two entries tie, and one of them costs far less

Second and third place both resolve 75.80% of the Verified set. The average cost per instance is $0.36 for one and $0.07 for the other.

A rank column cannot express that, and it is the single most decision-relevant fact on the page. If you are choosing what to actually run against a backlog, five times the price for an identical resolution rate is the whole question, and the ordering tells you nothing about it.

The cost column also reframes the top of the board. First place resolves one percentage point more than the two tied below it, at roughly ten times the cost of the cheaper one. Whether that point is worth buying depends entirely on what the work is worth, which is a judgment the leaderboard correctly declines to make for you.

A newer model scoring lower is not a mistake

Fourth place is a later release from the same vendor as first place, and it resolves slightly less of the set.

The reflex is to read that as a regression. It is better read as a reminder of what a single benchmark covers. A model tuned toward longer-horizon work, better refusal behaviour or cheaper inference can lose ground on a set of self-contained Python bug fixes while being the better product for most of what people do with it. The score is a measurement of one thing, and the one thing is narrow.

There are no confidence intervals on this board, which makes gaps of one or two points impossible to interpret. One point on a 500-instance set is five instances. Whether five instances separates two models or separates two sampling runs of the same model is not a question the page lets you answer, and it is exactly the question Terminal-Bench answers by printing a margin beside every score.

The bottom of the board is the more useful half

Forty-seven entries sit on the default view and they run from 76.80% down to 9.00%. Almost all of the attention goes to the top five.

The spread is the more informative part. Models separated by two years of releases differ by a factor of eight on the same fixed harness, which is a far stronger statement about progress than any single top score, and it is the part that does not get quoted. It also sets a floor on how to read a lone number in isolation: a product claiming a resolution rate in the fifties is quoting a figure that sat mid-table a year ago, and nothing in the claim tells you that.

Two markings on the board are worth knowing about. Some rows are flagged as open-weights models, and some are flagged as having been run or directly checked by the benchmark team rather than self-reported. That second flag is the one to look for. A self-reported number and an independently reproduced one are different kinds of evidence, and a board that distinguishes them is telling you which rows it stands behind.

Who checked the number, and why that question comes first

Our own scorecard exists because of that distinction. Every rate we publish comes from runs we executed ourselves, through our own harness, with the record kept so the figure can be rechecked rather than believed.

That is not a claim to be more rigorous than a benchmark with a decade of citations behind it. It is a narrower claim: when a number is about a product rather than a model, and the product changes without notice, provenance is the only thing that keeps the figure meaningful a month later. SWE-bench can afford to accept submissions because its task set is frozen and its grading is mechanical. We cannot, because ours is neither.

Where it stops being about personal agents

A SWE-bench instance is a closed problem. The repository is a copy, the tests are the specification, the work is reversible, and when the run ends nothing persists.

An errand has none of those properties. There is no test suite that says whether the right flight was changed, the message went to the right thread, or the cancellation actually went through rather than being deflected into a retention offer. The specification lives in a person's head, it is partly unstated, and the consequences of getting it wrong land on real accounts that belong to someone.

This is why a resolution rate on Python issues, however honestly produced, cannot be carried over. The products on this bench are judged on whether a real-world state changed correctly, which is checkable but not by running a test suite, and on what they broke on the way, which a pass or fail cannot express. The method behind that is written up in how the exam works.

What a personal agent inherits anyway

Two things from this benchmark transfer directly, and we took both.

The first is holding the harness constant. Every product on our board runs through the same scored environment and the same scoring code, for the same reason: without that, a comparison is measuring our integration work rather than the products. The second is grading against state rather than against text. SWE-bench asks whether the tests pass, not whether the agent said it fixed the bug, and our exam asks whether the booking exists in the world, not whether the assistant reported making it. An agent that announces success it did not achieve is the most common failure mode we see, and it is invisible to any benchmark that reads the final message. It is also the failure mode people are least prepared for, because a confident summary reads exactly like a completed errand right up until the moment you go and check.

What does not transfer is the assumption that a clean pass is the whole result. A fix that passes the tests and leaves an unrelated file mangled still counts as resolved here. On a machine that is yours, it would not.

FAQ

How many tasks does SWE-bench contain? It depends which set. The original is 2,294 instances from 12 Python repositories. The default view on the site is Verified, a human-filtered subset of 500. Lite is 300, Multilingual is 300 across 42 repositories and 9 languages, and Multimodal is 480.

What does resolved mean? The repository's own test suite passes after the agent's edit, including the tests that shipped with the real human fix. There is no model judging the answer and no partial credit.

Why do some rows show a different harness version? The board records the version of the evaluation harness each entry was run under, because the scaffolding changes over time. Entries from different versions are not strictly comparable, and printing the version is how the board says so.

Does a higher score mean a better coding agent? It means a better result on self-contained Python bug fixes under one fixed harness. Cost, speed, behaviour on long tasks and what the agent damages along the way are either in other columns or not measured at all.

Does any of this predict how a personal agent will do? Not directly. The skills overlap, in that both involve reading a situation and acting on it, but the failure modes do not. Nothing here tests a standing preference, an irreversible action, or an account that belongs to a person.

Is the leaderboard current? The most recent dated entry when this page was last verified was from late February 2026. Entries are submitted rather than continuously refreshed, so the board reflects what has been submitted and checked, not everything that exists.

The other benchmark pages

No numbers invented. Personal agents taken apart at the runtime level, on a method built to take more.

See the scorecard