Personal Agent Bench

Terminal-Bench, and the one habit the rest of this field should copy

Terminal-bench puts an agent in a terminal and asks it to finish real work there: fix the thing, run the thing, leave the machine in the state the task described. It is hosted by Stanford, Harbor and the Laude Institute, the harness is Apache licensed, and the public leaderboard is on version 4.0. It is also, as far as we can tell, the only widely cited agent benchmark that publishes a confidence interval beside every score, which is why it gets a page here even though it measures almost nothing a personal agent is for.

  • Every row on its leaderboard carries a plus or minus. That one habit separates a result from a number, and it is rarer in this field than it should be.
  • It measures work in a terminal. No inbox, no calendar, no saved logins, no consequences outside the box.
  • The gap between first and fifth place is smaller than it looks once the intervals are read, which is the entire argument for printing them.
  • We take the interval habit from it and nothing else, because the thing it measures and the thing we measure barely overlap.

What it actually measures

An agent gets a terminal and a task with a checkable end state. It works, and the harness inspects the machine afterwards to see whether the described state was reached. Resolution rate is the share of tasks it got there on.

The task set lives in the repository and the scored set is versioned, which matters more than it sounds: a benchmark whose task list quietly grows is a benchmark whose scores cannot be compared across months. Version 4.0 is the current scored set, named in the leaderboard definition in the repository rather than only on the site.

The leaderboard

#SubmissionAgentOverallCostDated
1Opus 5.5 (max)Claude Code64.8% ± 3.1%$4.7k2026-09-22
2Sonnet 5.5 (max)Claude Code61.8% ± 2.9%$7.3k2026-09-28
3GPT-6 Astra (max)Codex58.2% ± 2.8%$3.3k2026-09-03
4GPT-6.1 Sol (max)Codex58.2% ± 3.1%$0.6k2026-09-29
5Fable 5.1 (max)Claude Code57.9% ± 3.8%$6.2k2026-09-01
As of 2026-10-08source
Resolution rate on the version 4 task set, with the interval the leaderboard publishes beside each score. Read off the rendered board on 2026-10-08; version 4.0 is confirmed by leaderboard/leaderboard.yaml in the repository, which names the Hub leaderboard 4-0-0. The scored revision holds 66 tasks, per the dataset page on the Harbor hub, which labels it rev. 4 (4.0.0) and reports 66 of 66; the 67 task directories on the repository main branch are one more than that, so the hub figure is the one used here.

Read the interval before the rank

Look at the top of that table as a ranking and you get a clean story: one model wins, another is second, two tie for third.

Read the intervals and the story changes. First place is 64.8% with a margin of 3.1 points, second is 61.8% with 2.9. Those ranges touch. The two models tied at third are genuinely indistinguishable, and the fifth-place score sits inside the range of the two above it. What the table actually establishes is roughly two groups, not five positions.

A rank implies an order that the measurement cannot support. The interval is what tells you how much of the gap is real.

This is the same reason our own scorecard prints a range and a task count next to every rate and refuses to combine them into one number. We arrived at that rule separately, and finding it already in use here was the main reason to write this page.

What it costs to run, and why that is on the board

Two columns on that leaderboard are unusual: tokens and cost. The top entry burned about 8 billion tokens and roughly 4.7 thousand dollars to reach its score.

Publishing that beside the result is an honest move. It makes visible that a score is partly a statement about budget, and it lets you see when a higher rank was bought rather than earned.

The clearest example is not at the top. Two entries share third place at 58.2%. One spent about 3.3 thousand dollars getting there, the other about 0.6 thousand. Same resolution rate, same scored task set, roughly a fifth of the money. A rank column cannot say that, and most benchmarks in this section would not have let you see it, because they publish a score and stop.

Whether it matters depends on who is reading. A lab comparing capability can ignore it. Anyone deciding what to actually run cannot, since the thing they are choosing between is not two numbers but two bills. The cheaper of the two also posted a slightly wider interval, which is the trade stated plainly: fewer dollars spent, slightly less certainty about where the true rate sits.

It also quietly undermines the leaderboard as a leaderboard. Once cost is on the page, first place stops being a claim about which model is best and becomes a claim about which is best at some price, and the page does not pretend otherwise. Our own board has no cost column, because the four products we score are not billed in a way that compares, and leaving the column out is the more honest of the two options available to us.

Where it stops being about personal agents

Everything above is about a terminal. The task arrives cold, the agent has no history with you, and the machine it works on was created for the task and is destroyed after it.

A personal agent fails in ways that setup cannot reach. It sends the message to the wrong thread. It accepts a retention offer on your behalf. It books the second reservation because the first page was slow. It remembers a preference you stated once and applies it for months. None of those are terminal problems, and none of them show up as a resolution rate.

The products this site examines all have an inbox, saved logins and the ability to act on accounts that belong to you. The distance between that and a disposable container is the whole subject of how the exam works.

That is not a criticism of the harness. A disposable machine is the right design for the question it asks, and the reason its numbers are reproducible at all is that nothing survives between runs. The failure is in borrowing the number for a question it was never built to answer.

There is one place where the two do overlap, and it is worth naming because it is where a personal agent inherits terminal-bench's problem.

Several of the products on this bench give the agent a machine of its own: a sandbox, a container, in one case a Linux VM per user. Everything terminal-bench measures about operating a shell applies there. The difference is that their machine persists, holds your files, and carries whatever the agent learned yesterday into today, so a mistake does not vanish when the task ends.

That is the axis our teardowns call where it runs, and reading terminal-bench alongside them is genuinely useful. It tells you how well a model drives a shell. It cannot tell you what happens when the shell is yours.

The part that is easy to miss

The harness is open and the tasks are readable. Anybody can look at what a task actually asks and judge whether the scored behaviour is the behaviour they care about.

That is not a small thing. Several benchmarks in this section publish a number whose underlying task set is either partly held back or has changed since the figure was computed, which makes the figure unfalsifiable in practice. A task you can open is a claim you can check, and the same standard is why every run we publish ships with its own record.

Where the harness came from

The repository was created in January 2026 and has been in steady development since, under Apache 2.0, with the scored task set versioned alongside it. That history is worth a sentence because it explains the version number: a benchmark on its fourth scored set has already discarded tasks that turned out to measure the wrong thing, which is the normal and healthy path.

It also means the earlier numbers you may find elsewhere are not comparable to these. A score against version 2 and a score against version 4 are answers to different exams, and nothing on a chart makes that visible. We have the same problem on our own board, where one product's rate moved several points after we fixed defects in our adapter rather than because the product changed, and the scorecard says so in the row rather than in a footnote.

What the dates on each row are doing there

Every row carries the date its run was scored, and the five entries above span four weeks. That column is doing more work than it looks.

A leaderboard without dates implies all its rows were produced under the same conditions, which is almost never true. Models get updated behind a stable name, the harness gets patched, the provider's latency and rate limits move. A row scored on September 1 and a row scored on September 29 are two measurements taken on two different days of a moving system, and the gap between them is not necessarily a difference between the models at all.

This is the single most common way benchmark tables mislead without stating anything false. We hit it directly: our own figures were produced before several defects in our adapters were fixed, and the honest response was to keep the date beside every rate rather than quietly refresh the number and leave the impression that nothing underneath had changed.

How to read it if you came here about personal agents

Treat it as a measure of one capability, stated carefully, in a setting unlike yours.

If a product's agent runs commands on a machine, terminal-bench tells you something real about the model underneath it. It tells you nothing about whether the product asks before spending your money, whether it can be reached by a stranger who knows your email address, or what it retains about you. Those are the questions the scorecard is built around, and no terminal benchmark will answer them.

FAQ

What is Terminal-Bench? A benchmark that gives an agent a terminal and a task with a checkable end state, then measures how often the machine ends up in the state the task described. It is hosted by Stanford, Harbor and the Laude Institute, and the current public leaderboard is version 4.0.

Why does this site cover a terminal benchmark? For one habit worth copying: it publishes a confidence interval beside every score. That practice is what separates a result from a number, and it is the same reason our own figures ship with a range and a sample count.

Does a good Terminal-Bench score mean a good personal agent? No. It measures operating a shell on a disposable machine. A personal agent is judged on what it does with your accounts, your messages and your money, and those failures have no terminal equivalent.

How many tasks does it contain? 66, for the version that the current leaderboard scores. The number is not on the leaderboard itself; it comes from the dataset's own page on the Harbor hub, which lists revision 4 and says it is displaying 66 of 66 tasks. The harness repository's main branch holds 67 task directories, one more than the scored revision, which is why the hub figure is the one to use.

Where do the numbers on this page come from? The rendered leaderboard at tbench.ai, read on the date shown above it. The version is confirmed by the leaderboard definition file in the repository rather than taken from the site alone.

The other benchmark pages

No numbers invented. Personal agents taken apart at the runtime level, on a method built to take more.

See the scorecard