Personal Agent Bench

OSWorld 2.0, and the point where every agent stops finishing

OSWorld gives an agent a real desktop, real applications and a task that a person would recognise as work, then checks the machine afterwards to see whether the work got done. Version 2.0, published in June 2026, rebuilt the task set around long workflows: 108 of them, each taking a skilled human a median of about 1.6 hours. It is the closest thing in this section to what we measure, and it is also the benchmark that has published the most uncomfortable finding about how far agents currently get.

  • Its tasks are long. Nearly 70% take a skilled person more than an hour, and one shown on its own site runs past 500 agent steps without the claim ever being submitted.
  • It reports two numbers per entry, finished and partially finished, and they differ by more than thirty points at the top of the board.
  • Above a certain task length, every model it tested completes zero tasks outright. That is a published result, not an inference.
  • Its own scores move by double digits when only the benchmark version changes, which is the strongest available argument for printing dates beside numbers.

What it actually measures

An agent is placed in a working operating system with real software installed and given a task described the way a person would describe it. When the run ends, scripts inspect the machine: is the file there, does it contain the right thing, was the form submitted. Success is a state of the world, not a judgement about the agent's explanation of itself.

Version 2.0 spans seven professional domains and 21 sub-categories, covering research, creative production, engineering, personal services, business and finance, administration and compliance, and healthcare. The authors also map tasks onto occupation families to estimate what share of actual paid work the set covers, which is an unusual thing for a benchmark to attempt and a useful corrective to task sets assembled from whatever was easy to automate. The largest shares land on document preparation, software and databases, and finance and operations analysis, with a long tail behind them. Reading that distribution is a quicker way to understand what the benchmark is for than reading the task list.

The jump from the first version is the headline. A version 1.0 task took roughly 30 tool calls. A version 2.0 task averages 318. That is not a harder version of the same exam, it is a different exam, and the authors treat it as one by keeping the first version online at its own address instead of replacing it.

The leaderboard

#SubmissionBenchmark releaseOverallPartial
1Claude Opus 5 (Max, batch tool)v2.144.33%77.67%
2Claude Opus 5 (High, batch tool)v2.136.89%68.92%
3Claude Opus 5 (Xhigh, batch tool)v2.133.33%73.11%
4Claude Opus 5 (Medium, batch tool)v2.133.01%64.85%
5Claude Opus 5 (Max, batch tool)v2026.08.0831.43%68.31%
As of 2026-10-08sourcepaper
Binary accuracy on OSWorld 2.0, the share of the 108 long-horizon workflows completed outright, with the partial column giving the share of intermediate checkpoints reached. Read off the rendered board on 2026-10-08 under the Foundation E2E GUI view with no step-budget filter applied. The board carries no dates and no confidence intervals, and its cost per task column is empty for these entries. The release column is kept because it moves the score: the same model at the same setting scores 44.33% on v2.1 and 31.43% on v2026.08.08.

Finished and nearly finished are not the same column

Every entry carries two rates. Binary accuracy is the share of tasks completed outright. Partial is the share of checkpoints reached.

At the top of the board those two numbers are 44.33% and 77.67%. The gap is the whole story. Three quarters of the required steps get done, and fewer than half of the tasks actually land.

That gap is not a rounding problem, it is the shape of the failure. An agent that gets most of the way through a reimbursement claim and never submits it has not done 77% of the errand. It has done none of it, and it has probably left a half-filled form behind that someone now has to find and clean up. A benchmark that reported only the partial figure would describe these agents as nearly competent. A benchmark that reported only the binary figure would hide how close they get. Printing both is the honest option, and it is the one chosen here.

The point where completion goes to zero

The site publishes a breakdown of completion against task length, and it ends somewhere most benchmark pages do not go.

On tasks under 45 minutes of human-equivalent work, the better models complete 20% to 24% outright. In the 137 to 163 minute band, no model exceeds 10%. Above 163 minutes, binary completion falls to zero for every model retained in the plot.

Not low. Zero. There is a length of errand past which nothing currently available finishes the job.

The authors attribute it to compounding: execution load and state-management errors accumulate over a long run in a way that short tasks never expose. That matches what we see. The failures that end our own runs are rarely a single wrong click. They are an early misreading that survives twenty steps and quietly invalidates everything after it.

What a 500-step failure actually looks like

The site walks through one run in detail, and it is worth reading because it is the most legible account of long-horizon failure published anywhere in this section.

The task is an overseas reimbursement claim: gather evidence from a bank statement, an email and a guidelines document, then file it in the expense system before a deadline. The agent reads the rules correctly at step 65. By step 120 it is still searching for evidence. At step 220 it hits a dead API and turns that into a tooling detour. At step 300 it is still collecting. It finally opens the expense system at step 459, and the run ends with the agent reverse-engineering a date picker rather than with a submitted claim.

Nothing in that trace is a capability failure. Every individual step is competent. What is missing is any sense that the preparation has a budget, and that is not a skill a resolution rate measures.

It is also the failure we see most often in our own transcripts, under a different costume. An assistant asked to cancel a subscription reads the terms, finds the help centre, locates the policy, explains the policy back, and never reaches the button. Thoroughness and completion are separate abilities, and only one of them is what was asked for.

Version numbers move the score more than model choice does

One pair of rows on that board deserves more attention than the top one.

The same model, at the same reasoning setting, with the same tool configuration, scores 44.33% on one release of the benchmark and 31.43% on another. Nothing about the agent changed. Almost thirteen points of difference come from which version of the task set it was scored against.

This is the strongest available evidence for a rule we apply to our own numbers: a rate without a date and a method beside it is not a measurement, it is a rumour. Anyone quoting an OSWorld score without saying which release produced it is quoting a number that could be off by a third. The board itself handles this correctly, letting you filter by release version, which is more than most leaderboards offer and more than most readers use.

Buying the last few points costs more than it looks

The published findings include a curve that is easy to summarise and hard to argue with: scores climb, and the token budget required to climb them climbs faster.

One model reaches about 14% binary completion on roughly 39 thousand output tokens and then flattens out. Another gets to 18.2% at 150 thousand, and a later release of it reaches 20.6% at 244 thousand. Going from 14% to 20.6% costs more than six times the output.

This matters to anyone building on top of these models rather than reading about them, and it matters in a specific way for a personal agent. An errand is not worth an unbounded budget. Spending six times the tokens to raise the chance of a correctly rebooked flight from one in seven to one in five is a trade almost nobody would take knowingly, and it is exactly the trade a product makes silently when it is tuned for leaderboard position. The shape of that curve is an argument for measuring at a fixed budget, which is what both this benchmark and our own exam do.

Where it stops being about personal agents

OSWorld is the closest neighbour this site has, and the distance is still real.

Its tasks are professional work: engineering, compliance, finance, healthcare administration. They are long and they are hard, but they are assigned. Nobody on that desktop has a standing preference the agent is supposed to remember, an inbox whose history changes what the right action is, or a credential that makes a wrong action expensive rather than merely incorrect.

The machine is also fresh. Everything the agent needs is in the task, and everything it learns dies with the run. The products on this bench carry state between errands on purpose, and the interesting failures are the ones where yesterday's context produces today's wrong answer. Our teardowns call that axis what it remembers, and OSWorld has nothing to say about it by design.

What we take from it

Three things, and we took them before reading the 2.0 paper, which is part of why this page exists.

Report completion and progress separately rather than blending them. Keep the horizon in the measurement, because an errand that is abandoned at step 459 is a different failure from one that is refused at step 2, and a single rate erases the difference. Put the version and date of the task set next to every number, since this benchmark has now demonstrated what happens when you do not.

What we do not take is the step budget as a ceiling. A personal agent that runs out of patience mid-errand does not stop cleanly, it leaves a half-booked flight behind, and that consequence is something our scoring has to catch rather than treat as a timeout.

FAQ

How many tasks does OSWorld 2.0 have? 108 long-horizon workflows, spanning seven professional domains and 21 sub-categories.

How long are the tasks? The median takes a skilled human about 1.6 hours, and 69.6% take more than an hour. An average run uses around 318 tool calls, against roughly 30 in version 1.0.

What is the difference between binary accuracy and partial? Binary accuracy is the share of tasks completed outright. Partial is the share of intermediate checkpoints reached. At the top of the board they are 44.33% and 77.67%.

Is the older version still relevant? The first version remains online as an archive and was upgraded to a verified edition in July 2025. Scores from it are not comparable to 2.0 scores, and the site keeps the two on separate addresses rather than overwriting the old numbers.

Does it publish confidence intervals? No. The board reports point estimates with no margin, so small differences between adjacent rows cannot be distinguished from run-to-run variation. The cost per task column is also present but empty for the leading entries, so the board cannot currently be read for price.

Is a desktop benchmark relevant to a personal assistant? Partly. Several products on this bench do their work on a machine of their own, so the skill being measured is real. What is missing is memory, standing instructions, and the fact that the accounts being touched belong to someone.

The other benchmark pages

No numbers invented. Personal agents taken apart at the runtime level, on a method built to take more.

See the scorecard