AssistantBench, and why a similar name is not a similar thing
AssistantBench is an academic benchmark for web agents, built at Tel Aviv University with collaborators at Penn, the Allen Institute, Washington and Princeton. It holds 214 tasks drawn from more than 525 pages across 258 websites, and it asks a single question: can an agent go out onto the open web and come back with the right answer. It is on this list mostly because its name sits next to ours in search results, and because it does one thing with its scoring that we think is correct and almost nobody copies.
- It is a question-answering benchmark on the live web. The agent finds things. It does not change anything.
- It lets an agent decline to answer, and reports that separately, which rewards admitting failure instead of inventing a result.
- Its headline accuracy carries partial credit, and the exact-match figure beneath it is roughly half as large.
- Its public leaderboard currently mixes finished systems with leftover test submissions, which is worth knowing before quoting a position on it.
What it actually measures
A task is a realistic information-gathering request that would take a person a while: the kind of thing where the answer exists on the web but is spread over several pages and needs assembling. The agent browses, gathers, and produces an answer. Scoring compares that answer against a gold answer, with credit available for being partly right on structured answers rather than only for an exact string.
The tasks came out of the authors' own observation that most web-agent benchmarks were either synthetic or trivially short. 214 tasks is small by modern standards, and the set is fixed, which makes it cheap to run and easy to saturate. The accompanying paper also introduces an agent of the authors' own, built by adding planning and memory components to an existing web agent, which tells you what they believed the hard part was.
The leaderboard
| # | Submission | Organisation | Overall | Exact match |
|---|---|---|---|---|
| 1 | Autobot ChatGPT-HQ + GPT-5.6 Sol Ultra | Autonomous Production | 50.70 | 27.1 |
| 2 | Jina + Claude Code | Om Labs | 42.10 | 22.7 |
| 3 | Distyl Button | Distyl AI Inc. | 38.00 | 19.9 |
| 4 | TongAgents | BIGAI | 37.60 | 18.2 |
| 5 | test_Agent | 36.70 | 18.2 |
Letting an agent say it does not know
Three columns here do something unusual. Accuracy is scored over all tasks. Answer rate is the share the agent chose to answer at all. Precision is accuracy among the ones it answered.
Splitting those apart changes the incentive completely. On a benchmark that scores only accuracy, a guess is free: a wrong answer costs the same as no answer, so the optimal strategy is to always produce something. Reporting answer rate separately makes abstention visible and makes a confidently wrong answer look worse than an honest refusal.
That is the right design, and it is the single most transferable thing on this page. The equivalent failure for a personal agent is worse than a wrong answer to a question. An assistant that cannot find your booking reference and invents one has not merely failed, it has handed you something that looks like success and will not be discovered until it matters. Our own scoring treats a false claim of completion as its own category rather than folding it into a resolution rate, for exactly this reason.
Partial credit, and what the exact-match column says
The leading entry scores 50.70 on accuracy and 27.1 on exact match.
Both are real numbers about the same runs. The first says the agent produced answers that were substantially right about half the time. The second says it produced the precisely correct answer about a quarter of the time. Which one you should care about depends entirely on what you were going to do with the answer, and a headline that quotes only the first is not wrong so much as incomplete.
The difficulty breakdown underneath makes the same point more sharply. That same entry scores 88.0 on easy tasks, 67.7 on medium and 39.7 on hard. An average across those three is a statement about the mix of the task set as much as about the agent.
What the current board looks like
This part needs stating plainly, because a leaderboard position is the kind of thing that gets quoted.
The board is a public submission queue, and it has not been pruned. The entry sitting fifth is named test_Agent and lists no organisation. Below it are rows called reb_1, reb3, reb4, check_1, check_state, tt_api_1, wsg_v1 and several variants of SO_R3_Testing. Three separate rows carry identical accuracy, exact-match and difficulty-breakdown figures while differing only in answer rate, which is the signature of one system submitted repeatedly under different names.
None of that is misconduct. It is what an open submission form looks like after two years without a janitor, and the underlying scores are still produced by the benchmark's own scorer rather than self-reported. But a position on this board is a position among submissions, not among products, and the distance between those two things is large here.
The hosted leaderboard also did not render a table in a normal browser session when this page was last verified. The figures above were read out of the application's own configuration, which is the same data the page is built from. That is a presentation problem rather than a data problem, and it is recorded here because the alternative would be to quote numbers without saying how we got them.
A small fixed task set ages in a particular way
214 tasks on the live web is a set that was cheap to build and is now hard to keep honest, and the reason is not the benchmark's fault.
The answers live on real websites. Pages get rewritten, listings expire, prices change, and the gold answer recorded at publication slowly stops matching what is out there. A frozen answer key over a moving corpus drifts in one direction: it gets harder to score correctly, and some of what looks like agent error is the world having moved. Nothing on the board distinguishes those two.
The same small size makes the top of the board volatile. A single difficulty band holds enough tasks that a handful of items decides several points of accuracy, and with no confidence intervals published, adjacent rows cannot be told apart. This is the same objection we make to our own early figures, which is why the rates on our scorecard carry a range and a count rather than standing alone.
It also explains the shape of the leading entries. The top two are not research systems but products, scoring well above the academic baselines the benchmark shipped with. When a fixed public set becomes a thing worth winning, it stops measuring general capability and starts measuring attention paid to it, and that transition is usually invisible from the board itself.
Where it stops being about personal agents
AssistantBench is a retrieval benchmark wearing an assistant's name.
Every task ends with an answer. Nothing is booked, cancelled, sent, paid for or undone. There is no account to sign into, no message that reaches another person, no action that cannot be taken back. The hardest thing on the board is finding something difficult on the open web and reporting it accurately, which is a genuine skill and a small fraction of what an assistant is asked for.
The tasks are also one-shot. Each arrives with no history behind it and leaves no trace after it. The products this site examines are supposed to carry a preference you stated last month into a decision made today, and the most interesting things they get wrong come from that carry-over rather than from any single lookup. Our teardowns spend most of their length on it.
Why a page about it at all
Two reasons, and the first is uncomfortable enough to say out loud.
The names are close. Someone looking for a benchmark of personal assistants will find a benchmark of web question-answering, and the two produce completely different impressions of how capable these systems are. A 50% accuracy figure for finding information on the web does not mean half of your errands will get done, and we would rather the comparison be explicit here than be made silently by a reader.
The second reason is the abstention design. It is the one piece of scoring machinery on this list that directly addresses the failure mode we care most about, and it came from an academic group in 2024 rather than from anyone building products. It deserves to be copied more than it has been. The full method we run is written up in how the exam works, and the part that treats a confident false completion as its own outcome rather than as a near miss owes something to this benchmark's framing.
What we took and what we left
The abstention split is in our scoring. An assistant that reports it could not complete an errand is recorded as having not completed it, and an assistant that reports completing one it did not is recorded as something else entirely, because those two are not the same failure and averaging them together destroys the only distinction a reader actually needs.
What we left is the fixed public task set. Publishing every task would let anyone reproduce our numbers, which is the argument for it, and would also let a vendor tune against them, which is the argument against. Our compromise is to publish the method, the per-run records and the scoring code while keeping the live task set closed, so a figure can be rechecked without being gamed. That is a weaker form of reproducibility than this benchmark offers, and it is a deliberate trade rather than an oversight.
FAQ
How many tasks does AssistantBench have? 214, drawn from more than 525 pages across 258 different websites.
What is the difference between accuracy and exact match? Accuracy awards partial credit on structured answers. Exact match requires the precisely correct answer. At the top of the board those figures are 50.70 and 27.1 for the same runs.
What does answer rate mean? The share of tasks the agent chose to answer rather than abstain on. Precision is accuracy calculated only over the tasks it answered, so an agent that declines the ones it cannot do scores better on precision than on accuracy.
Can anyone submit to the leaderboard? Yes. Submissions are prediction files scored by the benchmark's own scorer. The queue has not been cleaned up, so finished systems and leftover test runs sit on the same board.
Is it still maintained? The leaderboard application was updated in September 2026 and new entries continue to appear. The task set itself has not changed since release.
Does it test anything a personal assistant does? It tests finding information on the live web, accurately, without inventing an answer. That is a real part of the job. It does not test acting, remembering, or anything irreversible.
The other benchmark pages
- GAIAEvery question arrives cold. Nothing carries over from the last one, so a system that remembers you scores exactly the same as one that doesn't.
- BrowseCompReading, not doing. Nothing is booked, cancelled, paid for or sent.
- WebArenaThe sites aren't yours and hold nothing about you. No saved cards, no order history, no account that remembers.
- OSWorldSkill, not memory. It asks whether an agent can operate the machine, never whether it knows why you'd want it to.
- Tau-benchThe user is simulated and the domains are narrow, but this is the closest thing on the list to what a personal agent actually does.
- AgentBenchNone of the environments is your life. No calendar, no inbox, no bill to dispute.
- Terminal-BenchAlmost nothing a personal agent does happens at a terminal.
- SWE-benchA closed problem with a written specification. An errand has neither, and it cannot be rolled back.
No numbers invented. Personal agents taken apart at the runtime level, on a method built to take more.
See the scorecard