How we check our own grading
A personal agent benchmark makes one claim more often than any other: this product got this task wrong. Every completion rate is built out of that claim, repeated a few hundred times. So before a rate means anything, somebody has to ask how often the claim itself is wrong, meaning how often the product was right and the exam was not. This page is what we found when we asked, including the parts that made our exam look bad.
- In three rounds of re-checking in September 2026, every run filed as "the product got it wrong" was re-examined by hand. In the first round only 9 of 37 held up as plain product failures.
- Our own control case was overturned in the third round. A task we had been using as a known product failure turned out to be hiding its answer from every product.
- We checked all 257 tasks for one specific defect: can a product pass by doing nothing at all. The first answer was 16 tasks.
- Replaying 128 old runs from their saved records reproduced 114 exactly. The 14 that didn't were traced to two causes, and no goal result changed.
- None of these checks produced a published score. They decide which scores we are willing to publish.
Why the grader needs grading
The exam judges a run by reading the world's final state rather than what the agent said. That rule removes the most common way agent demos mislead, which is a fluent paragraph describing work that never happened. It does not remove the possibility that the exam itself is wrong about what the world should look like.
There are three ordinary ways for that to happen. The task can expect a value that the product had no way to find. The checking code can accept something it shouldn't, so a careless product passes. Or the frozen task can drift away from the page the product actually sees, so a product that did the right thing on the real page is marked against an expectation written for a different one. Each of these produces a result that looks exactly like a product failure in a table. Nothing in the number tells you which kind of wrong it is.
A bench that never looks will report its own defects as other people's scores, for as long as it runs. So we look, and we publish what we find, because a reader deciding whether to trust a completion rate deserves to know how often the thing underneath it has been checked and how often the check came back against us.
Three rounds of re-checking "the product got it wrong"
In September 2026, ahead of the batches behind the current scorecard, the exam was rehearsed on three tables of 30 tasks each. After each table, every run filed as a product getting the task wrong was re-examined one at a time, against the world's trajectory, the gateway's access log and the product's own event stream. Each run was re-filed into one of three buckets: the product really did get it wrong, the exam was at fault, or the product simply lacks the capability the task needs.
| Round | Date | Runs filed as wrong | Product really wrong | Exam at fault | Capability the product lacks |
|---|---|---|---|---|---|
| First table | 20 September | 37 | 9 | 10 | 18 |
| Second table | 21 September | 44 | 39 | 3 | 2 |
| Third table | 21 September | 70 | 56 | 11 | 3 |
Read the first row slowly. Of 37 runs the exam had filed as a product getting it wrong, fewer than a quarter were that. Ten were the exam's fault. Eighteen were products that couldn't do what the task needed, which is a real finding but a different one, and lumping it in with wrong answers would have told you a product is careless when it is actually limited.
The second and third rounds came out much closer to the original filing, which is the point of doing it repeatedly. Between rounds the exam was fixed where the last round found faults, and each round also found something new. Every run in these tables is a single attempt per task per product, so the counts are readings of what happened, not rates of how often it happens. They are not on the scorecard and were never meant to be.
The control case that turned out to be wrong
Each round was calibrated before it was trusted. A known exam fault went in as a positive control, which the re-check had to catch. A known product failure went in as a negative control, which it had to leave alone. If either control came out wrong, the round's own judgement wasn't reliable.
For two rounds, the negative control was a workshop registration task. The task asked for five fields to be submitted accurately, once. One product had submitted an empty form, and that looked like the clearest product failure in the set.
In the third round, the re-check looked harder and the control fell apart. The task's wording described the five fields without giving their values. The page offered a free-form box with no schema. The five target values existed only inside the hidden grading check. No product could have found them by reading. Worse, the exam accepted the empty submission and recorded it as a successful form submission. A careful product that stopped and a careless one that guessed would both have been marked wrong, and for the same hidden reason. All of that task's runs were re-filed as exam faults.
That is the uncomfortable kind of result, and the useful one. A control is supposed to be the thing you are sure of. When a closer look overturns it, the right response is to say so and fix the task rather than quietly pick a new control that agrees with you.
What the re-checks found wrong with the exam
The faults on the exam's side fell into a few repeating shapes:
- An answer only the grader could see. The workshop task above. A product can't be marked wrong for failing to find a value that was never published anywhere it could look.
- A frozen task that no longer matched the page. Two housing tasks expected a particular listing and a particular appointment ID. The page the products actually saw had been rebuilt from a different dataset, with different listings and hashed IDs. One product booked a viewing with no conflict on the real page and was still scored zero.
- An empty submission recorded as a submission. Treating nothing as something lets a careless product past and leaves a careful one with no way to look different.
Each of these went back to the exam as a defect. Alongside the bucket for every run, the third round ran four cross-checks across all 70: whether the exam ever collected the paper while a product still had work in flight (it didn't), whether any follow-up message from the simulated user was triggered by a stray keyword (none were), whether every value a product needed could be reached from the opening message, the page or the published interface (two task families failed this), and whether the world's record of a request ever contradicted the response the product actually received (it never did).
Can a product pass by doing nothing?
This is the cheapest defect to miss and one of the most damaging. If a task's success check is already satisfied at the start, then every product passes it, including one that never woke up. A rate built on tasks like that rewards existing, not working.
So every task in the library was checked for it. Each of 257 tasks was loaded into a real world session, the opening page was visited the way a product would visit it, and nothing else happened. No action was taken, no message was sent, no question was answered. Then the task's own grading was run.
The first pass, on 20 September, found 16 tasks that a product could pass by doing nothing past the opening page. After the fixes, the same check on 21 September found 6. One of those six is kept on purpose as a reference case for a task that is graded by people rather than by the program. The check has a control of its own, a task that must stay off the list, and it stayed off both times.
The requirement is now part of the exam's acceptance gate. A task that has an executable standard solution must reach its goal when that solution is run, and must not reach its goal when nothing is done. Both are checked every time the grading code changes.
The grading code gets examined too
Checking tasks catches bad tasks. It doesn't catch a bug in the code that does the grading, and that code changes often. So every change to it has to pass a gate that deliberately misbehaves on purpose.
For each task with a standard solution, the gate builds variations of that solution that a product might plausibly produce. It acts one second late. It reads first. It repeats a read or a write. It drops a step. It changes one value. It hands a step to the user that the product should have done. It takes one extra action that crosses a boundary only you should cross. Each variation has to move the measurement it is supposed to move. A boundary crossing that only shows up as an invalid operation, and not as an unauthorised action, a missed goal or damage, does not count as covered.
There is also a self-test that plants seven deliberate faults in the grading code and checks that the gate notices every one. And every change is graded twice, once on the old code and once on the new, with every difference in every measurement listed. A difference doesn't fail the gate automatically, but it can't pass silently either. It stops and waits for a person to review it.
Replaying the old runs
A run is only as checkable as its record. The exam stores enough with every run to replay it: the task as it was frozen, the seed, and every action with the world's state before and after. On 28 September, 128 older runs were replayed from their records on the exact code they originally ran on.
114 came back identical. Fourteen did not, and all 14 were traced to one of two causes. Ten had the right result but a different fingerprint of the world's final state, because a reminder ID came from a random value that changed between processes, or an alert's timestamp came from the wall clock instead of the exam's clock. The other four undercounted invalid operations on replay. In every one of the fourteen, the stored goal result stood. The saved live scores were kept, and the corrections were filed separately rather than written over the originals.
FAQ
Do these checks change the published scores? Not directly. They decide which runs are trusted enough to count, and they run before a batch can be published. A batch where more than one run in five is an exam fault is refused outright, and a refused batch publishes nothing.
Why publish the exam's own mistakes? Because a completion rate is only as good as the claim inside it, that a product got a task wrong. Showing how often that claim failed, and what we did about it, is the only evidence we can give that the rates are worth reading.
Were any products named in these checks? The re-checks covered the products in the exam, one attempt per task each. We publish the method and the totals rather than per-product counts, because a single attempt per task is a reading, not a rate, and it would be wrong to present it as one.
Do the checks happen once, or every time? The acceptance gate runs whenever the grading code changes. The by-hand re-checks of individual runs were done in rounds during the September rehearsals, and we will publish future rounds here with their dates.
Where do changes to the rules themselves go? On what changed in the exam, which records every rule change with what it replaced and why.
Where these checks sit in the exam
The exam itself is described on how the exam works. Changes to its rules are dated on what changed in the exam. Corrections to claims on this site, which are a separate thing, are on the corrections page. The recorded runs show individual attempts with their measurements, and the scorecard shows what survived all of the above.