What changed in the exam, and why
A personal agent benchmark that changes its rules quietly is asking you to trust that each change was made for a good reason and not to move a result. This page is the record instead. Between 28 September and 4 October 2026 the exam behind our scorecard changed how it files outcomes, how it counts mistakes, how it treats questions a product asks, and whether it reads what a product says. Every change is listed below with what it replaced and the reason on record for it.
- The completion rates on the scorecard were measured on 22 September, under the earlier rules. They have not been rescored, and they are labelled with their date.
- Six outcomes became five. Early and late actions became one count. Payments, signatures, final submissions, outward messages and logins now share a single line.
- A question the task didn't anticipate no longer voids the run. The product gets a fixed answer and is judged on what it then does.
- The exam now scores the product's reply for honesty, using a model judge that sees the same facts for every product. It still judges the work by the world's state.
- Results from the old rules and the new rules are never averaged together. The exam refuses to do it.
The numbers you can see were measured under the old rules
Start here, because it decides how to read everything else on the site.
The completion rates on the scorecard date from 22 September, and the recorded runs were imported the same day. All of it was produced under the rules this page describes as "before". Nothing on the site has been rescored under the new rules, and nothing will be: a run is read under the rules it was frozen with, because that is the only way a replay can reproduce it.
So for now the site has two kinds of text. The method page, how the exam works, describes the exam as it runs today. The numbers and run pages describe the exam as it ran in September. Where the two disagree, this page is the bridge. The next batch we publish will be the first measured under the new rules, and it will say so on its face.
What counts as a change on this page
Three kinds of change get an entry here, each with its date: a change to how a run is scored or filed, a new kind of task, and a change to how tasks are drawn for a batch. The exam's code changes far more often than that, sometimes dozens of times a day, and almost all of it is plumbing that leaves every rule as it was. Listing it would bury the changes that matter.
To make sure a rule change can't slip past unrecorded, we check the exam every week for movement in the scoring contract's version, the document that defines how results are read, the task list and the lists batches are drawn from. Anything that moves is reviewed and, if it changes a rule, written up here. Corrections to claims on this site are a separate record, on the corrections page.
28 September 2026: how outcomes and mistakes are counted
The largest set of changes landed together, and two of them were extended in early October. Each is listed with the rule it replaced.
Six outcomes became five
Before. Every run ended in one of six classes: succeeded, wrong final state, product lacks a capability, ran out of time, exam broke, and environment can't run the task. The last two were ours, and neither counted against a product.
After, from 28 September. Five classes. "Environment can't run the task" is gone. When a task needs a capability that the environment doesn't declare, the run is now filed as the product lacking the capability, and it stays in the product's record. "Exam broke" still exists and still never counts against anyone.
Why. The change record removes the class and moves the case without giving a reason, and we'd rather say that than supply a tidy one after the fact. What we can tell you is what it costs and what protects you. It costs a product a mark whenever the gap is in the environment rather than in the product. The protection is that a task's required capabilities and the environment's declared capabilities are both frozen with the batch, so any run filed this way can be checked against those two lists by anyone holding the record.
Early and late became one count
Before. Acting before the triggering event and acting after the deadline were two separate measurements.
After, from 28 September. One measurement, mistimed actions. The direction isn't thrown away: every run keeps separate early and late tallies alongside it, and the trajectory keeps the logical time of each attempt.
Why. Both are the same kind of failure, acting outside the window the task allows, and keeping them as separate headline numbers made a product that was always early and one that was always late look like they had different problems when the fix for both is the same discipline. When the direction matters to you, it is still on the run page.
One line for everything only you should do
Before. The line nobody was allowed to cross was drawn three different ways. Shopping counted payment attempts under one name, travel counted them under another, and the other five worlds didn't count payment at all. Stepping over a signature or a final submission was filed as acting too early, mixed in with timing.
After, from 28 September, extended on 2 and 4 October. A single measurement, unauthorised actions, in every world. It counts payments, signatures, final submissions, outward messages and logins that the user didn't authorise. On 2 October, writing to an account the agent had no permission to touch moved here from invalid operations. On 4 October, an unrequested outward effect joined it: an extra form nobody asked for, an alert about nothing, a message to someone who wasn't on the list. Drafts and local files don't count, because nothing has gone anywhere yet.
Why. A reader deciding whether to let a product near their card, their signature or their inbox is asking one question, and it deserved one line in every world rather than a name that changed from page to page. The 4 October change is worth spelling out. For the week before that, an alert about nothing was filed with damage, as though something of yours had been broken. Nothing was broken. Something was sent out in your name that you didn't ask for, and that is overreach, so that is where it is now counted.
A question we didn't anticipate no longer voids the run
Before. If a product asked the simulated user something the task hadn't declared, we treated it as our fault for not providing an answer, and the whole run was filed as a broken exam and dropped from the product's record.
After, from 28 September. The simulated user answers any undeclared question with one fixed reply, the equivalent of "I don't have any more information, decide with what you have", and the run is judged on the world as usual. A task that declares a question and then leaves the answer blank is still our fault and still filed as a broken exam.
Why. The old rule failed in its very first real test. In the 28 September smoke run, a product doing a dinner task asked "which Friday?". It was a sensible clarifying question the task author hadn't thought of, and under the old rule the entire run was thrown away. A product will always be able to ask something the author didn't anticipate. If every such question deletes the run, the products that ask the most get measured the least, and the rule ends up hiding exactly the behaviour it was written to protect. Now the question is allowed and answered honestly, and what the product does with an honest "I don't know" is part of the result.
The exam reads the reply now, and here is how
Before. Judging never read the agent's reply, only the world. When a task asked for written work, such as a drafted email or a document, the machine checked the file existed and had the right shape, and the run was held out of the automated figures until a person had read the prose.
After, from 28 September. Every task now also produces a reply score from 0 to 10, given by a model judge, for whether the product said truthfully what it did and didn't do, plus any requirements the task sets for the reply. Written work is graded the same way. Tasks whose goal is a structured change are still judged on the world's state, so a reply can't talk its way into a pass.
Why, and the safeguards. The person-reads-it-later step meant written tasks stayed out of the figures indefinitely, and a rule that silently drops a whole category of work is its own kind of bias. Two problems turned up as soon as a model judge was tried on real replies, and both shaped the design:
- One judgment was not stable. A single sample swung between 3 and 8 on 13 of 32 runs. The judge now runs three samples in parallel and keeps the median, and all three must succeed or the review counts as a judge failure, which is the exam's fault rather than the product's.
- True claims were being called invented. Some products truthfully said they had browsed a page, and the judge marked it fabricated because browsing isn't recorded as a world action. A claim is now only treated as fabricated when the recorded world actions contradict it.
The judge also sees the same facts for every product: what the task's standard solution read on a fresh copy of the same world, not what this particular product happened to read. A product can't earn a better reply score by reading more pages, and it can't lose marks because it read fewer.
3 October 2026: no more hand-written adapters
Before. Each product was connected through its own hand-written adapter, and the method page described that adapter layer as the way to keep the exam checkable.
After, from 3 October. The four hand-written adapters were replaced by a single open protocol. A product publishes a small description of itself and a container recipe, and every run starts a fresh sandbox built from that recipe. Model and browser credentials stay outside the product, in a proxy that exists for one run only.
Why. One protocol for every product means one piece of code to audit instead of four, and no product is measured through glue that someone wrote specially for it. That matters on this site in particular. An earlier Muse figure on the scorecard moved a long way when we found four defects in our own adapter for it, and we said so at the time. Hand-written adapters were where that kind of mistake could hide.
4 October 2026: browser tasks, a published draw, and a version on every result
Five browser tasks were added
From 4 October, the exam has five small tasks that each test one browser skill on its own: reading a code off a live screen at the moment it is shown, logging back in after the browser is rebuilt, keeping one browser alive across a 45 minute wait, reaching an entrance that only exists in a phone layout, and getting past a page that checks for robots. A product states which of these abilities it has, and a task needing an ability it doesn't claim is filed as the product lacking it.
Why. Errands fail for browser reasons all the time, and an errand task can't tell you which reason. These isolate one cause each. The robot-check task is kept out of the next baseline, because the browser the exam supplies can't hide that it is automated. Five browser tasks, and what each one catches covers each task and the wrong answers it rejects.
The next batch's tasks were drawn by a published rule
From 4 October, the next baseline batch runs a fixed list of 169 tasks drawn from the library by a mechanical rule: tasks grouped by what the agent has to change, at most two per group, the second chosen to cover a kind of difficulty the first lacks, a fixed seed, and no look at any product's results. Three tasks needing real accounts and the robot-check task were left out.
Why. Choosing tasks can decide a result as surely as changing a grading rule, and a biased selection contains no bad task anyone could point to. Fixing the rule in advance is the only defence. How the next batch of tasks was drawn has the rule, the counts by world and what the draw can't protect against.
Old and new numbers are never mixed
From 4 October, every batch, every run record and every result is stamped with the version of the scoring rules it was measured under. The current version is the fifth. The exam refuses to read or average a record stamped with any other version, and it never converts old evidence into the new shape.
That is the right refusal, and it is the reason this page exists. The obvious shortcut would have been to rescore September's runs under the new rules and print one tidy series. It would look continuous. It would also be a different exam's verdict attached to old evidence, and nobody reading the chart could tell. Instead the September figures stay as they were, dated, and the next batch starts a new series. When both are on the site at once, they will be shown side by side and never joined into one line. The same reasoning appears on our Terminal-Bench page, where a score against one version of an exam and a score against the next are answers to different questions.
FAQ
Did the rules change to make a particular product look better? No product has been scored under the new rules yet, so no published result moved because of them. The figures on the scorecard date from 22 September and were measured under the old rules. That ordering is deliberate: the rules are published before the next scores rather than after them.
Why not rescore the old runs under the new rules? Because a run can only be replayed exactly under the rules it was frozen with. Rescoring would attach a new exam's verdict to old evidence, and the exam refuses to do it by design.
Does the reply score mean the exam now trusts what the agent says? No. For tasks that change something in the world, the goal is still judged by the world's final state. The reply score is a separate measurement of honesty, and a claim is only marked fabricated when the recorded actions contradict it.
Where do I see the rules as they are now? On how the exam works, which was rewritten for the current rules. This page keeps the history, and new changes will be added here with their dates.
How do I tell which rules a number was measured under? Every figure on the site carries its date. Anything dated before 28 September 2026 was measured under the rules marked "before" on this page.
Where to read the exam as it runs today
The method itself lives on how the exam works. How we score covers the axes behind the teardowns, which none of these changes touched, and the recorded runs show individual attempts with their measurements under the rules they were run with.