Personal Agent Bench

How the next batch of tasks was drawn

A personal agent benchmark can decide its results without changing a single grading rule, simply by choosing which tasks to run. Pick the errands one product is good at and it wins. Pick the ones another product trips on and it loses. So this page publishes the rule that chose the tasks for our next baseline batch, before any results from it exist, along with what the rule can and can't protect against.

  • The draw started from the exam's full library of 312 tasks and removed only the tasks the exam can't run fairly for every product.
  • Tasks were grouped by what the agent actually has to do, and no group contributes more than two tasks. Nobody wins by being good at one errand repeated thirty times.
  • The draw used a fixed random seed chosen on 4 October 2026 and never looked at any product's results.
  • The list holds 169 tasks, each to be run twice for every product in the exam. That's enough for a rate with an honest interval, not enough to rank anyone.
  • Every limit of the draw is listed below, including the worlds it ended up weighting most heavily.

Why the draw is published before the results

The scorecard's rule has always been to publish how something is measured before publishing the measurement. Grading rules are the obvious case, and they are on how the exam works. Task selection is the less obvious one, and it is at least as powerful.

Every task in the library is a legitimate errand. That's what makes selection dangerous: a set chosen to flatter one product doesn't contain a single bad task, so nobody looking at it can point to the trick. The only defence is to fix the rule in advance, apply it mechanically, and say what it was. If the results then favour someone, they favour someone under a rule that was written down before anyone knew which way it would cut.

Start from everything, then remove what can't be run fairly

The library held 312 tasks when the draw was made. Three were removed because they need real accounts on real services and an open connection to the public internet. Those can't be reset to the same starting point for every product, and a task that starts differently for each product isn't the same task.

One more was removed after the draw: the task that has to get past a page checking for automated browsers. The browser the exam supplies can't disguise that it's automated, so every product using it would fail that task for a reason that belongs to us. The browser tasks page covers that task and the four browser tasks that stayed in.

Nothing else was removed. In particular, no task was removed because a product did badly on it in rehearsal, and no task was added because a product did well.

Group tasks by what the agent has to do

The library has many near-relatives. One booking task asks for a table for four, another for a table for two on a different night, a third adds an allergy. They exercise the same skill against slightly different facts. If they were drawn freely, whichever skill has the most variations would dominate the batch, and a product strong at that one skill would look strong overall.

So before drawing, every task was placed in a group by what the agent actually has to do. Two tasks are in the same group when they run in the same world and the agent's standard solution makes the same set of changes to it. Reads don't count toward this, because nearly every task starts by reading something. What defines the group is what the agent is supposed to change.

Tasks graded only on the agent's reply, with no changes to make, were grouped by world and then split further by what they need from the agent, such as remembering something across sessions, and by which faults they inject. Without that split they would all collapse into one huge group, and memory tasks in particular would crowd each other out.

The 309 remaining tasks fell into 128 groups.

At most two per group

From each group, at most two tasks were taken. A group with only one task contributes that one. A group with forty contributes two.

The cost of this is deliberate. A skill that the library tests forty different ways now counts for as much in the batch as a skill it tests twice. That under-weights whatever the library happens to have the most of, which is exactly the point: how many variations someone wrote for a skill says more about the authors than about how often you'll need it.

Two rather than one gives each group a second task with different facts, so a product doesn't pass a whole skill on the strength of one lucky reading. Raising the limit to three was considered as a way to grow the sample later, and it is recorded as an option rather than the current rule.

The second pick fills a gap

Within a group, the first task is simply the first one after a seeded shuffle. The second is chosen to cover something the first didn't.

The exam tags tasks with the kinds of difficulty they carry: graded only on the reply, has an injected fault, uses a connector, requires the agent to act without being asked, spans more than one session, needs answers from the simulated user, needs approval, or tests a safety boundary. If the first task in a group lacks one of these and another task in the group has it, that task is taken second. If no task fills a gap, the second task is again just the next in the shuffle.

The effect is that the two tasks drawn from a group are as different as the group allows, without anyone choosing them by hand.

A fixed seed, and no look at the results

The shuffle used a single fixed seed, the date of the draw written as a number. The seed, the grouping rule and the gap-filling rule are written at the top of the task list itself, so the list can be checked against the rule that produced it.

What the rule never reads is any product's result. It doesn't know which tasks were hard for whom in rehearsal, which groups any product passed, or which world anyone is strongest in. That is the property that matters most, and it's the reason the rule is mechanical rather than a judgement about which tasks are representative. A judgement can be honest and still lean, and nobody can tell afterwards whether it did.

What the draw looks like

The list holds 169 tasks. Ten of the exam's eleven worlds are represented, the eleventh being the robot-check world left out above, and the distribution is uneven:

WorldTasks drawn
Work files and documents67
Email and calendar30
Bookings and tickets20
Life admin and bills15
Shopping and after-sales15
Travel13
Live data streams5
Browser skills, across three small worlds4

Work files make up 40% of the draw, and that follows directly from the rule. The cap is per group, not per world, so a world with many distinct kinds of change gets many groups and therefore many tasks. Work file tasks involve many distinct kinds of change. Live data tasks involve few. We could have forced the worlds into equal shares, and that would have meant overriding the mechanical rule with a judgement about which worlds matter, which is the thing the rule exists to avoid. Instead the shares are printed here so you can discount them yourself. A product strong at work files will look better overall in this batch than in one balanced by world, and that should be read as a property of the draw.

What this draw can and can't support

Each task will be run twice for every product. Two attempts at each of 169 tasks gives several hundred attempts per product, enough for a completion rate with a confidence interval tight enough to be worth printing. It is not enough to separate products that are close, and the scorecard will say so in the interval rather than in a ranking.

It also can't tell you about any one task. Two attempts at a single errand is an anecdote, and per-task results will be shown as what happened, never as a rate. And it can't tell you about errands the library doesn't contain. The groups cover what the exam knows how to test, which is a subset of what you'll actually ask an agent to do.

How this relates to the numbers on the scorecard now

The completion rates on the scorecard today were measured on 22 September. That was before this rule existed, on a different set of tasks, under the earlier grading rules. They are not a smaller version of this batch, and they shouldn't be read as a preview of it.

When the batch drawn by this rule reports, it will start a new series rather than extend the old one. The two will be shown side by side with their dates, their task sets and their rules, and never joined into a single line. A product that moves between them may have changed, or the tasks may have, or the rules may have, and a chart that joined them would hide which. Keeping them apart is less satisfying to look at and is the only version that doesn't mislead.

FAQ

Could the draw still favour a product by accident? Yes, and the world shares above are the clearest way it could. The rule guarantees the selection wasn't steered by results. It doesn't guarantee the result is neutral between products with different strengths, and no selection rule could.

Why not run every task? Because the library is uneven in the way described above, and running all of it would make the batch a measure of which skills have the most variations. Running every task twice for every product would also cost nearly twice as much, for a result that's more lopsided rather than more accurate.

Will the draw change? When it does, the new rule will be published before its results, and the change will be dated on what changed in the exam.

Can I see which tasks were drawn? The list itself isn't published on this site. What's published here is the rule, the counts and the shares. The errand pages describe the kinds of work the tasks cover.

Where the draw sits in the exam

The grading rules are on how the exam works, how the exam's own judgements are checked is on how we check our own grading, and every change to either, including this draw, is dated on what changed in the exam.