Five browser tasks, and what each one catches
On 4 October 2026 our personal agent benchmark added five tasks that test the browser rather than the errand. Each is small on purpose: a code to read off a live screen, a login to recover, a page to check twice 45 minutes apart, an entrance that only exists on a phone, a page that turns away robots. These are the places where real agents get stuck in public, and until now the exam had no task that isolated any of them.
- Each task targets one browser skill an errand silently depends on. When a product fails one, you know which skill failed, not just that the errand didn't finish.
- Every task ships with the wrong answers it must reject: a code read a minute too late, a value reused without rereading, a form filed on the desktop layout.
- A product states which browser abilities it has. If it doesn't claim one, a task that needs it is filed as "the product can't do this", never guessed.
- The page that checks for robots is left out of the next baseline. Our own browser can't hide that it's automated, and we won't blame a product for our limitation.
- No product has a published result on these yet. This page describes the tasks, not anyone's score.
Why browser work needed its own tasks
Most errands on this bench run through a browser at some point, and a surprising share of the public failure stories are really browser stories. The method page collects several: an agent locked out of a cable account even with browser takeovers, an account blocked after a couple of hundred scans, a takeover where the person had to step in at exactly the right moment. In each case the errand failed, but the errand wasn't the problem. Something underneath it was.
An errand task can't tell you which underneath thing broke. If a product fails to book a table, it might have misread the menu, lost its session, run out of time or been blocked as a bot, and the final state looks the same in every case. So the new tasks each strip an errand down to one browser skill, with a small piece of business at the end to prove the skill actually got used. A failure on one of these points at one cause.
They are also deliberately cheap to grade. Each runs in its own small world, uses the same scoring as every other task, and comes with an explicit list of wrong ways to finish that the grading must reject. Those lists are the most useful part of the design, so each section below includes them.
Reading a live code while you hold the browser
The task. A verification step shows a one-time code on a live screen. The agent has to open the live view, read the code that is showing now, enter it, and then finish the business step that the verification unlocks.
This is the shape of every handover that involves a person: the agent can get you to the door, and something about the door needs a fresh reading. The interesting failures are about freshness, not about reading.
Wrong ways to finish, all graded as not reaching the goal:
- Read the code, wait a minute, then enter it. The screen has moved on and the code is stale.
- Enter a code without ever opening the live view. Even if the digits happen to be right, the agent didn't look.
- Do nothing.
The reply is also checked: the product has to say which code it saw and what happened, and must not invent a code it never read.
Staying logged in after the browser is rebuilt
The task. The agent logs in to a members' site, looks up the most recent renewal order number, and files it on the site. Partway through, the browser is shut down and rebuilt from nothing. The agent is told this may happen and is expected to log back in, read the order number again, and only then file it.
Real browsers die. Sessions expire, sandboxes get recycled, a long task outlives the machine it started on. What matters is whether the agent notices that its picture of the world went stale when the browser did.
Wrong ways to finish:
- Reuse the order number it read before the rebuild without reading it again. The number happens to be the same here, and it still fails, because the next time it won't be.
- Restore the browser without logging in again, then try to carry on.
- Do nothing.
The first of those is the one worth dwelling on. A grader that only checked the final value would pass it. This one checks that the value was read after the rebuild, which is the difference between an agent that recovered and one that got lucky.
Holding one browser for 45 minutes
The task. A process takes 45 minutes to produce a result. The agent has to read a page at the start, read it again at least 45 minutes later in the same browser, and then file the confirmation number that appears.
This is the "keep an eye on it for me" task in its barest form, and it's the one we expect to be hardest. In our September rehearsals, every task built around watching something later failed for every product, the same way: the agent ended its work as finished before the event it was meant to wait for. That finding is on how we check our own grading. This task asks the underlying question directly: can the product keep a browser alive and come back to it?
Wrong ways to finish:
- Read twice, but only a minute apart.
- File the confirmation after the first read, without waiting for the second.
- Do nothing.
One variation is explicitly allowed: doing everything a second late still passes, and causes no damage. A grader that punishes a one-second slip is measuring the clock, not the agent. While a product is working, the exam's clock runs at real speed, so the 45 minutes are real minutes.
The entrance that only exists on a phone
The task. The user says they're on their phone and want to book through a mobile-only reservation channel. The channel only appears in the mobile layout. The agent has to switch to a phone-sized view, find the entrance and file the booking through it.
Sites increasingly hide things behind layout: a phone-only offer, a desktop-only export, a button that moves when the window narrows. An agent that only ever sees the desktop layout will tell you the option doesn't exist, and from where it's standing, it's right.
Wrong ways to finish:
- File the booking through the desktop path.
- Switch to the phone layout but book through an entrance that isn't the mobile one.
- Do nothing.
The exact screen size is the product's choice. A different phone size that still produces the mobile layout passes. The task cares about reaching the mobile entrance, not about matching a particular device.
The page that checks whether you're a robot
The task. A partner's public price page runs a check for automated browsers before it shows anything. The agent has to get through it, read the monthly renewal quote, record it on the site and report it.
Wrong ways to finish:
- Record the value without opening the page at all.
- Read the page but record a number off by one.
- Get through with an ordinary automated browser that announces itself as automated. The page sees it, and the quote is never legitimately read.
- Do nothing.
This task is real and it is graded like the others, but it is not in the next baseline batch. The reason is ours. When the exam supplies the browser itself, that browser can't disguise the fact that it's automated, so a product using our browser would fail this task for a reason that has nothing to do with the product. We left it out rather than record our own limitation as someone else's failure. Products that bring a browser able to pass the check can still be run on it separately, and any such result will be labelled as such.
A product says what its browser can do, and is held to it
Each product connected to the exam publishes a short description of itself, and part of that description is which of three browser abilities it has: getting past robot checks, changing the screen size, and keeping one browser alive over a long wait. Logging back in after a rebuild and handing a live screen to a person are declared the same way, through the parts of the description that cover sessions and takeovers.
The rule is strict in both directions. Only an explicit yes counts. A missing answer or a no means the ability isn't there, and a task that needs it is filed as "the product can't do this". It is never inferred from a product's name, its marketing, or whether its browser is its own. And a product that claims an ability it doesn't have will fail the task on the evidence, because the task checks what actually happened.
That split matters for reading results. "Can't do this" and "got it wrong" are different findings, and on this site they're always reported as different things. A product that honestly declares it can't hold a browser for 45 minutes has told you something true about itself, which is more than a product that claims it can and then quits after one read.
What these tasks can't tell you yet
There are no results on this page, because none have been published. The next baseline batch is drawn to run each task twice per product, and when results land each one will be reported with its sample size, under the rules it was run with.
Two limits are worth knowing in advance. First, our own browser can't be resized after it starts, so the phone task depends on a product being able to start in, or switch to, a mobile layout some other way, and any comparison has to say where that mattered. Second, a task this small is a test of one skill, not of a whole errand. Passing all five doesn't mean a product can do your errands. Failing one tells you where it will probably break.
FAQ
Why only five browser tasks? Because each one isolates a single skill, and five covers the browser failures we see most often in public: stale handovers, lost sessions, long waits, layout-gated options and bot checks. More can be added, and when they are they'll be recorded on what changed in the exam with their dates.
Is getting past a bot check something an agent should do? The task uses a public price page that the user is entitled to read, and the agent records a public number. It tests whether an agent can do on the user's behalf what the user could do by hand. It is also the one task we left out of the baseline, for the reason explained above.
Does a product fail if it doesn't declare a browser ability? It's filed as "the product can't do this", which is reported separately from getting a task wrong. Not having an ability and doing a task badly are different facts about a product.
Are the 45 minutes real? Yes, while a product is working the exam's clock runs at real speed. Only our own grading checks fast-forward through the wait, and those test the grader, not the product.
Where these tasks sit in the exam
The rest of the exam is described on how the exam works, and the date these tasks were added is on what changed in the exam. How each task's wrong answers are checked against the grading code is on how we check our own grading. The errands these skills sit underneath, such as booking a table or live alerts, have pages of their own.