Personal Agent Bench

Watching something live, and knowing when not to tell you

Notification overload is usually written about as a problem with you. Turn on a focus mode, batch your alerts, unsubscribe from the noisy apps, build better habits. All of that works up to a point, and the point it stops working is the moment something genuinely does need you: a flight that just moved, a price that just fell below what you said you would pay, a score that changes in the last ten minutes of a match. The discipline that keeps the noise out keeps that out too. A personal agent proposes a different answer, which is that something else watches the stream and decides what reaches you. The personal agents on this bench all claim some version of that. Here is what the errand actually involves, and what we measure when we make them try.

  • The usual framing makes filtering your job. An agent proposes moving that job somewhere else, which is a design claim rather than a productivity tip, and a design claim can be tested.
  • Live data worlds in our exam carry a measurement the other worlds do not have: whether the agent interrupted you over something that did not warrant it.
  • Urgent events and noise events are judged separately and never cancel out. Alerting on everything does catch the urgent one and earns nothing for it.
  • Our exam judges the world, not the wording. Whether an interruption was warranted is read from structured records, not from how carefully the agent phrased the message.

Which of them can actually do this

ProductCompletedWhat it brings to this errand
Instinct43%36 to 50 · 159 tasksA clock or an event can start it, so it does not need you to ask first. That is the capability this errand is about, and it is also what makes unwarranted interruption possible.
Pine20%14 to 25 · 167 tasksReaches you by phone, which is the loudest channel on this bench and therefore the most expensive one to be wrong on.
Muse19%13 to 25 · 90 tasksWeb is where the product lives, so an alert has somewhere to sit rather than having to reach you. Its approval card puts a send in front of you before it happens.
TownWithheldWeb and a native Mac app with real system access, so it can reach notification surfaces on your own machine. It shows its assistant roughly 177 of the 1105 tools it lists.
GrokNo figure yetNot taken apart at the runtime level yet, so we have nothing established to tell you about how it would handle this.
The completion column is every task in the exam, not this one. We do not have per-errand figures yet, and a single rate presented as one would be the kind of thing this site exists to catch. As of 2026-09-22.

Two things that table does not say, and both matter more than the percentages.

None of those figures is specific to watching a live stream. Each is a product's share of every task in the exam, which tells you how often it finishes what it is given, not how it handles this particular errand. The per errand runs are not done.

And a completion rate says nothing about how a product finishes. An agent that never misses an urgent event because it forwards everything has a perfect record on the half of this problem that is easy to measure. The rest of this page is about the other half.

Notification overload is not really a volume problem

Count the alerts on an ordinary phone and the number is large, but the number is the symptom. The problem is that they arrive undifferentiated. A delivery driver two streets away, a colleague reacting to a message from yesterday and a gate change forty minutes before boarding all land with the same weight, because each app decides its own importance and every app thinks it is important.

The tools built to fix this work at the wrong level. A focus mode filters by app, by sender or by time of day, and none of those is the thing you care about, which is whether this particular event is worth your attention right now. The message that matters is often inside the channel you muted, which is why people who organise their notifications carefully still check compulsively.

Muting is a blunt instrument used on a problem that needs judgement. That is the gap a personal agent is aiming at.

Live data makes the gap obvious because the stream never stops. A flight status, a match score, an order in transit and a price that moves all produce a continuous feed of small changes, almost all of which are meaningless, a few of which change your afternoon.

What changes when something else does the watching

An agent that can watch a stream and decide what to pass on is making two separate claims, and they are not equally hard.

The first is that it can watch at all. This is mostly an engineering question: something has to run when you are not looking at it. Several of these products can be started by a clock, an inbound email or a calendar entry rather than by you typing, which is the machinery this errand requires.

The second is that it can judge. That is not an engineering question. It requires knowing that a two minute delay on a four hour layover is nothing and a two minute delay on a fifty minute connection is everything, which depends on facts about your day that live outside the data stream. This is where the pitch either holds up or quietly becomes a rebranded alert rule.

Our exam has a participant that answers on the user's behalf for exactly this reason. It is deterministic and strictly limited to facts the task declared as hidden, and anything outside that list gets the same refusal every time. An agent that works out what matters by asking has demonstrated something real. An agent that guesses and happens to be right has demonstrated nothing.

The measurement built for live data

Every world in our exam produces the same five numbers: whether the goal was reached, whether the agent acted before it was allowed to, whether it acted after the deadline, whether it damaged something the task was watching, and how many invalid operations it attempted.

Some worlds add one of their own. The shopping worlds track whether the agent crossed the payment line. The live data worlds track whether it interrupted you over something that did not warrant it.

Those are the only two world specific measurements the method names, which tells you how the exam was weighted. Unwarranted interruption is a column, published on its own, for the same reason the payment line is: a reader deciding whether to let a product near their attention is entitled to see that cost rather than have it averaged into a figure that hides it.

Why urgent and noise never cancel out

Suppose you score alerting the obvious way, by asking whether the agent told you about the thing that mattered. Under that rule the optimal strategy is to tell you about everything. An agent that forwards every change in the stream will catch the urgent one every single time and score perfectly, while being exactly the product nobody can live with. The naive way to measure this rewards the behaviour that makes an assistant unbearable.

So the exam judges the two things separately and never lets them cancel out. An agent that messages you about everything will catch the urgent event, and it gets no credit for that, because it also interrupted you six times over nothing. Both figures stay on the page. Neither is allowed to pay for the other.

A run that caught the urgent event and generated five needless interruptions is not a partial success. It is one success and five failures, printed next to each other.

Both are read from structured records rather than from the wording of what the agent said. That detail closes the obvious escape route. An agent cannot soften a needless interruption by opening with a polite apology for the interruption, and it cannot upgrade a missed event by describing in convincing prose how closely it was monitoring. The record says an interruption happened and says whether the event warranted one. Tone is not evidence.

Two clocks, and missing either one is its own failure

Several tasks in the exam are built around an event that starts a clock: a reschedule notice arrives, a price drops, a flight is delayed. The agent then has a deadline measured from that moment, and there are two separate ways to get the timing wrong.

Acting before the event is recorded on its own, because an agent that tells you the flight is delayed before the airline has moved anything has guessed correctly rather than worked correctly. On a live stream this is a real temptation. Aircraft turn late, boards flicker, and an agent pattern matching towards a delay that has not been declared is inventing news.

Acting after the deadline is also recorded on its own, even when the final state is right. A correct alert about a gate change that arrives once the gate has closed is not a late success. It is the wrong answer delivered politely.

The jobWhat it actually meansWhere it breaks
Watch the streamKeep looking after the first lookIt read once, cached, and never checked again
Recognise the eventThe change that crosses your threshold, not every changeEvery tick looks like an event when you have no threshold
Decide whether to speakWeigh this event against interrupting youNo threshold, so everything gets forwarded
Reach you in timeBefore the window closes, not afterCorrect, complete and half an hour late

The exam pushes on the first row deliberately. A task can declare faults, and a shared controller fires them at a set logical time so the web page and the API see the same breakage at the same moment. Prices and statuses can change partway through, which turns a correct answer into a stale one if the agent read early and never looked again. On a live data task that is not an edge case, it is the whole job.

Proactivity is the capability and the hazard

An agent that can wake itself up is the only kind that can alert you. It is also the only kind that can bother you unprompted. Those are not two features that happen to sit near each other, they are the same feature described twice.

That symmetry is why this errand is the sharpest test of a personal agent rather than a nice extra. Every other task asks whether an agent can finish something you told it to do. This one asks whether it can decide, unprompted, that you should be pulled away from whatever you are doing.

There is a security shape to it as well. An agent that can be started by an inbox is an agent a stranger can address, which makes the check between something happened and act on it more consequential than any feature on a comparison table. The teardowns cover which surfaces can start each product without you typing, and what sits between the trigger and the action.

Questions worth asking any product that claims this

AskA good answer sounds likeBe wary of
What is the threshold?Something you set, in your termsIt says it uses judgement
How often does it look?A stated interval, and what happens betweenVagueness about when it is actually running
Can I see what it decided not to send?A log of suppressed eventsNo such record exists
Who can trigger it?A named list of surfacesAnything that treats an inbound message as an instruction

What the exam actually checks

Judging reads the world's final state, never the agent's account of what it did.

For a live data task that means checking which events the world produced, which of them the agent raised, and which it raised that nothing warranted. The summary is especially tempting to believe here, because a message saying it has been keeping an eye on your flight reads the same whether or not anything was watching.

MeasuredOn this errand that means
Goal reachedYou were told about the event that actually mattered
Acted too earlyIt raised an alarm before the event had happened
Acted too lateThe alert arrived after the window had closed
Collateral damageSomething else the task was watching changed
Invalid actionsWrites attempted without the authority to make them
Unwarranted interruptionIt pulled you away from your day over nothing

These are published side by side and never combined, because deciding how many needless interruptions equal one missed connection is your call rather than ours. The full method is in how the exam works, and the reasoning behind refusing a combined score is in how scoring works.

What we can and cannot tell you today

The live data worlds are built and they run. The per product results for this specific errand are not done.

We could fill this page with plausible verdicts and nobody would immediately know. That is exactly why we will not. The scorecard carries the completion rates we do have, with intervals and sample counts attached, and one product's cell is empty because the figure we were given disagreed with our own evidence. A task page claiming results it does not have would undo all of that.

One thing has changed since this page was drafted. The live data world now has a recorded attempt you can read: Instinct, four steps over about eleven minutes, finishing with the world in the wrong state. Every measurement is shown separately, including the interruption count, which came back zero. Read it as a record and not as a result. One attempt by one product is not a rate, and no batch level quality gate was written for it, so the run is not admitted to the comparison.

When the rest are done, each one gets a page showing the conversation as the product itself rendered it, the world trajectory, the pages it opened, and every measurement. Until then, what this page offers is the shape of the problem, so you can judge any product's claim about it yourself. The nearest errand written up in the same detail is claiming for a delayed flight, which starts from the same triggering event and then asks for work rather than a message.

FAQ

Can an AI agent actually reduce notification overload? It can move the filtering decision, which is a different thing from reducing the volume. Whether that helps depends on how often it interrupts you over something that did not warrant it, which is why our exam measures that separately from whether it caught the event that did. A product that is right about the urgent one and wrong six times a day has not solved anything.

Why not just use a focus mode or an alert rule? Both work at the level of app, sender or time, and the thing you care about is one particular event. An alert rule that fires when a price drops below a number is useful and it is also the easy half. The hard half is knowing that this delay matters and that one does not, which depends on facts about your day the rule does not have.

How does the bench decide an interruption was unwarranted? From structured records rather than from the agent's wording. The task declares which events were urgent and which were noise, and the exam reads what the agent raised against that. Phrasing an interruption apologetically does not make it warranted.

Does an agent that alerts me about everything score well? On the goal it does, and that is precisely why the goal is not published on its own. Urgent events and noise events are judged separately and never cancel out, so an agent that forwards everything catches the urgent one, gets no credit for catching it, and carries every needless interruption on a column of its own.

Does the bench watch real flights and real match scores? No, and deliberately so. The tasks run in simulated worlds that can be reset and replayed, so two runs are comparable and anybody can check the result. A live stream that never repeats itself would mean results nobody can reproduce.