Personal Agent Bench

Claiming for a delayed flight, and the offer an agent should bring back to you

Flight delay compensation is the errand people most want to hand to somebody else, because it is tedious, adversarial, and easy to abandon. It is also one of the few household admin jobs where the counterparty has a direct financial interest in you giving up. There is more than one pot of money involved, they pay out under different rules, and accepting from one can quietly close the others. The personal agents on this bench claim to handle errands of this shape. Here is what the job actually involves, and what we measure when we make them try.

  • There are usually three separate pots: the airline's own obligation, your travel insurance, and whatever your card provided. They pay under different rules and they are not additive by default.
  • The reason for the delay is the whole case. Airline controllable and outside the airline's control are different worlds, and the airline is the one writing down which it was.
  • The expensive failure is accepting a voucher. Taking the offer in front of you can close the claim that was worth more, and it looks like success from the outside.
  • Our exam judges the world, not the wording. The travel worlds count crossing the payment boundary as its own measurement, separate from whether the goal was reached.

Which of them can actually do this

ProductCompletedWhat it brings to this errand
Instinct43%36 to 50 · 159 tasksIts vault fills a password without being able to read it back, so it can reach a booking portal you never hand over. It lives in a message thread, so there is no confirmation screen standing between it and an accepted offer.
Pine20%14 to 25 · 167 tasksPhones the airline for you, which is the route most likely to produce an offer on the spot. It keeps the merchant, the number, the duration and a summary, but no transcript, so what was actually offered on the call is not something you can go back and read.
Muse19%13 to 25 · 90 tasksAssembles the claim on a Linux machine that is yours, and keeps memory as files you can open, so the evidence it gathered stays readable. Its approval card puts a spend or a send in front of you before it happens.
TownWithheldWeb and a native Mac app with real system access, the broadest local reach on this bench, so it can reach boarding passes and receipts sitting on your own machine. It shows its assistant roughly 177 of the 1105 tools it lists.
GrokNo figure yetNot taken apart at the runtime level yet, so we have nothing established to tell you about how it would handle this.
The completion column is every task in the exam, not this one. We do not have per-errand figures yet, and a single rate presented as one would be the kind of thing this site exists to catch. As of 2026-09-22.

Two things that table does not say, and both matter more than the percentages.

None of those figures is specific to claiming for a delay. Each is a product's share of every task in the exam, which tells you how often it finishes what it is given, not how it handles this particular errand. The per errand runs are not done.

And a completion rate says nothing about how a product finishes. An agent that closes your claim by accepting a travel voucher has completed the task and may have given up the cash. The rest of this page is about that gap.

Three pots of money, and only one is the airline's

The airline may owe you something under the rules of wherever you flew from or to. Those rules differ by jurisdiction, they turn on the cause and the length of the delay, and the amounts are set by regulation rather than by goodwill. This is the pot people mean when they say compensation.

Your travel insurance may owe you something under a policy you bought, on terms that are entirely contractual and usually require receipts. Your card may have provided cover as a benefit, with its own claim process and its own limits.

These are separate claims against separate parties. Some policies reduce their payout by whatever the airline already gave you, which means the order you claim in can change the total. Nobody involved is going to point that out.

The practical question for delegation is not whether an agent can fill a form. It is whether it knows there are three forms and which one to file first.

Because the rules genuinely vary by route and by policy, an honest agent should be telling you which pot it is working on and on what basis. Anything that just says it is handling the claim has skipped the only decision that mattered.

The reason for the delay decides the case

Almost every regime draws the same line somewhere: delays within the airline's control are compensable, and delays outside it are not. Weather, air traffic control, and security events generally sit on the other side of that line.

The awkward part is that the airline records the reason. The stated cause is the input to your claim and it comes from the party paying the claim. Cases turn on it, and challenging it means finding evidence from somewhere else about what happened that day.

For an agent this is the difference between a data entry task and an actual case. Reading the airline's stated reason and repeating it back is the easy path and it concedes the argument before it starts.

What the errand really involves

Strip a delay claim down and there are four jobs inside it. Products tend to be confident about the first two and quiet about the last two.

The jobWhat it actually meansWhere it breaks
Establish the factsScheduled time, actual time, the stated causeThe booking is under a reference nobody kept
Pick the potAirline obligation, insurance, or card benefitClaiming in the wrong order reduces the total
File itA portal, a form, a document uploadThe portal wants a boarding pass you never printed
Handle the responseA refusal, a request for more, or an offerThe offer is the decision point, and it is yours

The first two are research problems. The last two are permission problems, and permission problems are where delegating your admin either works or quietly does not.

The voucher that closes the claim

Here is the version of this errand that costs people money.

Airlines frequently respond with an offer that is not the regulated amount: a travel voucher, loyalty points, a credit toward a future booking. These are often worth less than the cash entitlement, they frequently expire, and accepting one can be treated as settling the claim.

An agent told to get compensation, offered a voucher, and reasoning that a voucher is better than nothing has substituted its judgement for yours on a question about your money. It will report success. The record will show a closed case. You will find out what it cost when you try to claim the rest.

This is exactly why the travel worlds in our exam carry a payment boundary measurement on top of the universal five. Committing you to a settlement, accepting compensation in a form you did not choose, or agreeing to terms that close a claim all sit on the far side of that boundary. An agent can prepare every part of the claim and bring the offer to you. Stepping over it on its own is recorded as acting outside its boundary, and that stays separate from whether the claim was resolved.

There is a second, quieter version. Some third party claim services take a substantial percentage of whatever they recover. An agent that routes your claim through one has made a pricing decision on your behalf.

Claiming twice from two places

Most people assume the risk here is getting nothing. The more awkward version is claiming the same loss from two parties at once.

An agent that files with the airline, hears nothing for a week, and then files an insurance claim for the same delay has two processes running against one event. If both pay, you may have been compensated twice for a single loss, which insurers treat as something to unwind rather than a windfall. The paperwork lands on you.

Every world in our exam counts duplicate side effects as a measurement of its own, alongside whether the goal was reached. A run can be recorded as a success and still carry a duplicate, and both facts appear on the page. Collapsing them into one verdict would hide precisely the behaviour you would be cleaning up after.

It is the failure most likely to be invisible in a demo, because the agent's summary describes the outcome it intended rather than how many claims it opened to get there.

When the status changes under you

The instructive version of this task is the one where the facts move while the agent is working.

A task in our exam can declare faults, and a shared controller fires them at a set logical time so that the web page and the API both see the same breakage at the same moment. A page can start returning errors for a while. Content can come back truncated, with the truncation marked, so an agent that fills in the missing part from general knowledge is making it up rather than reading it. Prices and order statuses can change partway through, which turns a correct answer into a stale one if the agent read early and never looked again.

Delays are built for this. A flight status is not a fixed fact, it is a running one. An agent that reads a two hour delay, assembles a claim on that basis, and never looks again will file against a number that stopped being true. The world knows. The agent's summary does not.

There is also a clock. Several tasks are built around an event that starts one, and the agent has a deadline measured from that moment. Acting after the deadline is recorded as its own failure even when the final state is right. Acting before the triggering event is recorded separately, because an agent that files a claim before the delay is confirmed has guessed rather than worked.

What the exam actually checks

Judging reads the world's final state, never the agent's account of what it did.

For a delay claim that means checking whether a claim record exists, which party it was filed against, what was accepted, and whether anything else changed that should not have. The agent's closing message is not evidence. That single rule removes the most common failure in agent demos, which is a fluent paragraph describing work that did not happen.

MeasuredOn this errand that means
Goal reachedA claim actually exists in the record, filed against the right party
Acted too earlyFiled before the delay was confirmed, or before you approved
Acted too lateThe filing window closed while it was still working
Collateral damageSomething else in the booking changed that nobody asked about
Invalid actionsWrites attempted without the authority to make them
Payment boundaryIt accepted a voucher or a settlement on your behalf
Duplicate side effectsTwo claims opened for one delay

These are published side by side and never combined, because deciding how many unauthorised settlements equal one missed deadline is your call rather than ours. The full method is in how the exam works, and the reasoning behind refusing a combined score is in how scoring works.

Questions worth asking any product that claims this

AskA good answer sounds likeBe wary of
Which pot are you claiming from?It names the party and the basisIt says it will handle the claim
What will you do with an offer?It brings it back to you unacceptedIt optimises for closing the case
Where is the proof?It points at a claim reference or a paymentIt describes its own diligence instead
Can it act without reading my secrets?A specific answer either wayVagueness about where credentials live

Some of these products fill a credential without ever being able to read it back. That is a real design choice with real consequences, and the teardowns cover where each one keeps what it knows.

What we can and cannot tell you today

The travel worlds are built and they run. The product results for this specific errand are not done.

We could fill this page with plausible verdicts and nobody would immediately know. That is exactly why we will not. The scorecard carries the completion rates we do have, with intervals and sample counts attached, and one product's cell is empty because the figure we were given disagreed with our own evidence. A task page claiming results it does not have would undo all of that.

One thing has changed since this page was drafted. The travel world now has a recorded attempt you can read, on a different errand inside the same world: Instinct, sixteen steps, about ten minutes, finishing with the world in the wrong state. Every measurement on it is shown separately, including the ones that came back zero. Read it as a record and not as a result. One attempt by one product is not a rate, and no batch level quality gate was written for it, so the run is not admitted to the comparison. The run page says so itself.

When the runs are done, each one gets a page showing the conversation as the product itself rendered it, the world trajectory, the pages it opened, and every measurement including the runs that went nowhere. Until then, the useful thing this page offers is the shape of the problem. The nearest errand we have written up in the same detail is disputing a charge, which shares the two routes problem and most of the payment risk.

FAQ

Am I owed compensation for a delayed flight? It depends on where you flew from and to, how long the delay was, and what caused it. Most regimes draw the line at whether the delay was within the airline's control, with weather and air traffic control generally falling outside it. Because the rules vary by route, check the regime that applies to your specific flight rather than a general figure.

Can an AI agent claim flight delay compensation for me? Several claim to handle errands of this shape. The part worth checking before you delegate is what the product will do when the airline responds with a voucher instead of cash. Accepting it can close the claim that was worth more, so an agent should bring that offer back rather than resolve it, and our exam measures acting outside that boundary as its own result.

Can I claim from the airline and my travel insurance? Sometimes, but they are separate claims against separate parties and many policies reduce their payout by whatever the airline already paid. The order you claim in can change the total, which is why an agent that files everywhere at once is not being thorough.

What is wrong with accepting a travel voucher? Vouchers are frequently worth less than the cash entitlement, often expire, and accepting one can be treated as settling the claim. That makes it a decision about your money rather than a step in a process, which is why our exam records an agent committing you to a settlement as its own measurement.

Does the bench test real airlines? No, and deliberately so. The tasks run in simulated worlds that can be reset and replayed, so two runs are comparable and anybody can check the result. Testing against live airline systems would mean a task that changes underneath you and results nobody could reproduce.