Personal Agent Bench

Disputing a charge, and the decision an agent should not make for you

How to dispute a charge looks like one question and is actually two. There is the merchant route, where you ask the company that took your money to give it back. And there is the bank route, a chargeback, where you tell your card issuer the charge was wrong and let it claw the money out of the merchant's account. Both end with money returned. They are not interchangeable, they are not equally reversible, and picking the wrong one can cost you the account you were trying to fix. The personal agents on this bench claim to handle errands like this. Here is what the errand actually involves, and what we measure when we make them try.

  • A refund and a chargeback are different instruments. One is a request. The other is a claim against the merchant, and some companies close accounts that receive them.
  • The hard part is not finding the transaction. It is deciding which route to take, and that decision is about your money and your relationship with the merchant.
  • Doing both at once is the expensive failure. A merchant refund plus a bank chargeback is a double recovery, and sorting it out lands on you.
  • Our exam judges the world, not the wording. The shopping worlds count crossing the payment line as its own measurement, separate from whether the goal was reached.

Which of them can actually do this

ProductCompletedWhat it brings to this errand
Instinct43%36 to 50 · 159 tasksIts vault fills a password without being able to read it back, so it can reach an account portal you never hand over. It lives in a message thread, so there is no confirmation screen standing between it and a filed dispute.
Pine20%14 to 25 · 167 tasksPhones the company for you, which is the merchant route rather than the bank one. It keeps the merchant, the number, the duration and a summary, but no transcript, so what was actually agreed about the money is not something you can go back and read.
Muse19%13 to 25 · 90 tasksAssembles the case on a Linux machine that is yours, and keeps memory as files you can open, which means the evidence it gathered is readable afterwards. Its approval card puts a spend or a send in front of you before it happens.
TownWithheldWeb and a native Mac app with real system access, the broadest local reach on this bench, so it can reach receipts and statements sitting on your own machine. It shows its assistant roughly 177 of the 1105 tools it lists.
GrokNo figure yetNot taken apart at the runtime level yet, so we have nothing established to tell you about how it would handle this.
The completion column is every task in the exam, not this one. We do not have per-errand figures yet, and a single rate presented as one would be the kind of thing this site exists to catch. As of 2026-09-22.

Two things that table does not say, and both matter more than the percentages.

None of those figures is specific to disputing a charge. Each is a product's share of every task in the exam, which tells you how often it finishes what it is given, not how it handles this particular errand. The per errand runs are not done.

And a completion rate says nothing about how a product finishes. An agent that recovers your money by filing a chargeback you did not authorise has completed the task and may have cost you the account. The rest of this page is about that gap.

The two routes, and why the choice is the whole errand

Ask the merchant and you are making a request. The company checks its records, decides, and either returns the money or does not. It is slow, it is refusable, and it leaves the relationship intact. If you want to keep shopping there, this is the route.

File with your bank and you are making a claim. The issuer takes the money back from the merchant, charges that merchant a fee on top, and asks it to prove the charge was good. It is faster and far more likely to succeed. It also tells the merchant that you went over its head, and a fair number of companies respond by closing the account and refusing future business.

The instruments are not ranked. Which one is right depends on how much you want the money versus how much you want to stay a customer, and nobody but you can weigh that.

That is the shape of the problem an agent walks into. It is not a lookup and it is not a form. It is a judgement call with a consequence attached, sitting behind two steps of ordinary admin.

What the errand really involves

Strip a dispute down and there are four jobs inside it. Products tend to be confident about the first two and quiet about the last two.

The jobWhat it actually meansWhere it breaks
Identify the chargeThe line on the statement is a descriptor, not a merchantThe descriptor is a payment processor in another city
Assemble the evidenceOrder number, date, amount, what you were promisedThe confirmation email is in an account nobody linked
Choose the routeMerchant refund or bank chargebackThis is a decision about your money, not a step
Follow it throughA response window, a request for more detail, a deadlineThe clock started when the charge posted, not today

The first two are engineering problems. The third is a permission problem, and permission problems are where delegating your admin either works or quietly does not.

The clock nobody mentions

Card network rules give you a window to dispute, commonly measured in months from the transaction or from the date you expected delivery. Miss it and the strongest route closes, leaving you with the merchant's goodwill and nothing else.

This matters for an agent because the deadline is not visible on the page it is looking at. It is a property of the charge, and it runs whether or not anybody is working on the problem.

Our exam has a matching shape. Several tasks are built around an event that starts a clock, and the agent has a deadline measured from that moment. Acting after the deadline is recorded as its own failure even when the final state is right, because an outcome that arrives too late is not the outcome that was asked for. Acting before the triggering event is recorded separately too, since an agent that files a dispute before the charge has actually settled has guessed rather than worked.

Where the payment line sits

Shopping is the one part of our exam with an extra measurement bolted on, and it exists for exactly this errand.

Every world counts the universal five: whether the goal was reached, whether the agent acted too early, whether it acted too late, whether it damaged something the task was watching, and how many invalid operations it attempted. The shopping worlds add one more, which is whether the agent crossed the payment line.

Crossing that line covers anything that moves money or commits you to moving it. Filing a chargeback is on that side. So is accepting a partial refund that closes the case, and so is agreeing to store credit in place of cash. An agent can prepare all of it and bring it to you. Stepping over on its own is recorded as acting outside its boundary, which stays separate from whether the money came back, because an agent that finishes by overstepping has not really done what you asked.

There is a hard line underneath that one. Without valid credentials or authorisation, an agent may read a login page and read the reason a login failed, and it may not write account data, balances, export requests or transaction records. Attempting it is recorded as an invalid action rather than quietly ignored.

Recovering the money twice

Most people assume the risk here is getting nothing back. The more expensive version is getting it back twice.

An agent that emails the merchant for a refund, does not see a reply quickly enough, and then files a chargeback has started two independent processes against the same charge. If both land, you have been paid twice for one purchase. The merchant sees a customer who took a refund and then disputed it anyway, the bank sees a claim that the merchant can now disprove, and the cleanup is yours.

The same pattern without the second route is milder and still real: two support tickets for one charge, two case numbers, and a company holding conflicting instructions about your account.

Every world in our exam counts duplicate side effects as a measurement of its own, alongside whether the goal was reached. A run can be recorded as a success and still carry a duplicate, and both facts appear. Collapsing them into one verdict would hide precisely the behaviour you would be cleaning up after.

It is the failure most likely to be invisible in a demo, because the agent's summary describes the outcome it intended rather than the number of routes it opened to get there.

When the page lies to you

The instructive version of this task is the one where the obvious route fails.

A task in our exam can declare faults, and a shared controller fires them at a set logical time so that the web page and the API both see the same breakage at the same moment. A page can start returning errors for a while. Content can come back truncated, with the truncation marked, so an agent that fills in the missing part from general knowledge is making it up rather than reading it. Prices and order statuses can change partway through, which turns a correct answer into a stale one if the agent read early and never looked again.

That last one is built for disputes. An agent that reads an order as unfulfilled, spends several minutes assembling a case around non delivery, and never looks again has written a confident and false claim. The world knows the status changed. The agent's summary does not.

What the exam actually checks

Judging reads the world's final state, never the agent's account of what it did.

For a dispute that means checking whether a case record exists, which instrument it was filed under, whether the balance moved, and whether anything else changed that should not have. The agent's closing message is not evidence. That single rule removes the most common failure in agent demos, which is a fluent paragraph describing work that did not happen.

MeasuredOn this errand that means
Goal reachedThe charge is actually disputed or reversed in the record
Acted too earlyFiled before the charge settled, or before you approved
Acted too lateThe dispute window closed while it was still working
Collateral damageSomething else in the account changed that nobody asked about
Invalid actionsWrites attempted without the authority to make them
Payment boundaryIt committed you to money or a settlement on its own
Duplicate side effectsTwo routes opened, or two cases for one charge

These are published side by side and never combined, because deciding how many unauthorised chargebacks equal one missed deadline is your call rather than ours. The full method is in how the exam works, and the reasoning behind refusing a combined score is in how scoring works.

Questions worth asking any product that claims this

AskA good answer sounds likeBe wary of
Which route will you take?It names merchant or bank and asks you to confirmIt says it will handle the dispute
Where is the proof?It points at a case record or a balance changeIt describes its own diligence instead
What if the merchant goes quiet?It reports back and waits for youIt escalates on its own initiative
Can it act without reading my secrets?A specific answer either wayVagueness about where credentials live

Some of these products fill a credential without ever being able to read it back. That is a real design choice with real consequences, and the teardowns cover where each one keeps what it knows.

What we can and cannot tell you today

The shopping worlds are built and they run. The four product results for this specific errand are not done.

We could fill this page with plausible verdicts and nobody would immediately know. That is exactly why we will not. The scorecard carries the completion rates we do have, with intervals and sample counts attached, and one product's cell is empty because the figure we were given disagreed with our own evidence. A task page claiming results it does not have would undo all of that.

One thing has changed since this page was drafted. The shopping world now has a recorded attempt you can read: Instinct, six steps, just under twelve minutes, finishing with the world in the wrong state. Every measurement on it is shown separately, including the ones that came back zero. Read it as a record and not as a result. One attempt by one product on one task is not a rate, and no batch level quality gate was written for it, so the run is not admitted to the comparison at all. The run page says so itself, at the top, rather than leaving you to find out.

When the runs are done, each one gets a page showing the conversation as the product itself rendered it, the world trajectory, the pages it opened, and every measurement including the runs that went nowhere. Until then, the useful thing this page offers is the shape of the problem, so you can judge any product's claim about it yourself. The nearest errand we have written up in the same detail is cancelling a subscription, which shares the retention flow problem and none of the payment risk.

FAQ

What is the difference between a refund and a chargeback? A refund is a request to the merchant, which it can refuse and which leaves your relationship intact. A chargeback is a claim filed with your card issuer, which pulls the money back from the merchant and charges it a fee. Chargebacks succeed more often and are far more likely to end with the merchant closing your account.

Can an AI agent dispute a charge on my behalf? Several claim to handle errands of this shape. The part worth checking before you delegate is whether the product will choose between the merchant route and the bank route without asking you. That choice has a consequence you carry, so an agent should bring it back rather than resolve it, and our exam measures acting outside that boundary as its own result.

How long do I have to dispute a charge? Card network rules give a window measured in months, usually counted from the transaction or from the date you expected delivery. The exact limit depends on your issuer and the reason for the dispute, so check the terms on your own card rather than a general figure. The practical point for delegation is that the clock runs whether or not anybody is working on it.

What happens if an agent files a dispute twice? You can end up recovering the same money through two routes, which is a problem rather than a bonus, and untangling it falls to you. Our exam counts duplicate side effects as a separate measurement from whether the goal was reached, precisely because a run can look successful and still leave that mess behind.

Does the bench test real banks and merchants? No, and deliberately so. The tasks run in simulated worlds that can be reset and replayed, so two runs are comparable and anybody can check the result. Testing against live financial services would mean a task that changes underneath you, results nobody can reproduce, and real money at risk.