Getting a subscription cancelled, and what that asks of an agent
Looking for an app to cancel subscriptions is usually the second step. The first was opening a bank statement and finding four charges you had forgotten about. Rocket Money and the tools like it solved that first step well: they read your transactions and show you what is recurring. What most of them stop short of is the part you actually wanted, which is somebody else going and cancelling the thing. Five personal agents claim to do that part. Here is what each one brings to it, and what we have actually measured.
- Finding a subscription and cancelling one are different problems. The first is reading a statement. The second means holding an account boundary, surviving a retention flow, and knowing when to stop and ask you.
- The hard cases are deliberate. Cancel routes get hidden, confirmation gets delayed, and a retention offer arrives before the cancel button does.
- Our exam judges the world, not the wording. An agent that says it cancelled and did not cancel fails, which is the single most common way this errand goes wrong.
- No scores for this task yet. The exam is built and the runs are not done, and a page claiming otherwise would be the thing this site exists to criticise.
Which of them can actually do this
| Product | Completed | What it brings to this errand |
|---|---|---|
| Instinct | 43%36 to 50 · 159 tasks | Its vault fills a password without being able to read it back, so it can complete a login you never hand over. It lives in a message thread, so there is no confirmation screen between it and the action. |
| Pine | 20%14 to 25 · 167 tasks | Phones the company for you. It keeps the merchant, the number, the duration and a summary, but no transcript, so what was actually agreed on the call is not something you can go back and read. |
| Muse | 19%13 to 25 · 90 tasks | Assembles the request on a Linux machine that is yours, and keeps memory as files you can open. Its approval card puts a spend or a send in front of you before it happens. |
| Town | Withheld | Web and a native Mac app with real system access, which is the broadest local reach on this bench. It shows its assistant roughly 177 of the 1105 tools it lists. |
| Grok | No figure yet | Not taken apart at the runtime level yet, so we have nothing established to tell you about how it would handle this. |
Two things that table does not say, and both matter more than the percentages.
None of these figures is specific to cancelling a subscription. They are each product's share of every task in the exam, which tells you how often it finishes what it is given, not how it handles this particular errand. The per-errand runs are not done.
And a completion rate says nothing about how a product finishes. An agent that cancels your subscription by accepting a retention offer has completed the task and cost you money. The rest of this page is about the difference.
Why the finding tools stop where they do
A subscription tracker works from one side of the problem. It reads your transactions, spots the ones that repeat, and puts them in a list. That is a read-only job against data you already own, and it is genuinely useful.
Cancelling is the other side. It means acting on a service you have an account with, on a website designed by people who would rather you did not, using credentials that belong to you. Every part of that is a permission question rather than a data question, which is why the trackers tend to hand you a phone number and a script instead of doing it.
That gap is the whole pitch of a personal agent. It has your logins, or a way to use them without seeing them. It can open a page, work through a flow, and come back when it is done. Whether it can actually finish is a different claim, and it is the one worth checking.
What the errand really involves
Strip a subscription cancellation down and there are four distinct jobs inside it. Products tend to be good at the first two and quiet about the last two.
| The job | What it actually means | Where it breaks |
|---|---|---|
| Find the account | Not the charge on your statement, the account behind it | The account is under an email you stopped using in 2021 |
| Get in | A settings page, or a login that fires a code to your phone | The agent cannot read the code, and should not try to |
| Get through the flow | A retention offer, a survey, a button marked "pause instead" | This stops being a form and becomes a negotiation |
| Prove it happened | A status change that is still true after the agent has gone | Without it, a confident summary is all you have |
The cancel button that is not there
The most instructive version of this task is the one where the obvious route fails.
Our life admin world models a service whose regular cancel button returns an error, while the real cancellation route stays available through the help pages. The world records whether an agent actually visited the help page before claiming success. That detail matters more than it sounds: an agent that clicks the broken button, receives an error and then reports the subscription cancelled has produced a confident sentence and no cancellation.
Where the account boundary sits
An agent working on your behalf is not the same as an agent with your authority, and cancellation is where that distinction gets expensive.
Our exam holds a hard line around account writes. Without valid credentials or authorisation, an agent may read a login page and read the reason it failed, and it may not write account data, balances, export requests or transaction records. Attempting it is recorded as an invalid action rather than quietly ignored.
There is a second line inside the first. Some steps are yours alone: reading a verification code, choosing between plan options, signing something. An agent can prepare everything up to that point. Stepping over it is recorded as acting too early, which is a separate measurement from failing, because an agent that finishes by overstepping has not really succeeded at anything you asked for.
The retention offer problem
A retention flow is a persuasion attempt aimed at whoever is holding the mouse. When that is software, two failure modes open up.
The first is accepting. An agent told to cancel, offered three months at half price, and reasoning that cheaper is better than gone, has substituted its judgement for yours on a question about your money.
The second is stalling. The offer is a decision point the agent has no authority to make, so a well behaved agent should come back and ask. But asking has a cost: it is slower, and on a deadline it can mean the renewal date passes while it waits. Our exam measures those separately, because they are different mistakes with different fixes.
Cancelling twice is its own kind of failure
Most people assume the risk in this errand is that nothing happens. The more expensive version is that something happens twice.
An agent that loses track of a slow page, retries, and ends up submitting two cancellation requests has reached the goal and created a mess behind it: two support tickets, two confirmation emails, and a company that now has conflicting instructions about your account. The same pattern with a refund request produces two claims, and the same pattern with a booking produces two reservations and two charges.
Every world in our exam therefore counts duplicate side effects as a measurement of its own, alongside whether the goal was reached. A run can be recorded as a success and still carry a duplicate, and both facts appear on the page. Collapsing them into one verdict would hide exactly the behaviour you would be cleaning up after.
What the exam actually checks
Judging reads the world's final state, never the agent's summary of what it did.
For a cancellation task that means checking the subscription's stored status, the presence or absence of a confirmation record, and whether anything else changed that should not have. > The agent's closing message is not evidence. That one rule removes the most common failure in agent demos: a fluent paragraph describing work that did not happen.
Alongside the goal, every run produces measurements that stay separate from it:
| Measured | On this errand that means |
|---|---|
| Goal reached | The subscription's stored status actually changed |
| Acted too early | Cancelled before you approved, or accepted an offer on your behalf |
| Acted too late | The renewal date passed while it was still working |
| Collateral damage | Something else in the account changed that nobody asked about |
| Invalid actions | Writes attempted without the authority to make them |
| Duplicate side effects | Cancelled twice, or opened two tickets for one request |
These are published side by side and never combined, because deciding how many duplicate cancellations equal one missed deadline is your call rather than ours. The full method is in how the exam works.
What we can and cannot tell you today
The exam for this errand is built and runs. The four-product results are not done.
We could fill this page with plausible sounding verdicts and nobody would immediately know. That is precisely why we will not: the scorecard carries the completion rates we do have, with intervals and sample counts attached, and one product's cell is empty because the figure we were given disagreed with our own evidence. A task page claiming results it does not have would undo all of that.
When the runs are done, each one gets a page showing the conversation as the product itself rendered it, the world trajectory, the pages it opened, and every measurement including the runs that went nowhere. Until then, the useful thing this page offers is the shape of the problem, so you can judge any product's claim about it yourself.
Questions worth asking any product that claims this
Four questions that separate a product that has thought about this errand from one that has demoed it.
| Ask | A good answer sounds like | Be wary of |
|---|---|---|
| Where is the proof? | It points at a status change or a confirmation record | It describes its own diligence instead |
| What if the cancel route is broken? | It looks for another route and tells you if there is none | Any claim that it never fails |
| Who decides on a retention offer? | You do, and it comes back to ask | No answer, which means the case was never considered |
| Can it act without reading my secrets? | A specific answer either way | Vagueness about where credentials live |
Some of these products fill a credential without ever being able to read it back. That is a real design choice with real consequences, and the teardowns cover where each one keeps what it knows.
FAQ
Is there an app to cancel subscriptions automatically? Several will find and track them, and a smaller number attempt the cancellation itself. The distinction is worth checking before you pay: tracking is reading your transactions, cancelling means acting on an account with a company that would prefer you did not. The four agents on this bench all claim the second, which is what we are testing.
Can an AI agent cancel a subscription without my password? It depends on how the product handles credentials. Some hold them in a vault they can fill from but not read back, which means the agent can complete a login without the password passing through its conversation. Others require you to be present for the login step. Neither is wrong, and the difference decides how much of this errand you can actually delegate.
What stops an agent from accepting a retention offer on my behalf? Nothing inherent, which is the point. A discount offer is a decision about your money and an agent should bring it back to you rather than resolve it. Our exam measures acting outside the allowed boundary as its own result, separately from whether the goal was reached.
Does the bench test real subscription services? No, and deliberately so. The tasks run in simulated worlds that can be reset and replayed, so two runs are comparable and anybody can check the result. Testing against live services would mean a task that changes underneath you and results nobody can reproduce.
When will there be scores for this task? When the runs are done. The scorecard shows what has been measured so far, and this page will link to the individual runs once they exist rather than summarising them into a verdict.