Personal Agent Bench

Muse vs Pine, and what a machine of your own is actually for

Muse vs Pine is the cleanest test of a question the whole category keeps dodging: does a personal agent need a computer of its own. Muse answers yes twice over, giving every user a Linux virtual machine and spawning a second one when browser work is needed. Pine answers no, running nothing of substance on your machine and leasing a remote desktop when it has to see a web page. Both products then go and do errands for you, so the arrangement is not decorative. It changes what each one can hold, what it can resume, and what it costs to run.

  • One of them has a persistent place to put things. The other has to carry everything in the request.
  • That shows up most clearly in memory: plain files in a home directory against a fact layer you cannot open.
  • Pine makes phone calls, which is the one capability here that leaves software entirely.
  • Neither will show us the instruction text, and in both cases that is the part that decides behaviour.

The two runtimes, side by side

AxisMusePine
Where it runsOne Linux VM per user. The request is assembled on that machine rather than fetched as a finished blob. Browser work spawns a second, separate machine.Server-side. Neither the web app nor the Mac app runs the model. Browser work is handed to a separate agent on your own machine; the cloud route leases you a remote desktop you watch over VNC.
How tools are exposedTwo layers kept deliberately apart: a permission sheet of 246 methods carrying no parameters at all, and the real definitions loaded a namespace at a time when needed.26 tool names reach the client, grouped by job. The schemas never do. Computer use bypasses that list entirely.
Where memory livesPlain markdown in a home directory, embedded straight into the request. Changing one of those files inserts a short note rather than rebuilding the whole package.No file you can edit. A server-side persona plus a structured fact layer queried on demand.
What you get to approvePermissions and approvals are not scattered through the runtime. A single gatekeeper is the only path to them, along with anything leaving the machine, and the interface carries an approvals tab beside the activity feed. Credentials are held outside the agent's reach by the same arrangement. Deletion is unusually concrete for this list: workspace files can be removed like files, and a forget capability sits among its skills.Confirmation is required before a billing action, and there are two separate kinds of it. A start confirmation gates the task and carries its own expiry. A different identity confirmation appears when the company on the phone wants to hear from the account holder: a card naming what to say, which number will be calling, and when the offer lapses. Both the request and the receipt land in the task history. There is nothing to delete, because there is no file holding what it knows, only a persona and a fact layer.
How you reach itLaunched on iOS, Android and the web, and reachable in WhatsApp. A Mac app since 17 September can act in apps on your computer. We tested the web surface only.Web and Mac, outbound phone calls, a voice copilot, and a REST endpoint other people's agents can call. Four request shapes under one brand.
What we could not establishThe raw traffic between the runtime and the model: structure known, contents never seen. Also credentials, live third-party sign-in, and an interactive terminal the owners chose not to open.Production instruction text (roughly 126 probes, all empty), the calling tool's schema, which carrier places the calls, and whether work is split across a fast and a slow model.
Completed, whole exam19%13 to 25 · 90 tasks20%14 to 25 · 167 tasks
The completion row is every task in the exam, not a head to head on one errand. Neither figure is a ranking and we never combine them. As of 2026-09-22.

Two machines, or none

Muse gives each user a Linux VM. The request is assembled on that machine rather than arriving as a finished blob, and browser work spawns a second, separate machine. One user, two computers, both of them Meta's.

Pine runs server-side and neither its web app nor its Mac app runs the model. Browser work is handed to a separate agent on your own computer, while the cloud route leases a remote desktop you watch over VNC.

The phrase to notice in the Muse description is "assembled on that machine". It means the prompt is built where the files are, rather than composed elsewhere and shipped in. That is only possible because there is a persistent place for the files to live, and it is the single structural advantage a machine of your own provides.

What persistence buys, and what it costs

Having somewhere to put things changes what memory can be.

Muse keeps plain markdown in a home directory, embedded straight into the request, and changing one of those files inserts a short note rather than rebuilding the whole package. That last detail is the one that gives the design away: the memory is files, the files are on a machine, and editing one has the cost of editing a file rather than the cost of regenerating a context.

Pine has no file you can edit. There is a server-side persona and a structured fact layer queried on demand. Facts come back when something asks for them, which works and is cheaper to operate, and it means there is no object with your preferences in it, only a service that answers questions about you.

Ask what happens when each one is wrong about you. On one, something is wrong in a file. On the other, something is wrong in a query result, and you are asking a product to stop believing something rather than editing a line.

The cost of the Muse arrangement is that a machine per user is a machine per user. That is a real operating bill, and it is the reason most products on this bench do not do it. It also means the product carries state it has to keep safe, back up and eventually delete, none of which a stateless fact service has to think about.

Permission and definition, kept apart

Muse separates two things most products ship together. There is a permission sheet of 246 methods carrying no parameters at all, and then the real definitions, loaded a namespace at a time when they are needed.

Pine lets 26 tool names reach the client, grouped by job. The schemas never do. Computer use bypasses that list entirely.

Both are withholding, and it is worth being precise about what each one withholds and from whom. Muse withholds definitions from the context until the moment of use, which is an efficiency decision, and the permission sheet still says what exists. Pine withholds schemas from the client, which is a decision about what an outside observer can learn, and the published names do not describe what the product can actually do because computer use goes around them.

One of them phones people

Pine phones companies on your behalf about bills, refunds and cancellations. Muse does not, and almost nothing on this bench does.

This is the widest functional gap between the two, and it is wider than it appears, because a call is not another tool call. It puts an agent into a conversation with a person who did not agree to speak to a machine, about an account that belongs to you, on a line somebody is recording. When it goes wrong there is no log to inspect, only a human on the other end with an impression of what you authorised.

It also explains Pine's shape. Four request forms sit under that one brand: web, Mac, outbound calls, a voice copilot, and a REST endpoint other software can call. A product built around phone calls needs to be startable by something other than a person typing, and the endpoint is where that thought ends up.

Where you find each one

Muse launched on iOS, Android and the web, is reachable in WhatsApp, and since 17 September has had a Mac app that can act in applications on your computer. We tested the web surface only, which is stated here because it bounds everything else we say about it.

Pine answers on web and Mac, by outbound phone call, through a voice copilot, and over that REST endpoint.

Muse's Mac app matters for this comparison and we have not examined it. The product we took apart was the web one, with its Linux machine behind it. If the Mac app reaches into applications the way the description says, then Muse now spans both ends of the axis this page is built on, and that is a reason to re-examine it rather than a finding.

What neither will show us

For Muse the gap is the raw traffic between the runtime and the model: we know the structure and have never seen the contents. Also credentials, live third-party sign-in, and an interactive terminal the owners chose not to open.

For Pine the gap is the production instruction text, after roughly 126 probes that all came back empty, plus the calling tool's schema, which carrier places the calls, and whether work is split between a fast and a slow model.

Two of those are worth flagging for anyone comparing these products seriously. Nobody has seen what Muse actually sends the model. And nobody knows who is on the other end of Pine's phone calls, which for a product whose signature feature is a phone call is a larger blank than it sounds.

The interactive terminal in the Muse list is its own small story. One exists on that machine, and the owners chose not to open it. That is a decision rather than an absence, and it tells you the VM is a real computer being deliberately kept from behaving like one.

Which one the choice comes down to

The two questions that decide it are not the ones a feature list asks.

The first is whether your errands end in a conversation with a company. Cancellations, refunds and billing disputes usually do, and Pine is built around exactly that, down to the shape of its four front doors. Muse has no answer here at all.

The second is whether you want the agent to accumulate. A machine with files on it gets better at knowing you in a way a fact service does not, because the files are there to be added to and the agent assembles its own request on top of them. If the work you have in mind is a long relationship rather than a sequence of discrete jobs, that is the difference. If it is a sequence of discrete jobs, the persistence is overhead you are paying for.

Those two questions pull in opposite directions, which is why these products are less substitutable than their marketing suggests. One is a service that goes and gets a thing done, including by telephone. The other is a resident that learns.

What the figures say here

Both products have a published completion rate on our scorecard, each with an interval and the number of tasks behind it, and the two intervals do not overlap.

That is a real difference and it is still not a ranking. It says one of them finished more of the errands in the set we scored. It does not say which of these two arrangements is better, because the set was not designed to reward having a machine of your own, and a rate cannot see the difference between an errand completed inside a leased desktop and the same errand completed on a persistent VM. Read the rate for what it measures, and where it runs for the part it cannot.

FAQ

What is the main difference between Muse and Pine? Muse gives every user a Linux virtual machine, and a second one for browser work. Pine runs nothing of substance on your machine and leases a remote desktop when it needs a browser.

Which one can I correct when it gets something wrong about me? Muse keeps plain markdown files in a home directory, and changing one inserts a short note rather than rebuilding everything. Pine has no file to edit: a server-side persona and a fact layer queried on demand.

Can either one make phone calls? Pine does, about bills, refunds and cancellations. Muse does not. We could not establish which carrier places Pine's calls.

Does Muse run on my computer? The product we examined does not: it runs on a Linux machine Meta provides. A Mac app released on 17 September is described as acting in applications on your computer, and we have not examined it.

How many tools does each have? Muse exposes a permission sheet of 246 methods with no parameters, loading real definitions only when needed. Pine sends 26 tool names and no schemas, and its computer use bypasses that list.

Is this from testing or from their documentation? From our own teardowns of Muse and Pine, including what we could not establish for each. For Muse we examined the web surface only.

Why do the two completion rates not settle it? Because the task set was not built to reward having a machine of your own, and a completion rate cannot tell an errand finished on a persistent VM from the same errand finished in a leased desktop. The rates are real and they are narrow in what they measure.