Personal Agent Bench

Pine vs Instinct, two products that look identical from one step back

Pine vs Instinct is the pairing where the headline axis fails. Ask where each one runs and you get the same answer twice: server-side, nothing of consequence on your computer, the model somewhere else. On the property this bench usually leans on hardest, these two are the same product. The differences are a layer down, and they are large. One of them hands browser work to your own machine and phones companies on the telephone. The other leases a browser from a third party and waits for you in WhatsApp.

  • They are the only pair here that agree on where the work happens and disagree about almost everything after that.
  • Pine touches your computer more than Instinct does, which is the opposite of what their descriptions suggest.
  • One reaches outward to companies, the other inward to your message threads. Both call that meeting you where you are.
  • Their tool surfaces are the two extremes on this bench: 26 names with no schemas, against 571 actions you must read about first.

The two runtimes, side by side

AxisPineInstinct
Where it runsServer-side. Neither the web app nor the Mac app runs the model. Browser work is handed to a separate agent on your own machine; the cloud route leases you a remote desktop you watch over VNC.Server orchestration with work running in a sandbox. Browser tasks use a leased cloud browser, not the Chrome on your desk.
How tools are exposed26 tool names reach the client, grouped by job. The schemas never do. Computer use bypasses that list entirely.A command-line surface rather than a function list: 43 families and 571 actions by the product's own count, and three different totals depending on who is counting. An action refuses to run until its own help text has been read, which makes the manual part of the prompt.
Where memory livesNo file you can edit. A server-side persona plus a structured fact layer queried on demand.Three stores that do different jobs: a read-only markdown set, an observation log, and a fill-only vault for secrets.
What you get to approveConfirmation is required before a billing action, and there are two separate kinds of it. A start confirmation gates the task and carries its own expiry. A different identity confirmation appears when the company on the phone wants to hear from the account holder: a card naming what to say, which number will be calling, and when the offer lapses. Both the request and the receipt land in the task history. There is nothing to delete, because there is no file holding what it knows, only a persona and a fact layer.An approval card with three answers rather than two: approve, deny, or dismiss. It carries the plan and names which device will run it. Approval binds to the exact text shown, not to the request that produced it, so a rewrite needs a new approval. Its own rules forbid an unseen assumption from riding along with an irreversible send, and it will produce a receipt on request: what it read, what it sent, what it did not touch. Deletion has a documented surface covering storage, training, location and desktop indexing.
How you reach itWeb and Mac, outbound phone calls, a voice copilot, and a REST endpoint other people's agents can call. Four request shapes under one brand.Messaging first. WhatsApp, SMS and iMessage are the default surface, plus an address on its own mail domain. The web app is mostly an admin panel.
What we could not establishProduction instruction text (roughly 126 probes, all empty), the calling tool's schema, which carrier places the calls, and whether work is split across a fast and a slow model.The body of the main instruction set (only its section headings surfaced), which host the main agent runs on, the production model id, and four named policy sections we found pointers to but never text for.
Completed, whole exam20%14 to 25 · 167 tasks43%36 to 50 · 159 tasks
The completion row is every task in the exam, not a head to head on one errand. Neither figure is a ranking and we never combine them. As of 2026-09-22.

The same answer, and then a split

Pine runs server-side. Neither the web app nor the Mac app runs the model. Browser work is handed to a separate agent on your own machine, and the cloud route leases a remote desktop you watch over VNC.

Instinct orchestrates on a server with the work running in a sandbox. Browser tasks use a leased cloud browser, not the Chrome on your desk.

The first sentence of each is the same. The second is not, and it inverts what you would guess from how the two products present themselves. Instinct is the one that lives in your messaging apps, so it sounds closer to you. Pine is the one with the formal web interface, so it sounds further away. In practice Pine is the product that puts an agent on your computer to do browser work, and Instinct is the one that never does.

Being easy to reach and being close to your machine are different properties, and the two products have them in opposite combinations.

Twenty-six names against five hundred and seventy-one actions

Pine lets 26 tool names reach the client, grouped by job. The schemas never do, and computer use bypasses the list entirely, so the published surface does not describe what the product can actually do.

Instinct ships a command line instead of a function list: 43 families and 571 actions by its own count, with three different totals depending on who is counting. An action refuses to run until its own help text has been read, which makes the documentation part of the prompt and part of the bill.

These are the two extremes on this bench, and they are extremes of different things. Pine is minimal about what it tells the client and unbounded about what it can do through computer use. Instinct is maximal about what it exposes and strict about the conditions. One is a small door with a large room behind it. The other is a large door with a reading requirement at the threshold.

The practical consequence is about predictability. With Instinct you can enumerate the action surface and know the agent has loaded a description before acting. With Pine you cannot enumerate anything, because the interesting capability is the one that goes around the list.

It also changes who absorbs a mistake in the catalogue. Instinct's help text is inside the loop, so a stale description produces a wrong action that looks like a model error. Pine has no published description to be stale, and a wrong action through computer use looks like the agent simply doing something nobody expected. The first is diagnosable from a transcript. The second often is not.

Outward and inward

Pine phones companies about bills, refunds and cancellations. Four request shapes sit under that one brand: web, Mac, outbound calls, a voice copilot, and a REST endpoint other software can call.

Instinct is messaging first. WhatsApp, SMS and iMessage are the default surface, plus an address on a mail domain it provides. The web app is mostly an admin panel, and on the account we examined its chat route was switched off.

Both describe this as meeting you where you already are, and both are telling the truth about different directions. Instinct's version is that you never open anything: the agent is in the thread you were already in. Pine's version is that the errand reaches the company without you being in the conversation at all.

That second one is the more unusual capability and the less examined. A call puts an agent in front of a person who did not agree to talk to a machine, about your account, on a recorded line. When it goes wrong, the artefact is somebody's memory of what was agreed.

It is also the only capability on this bench that cannot be rolled back by software. A wrong message can be deleted and a wrong booking can be cancelled. A conversation that happened, happened.

Three stores against a fact service

Instinct keeps three stores doing three jobs: a read-only markdown set, an observation log, and a fill-only vault for secrets built to accept credentials and never return them.

Pine has no file you can edit. A server-side persona, and a structured fact layer queried on demand.

Neither gives you a document to correct, which separates both of them from Muse and Town. The difference between these two is what kind of thing is holding the state. Instinct has objects: a set, a log, a vault, each with its own rules, and you can at least reason about which is which. Pine has a service that answers questions about you, and reasoning about it means reasoning about query results.

If the product has the wrong idea about you, Instinct gives you a place to point at and Pine gives you a behaviour to complain about.

What wakes each one

Instinct can be started by a clock or an event. You are not the only thing that wakes it up.

Pine's REST endpoint means other people's software can drive it, which is the same property arriving from a different direction: the trigger is not necessarily you, and in Pine's case it is not necessarily a trigger you can see.

Put that next to the browser arrangement and Pine's shape comes into focus. Something that is not you can start it, it can place a telephone call, and when it needs a browser it may use an agent on your own machine. Each of those is defensible alone. Together they describe a product with more reach into the world than its quiet web interface implies.

The same wall in both teardowns

For Pine we could not get the production instruction text, after roughly 126 probes that all came back empty, nor the calling tool's schema, nor which carrier places the calls, nor whether the work is split between a fast and a slow model.

For Instinct we could not get the body of the main instruction set, only its section headings, nor which host the main agent runs on, nor the production model id, nor four named policy sections we found pointers to but no text for.

Both lists stop in the same place. We can describe the machinery around the instructions in detail and we cannot read the instructions. That is the boundary of examination from outside, and the honest thing a comparison can do is mark it rather than write past it.

The Pine gap worth repeating is the carrier. The product's distinguishing feature is a phone call, and we do not know whose network carries it. That is not an exotic detail. Whoever carries the call is a party to every conversation the product has on your behalf, and the product does not name them.

Which one the choice comes down to

One question separates them, and it is not about capability.

Does your errand end with a company, or with you? Cancellations, refunds and billing disputes end with a company, often on the phone, and Pine is built end to end for that. Scheduling, triage, following something up and keeping track end with you, and Instinct is built to live in the place where that conversation already happens.

The second question is tolerance for an unenumerable surface. Instinct will tell you what it can do and make the agent read about it first. Pine will not, and its most capable path is the one that does not appear in any list. If you want to reason about what the thing might do, that difference matters more than any number on our scorecard.

What the figures say

Both carry a published completion rate with an interval and a task count, and the gap between them is the widest on the board among products with usable figures.

It is still not a ranking, and this pairing is where that caveat bites hardest. The scored set mixes errands that end with a company and errands that end with you, and these two products are built for different halves of it. A single rate across the whole set is a statement about the mix as much as about the products.

FAQ

What is the main difference between Pine and Instinct? What each one reaches for. Pine phones companies about bills, refunds and cancellations. Instinct lives in your messaging apps and works from the thread you are already in.

Do either of them run on my computer? Neither runs the model locally. Pine hands browser work to a separate agent on your own machine and leases a remote desktop on its cloud route. Instinct uses a leased cloud browser and never touches yours.

Which exposes more tools? Instinct, by a wide margin on paper: 571 actions across 43 families, with the help text required before an action runs. Pine sends 26 names and no schemas, and its computer use goes around that list, so its real surface is not countable.

Can I correct what either one knows about me? Not directly. Instinct keeps a read-only markdown set, an observation log and a fill-only vault. Pine keeps a server-side persona and a fact layer queried on demand. Neither gives you a document to edit.

Can something other than me start them? Both. Instinct can be woken by a clock or an event. Pine exposes a REST endpoint other software can call, as well as placing outbound calls.

Is this from testing or from their documentation? From our own teardowns of Pine and Instinct, including the list of what we could not establish for each.