Personal Agent Bench

Instinct vs Town, and two opposite answers to the same problem

Instinct vs Town is worth reading as a disagreement rather than a comparison. Both products face the same problem, which is that a language model cannot be handed a thousand tools and asked to choose well. They solve it in opposite directions. Instinct exposes its entire surface as a command line, 571 actions deep, and refuses to run any of them until the agent has read that action's own help text. Town catalogues 1105 tools and then hides five sixths of them, showing its main assistant 177. Everything else that differs between these two follows from that choice.

  • One makes the agent read the manual before acting. The other decides in advance which manuals the agent will ever see.
  • Their execution places are nearly as far apart: a leased cloud browser against a native Mac app that reaches your messages and your screen.
  • Town's agent has a personality you can open and edit. Instinct's memory is three stores and one of them is designed never to give anything back.
  • Both can be started by something other than you, which is where the difference in reach stops being abstract.

The two runtimes, side by side

AxisInstinctTown
Where it runsServer orchestration with work running in a sandbox. Browser tasks use a leased cloud browser, not the Chrome on your desk.Server-side sync and session handling. The Mac app is native and genuinely reaches into the system: messages, contacts, screen and audio.
How tools are exposedA command-line surface rather than a function list: 43 families and 571 actions by the product's own count, and three different totals depending on who is counting. An action refuses to run until its own help text has been read, which makes the manual part of the prompt.A catalogue of 1105, narrowed in four stages before anything is offered. The main assistant sees 177 of them; one background writer is allowed 17.
Where memory livesThree stores that do different jobs: a read-only markdown set, an observation log, and a fill-only vault for secrets.Server-side: a wiki, a memory store, and a people model. The personality is a document you can open and edit, not a line buried in the prompt.
What you get to approveAn approval card with three answers rather than two: approve, deny, or dismiss. It carries the plan and names which device will run it. Approval binds to the exact text shown, not to the request that produced it, so a rewrite needs a new approval. Its own rules forbid an unseen assumption from riding along with an irreversible send, and it will produce a receipt on request: what it read, what it sent, what it did not touch. Deletion has a documented surface covering storage, training, location and desktop indexing.Approval is configuration rather than a prompt. The run mode, the account scope, the tool list and the approval mode together decide whether a task may start at all and whether it may write to anything outside, and an individual tool can carry extra approval on top of that. Approval records are kept alongside the people the assistants deal with. Memory has its own add, update and delete tools rather than a setting.
How you reach itMessaging first. WhatsApp, SMS and iMessage are the default surface, plus an address on its own mail domain. The web app is mostly an admin panel.Web and Mac carry equal weight, and routines can be started by a clock, an email or a calendar entry rather than by you typing.
What we could not establishThe body of the main instruction set (only its section headings surfaced), which host the main agent runs on, the production model id, and four named policy sections we found pointers to but never text for.The full instruction set (the member-facing interface refuses it), where the personality document is spliced in, and the execution bodies of four newer routines.
Completed, whole exam43%36 to 50 · 159 tasksWithheld
The completion row is every task in the exam, not a head to head on one errand. Neither figure is a ranking and we never combine them. As of 2026-09-22.

Reading the manual against being handed a short list

Instinct does not ship a function list. It ships a command-line surface: 43 families and 571 actions by the product's own count, and three different totals depending on who is counting. An action refuses to run until its own help text has been read, which makes the documentation part of the prompt and therefore part of the bill.

Town ships a catalogue of 1105 tools narrowed in four stages before anything is offered. The main assistant sees 177. One background writer is allowed 17.

Both are answers to the fact that a model given too many options picks badly. Instinct's answer is to make the agent pay attention: everything is available, but you cannot touch it without first loading its description. Town's answer is to make the choice earlier, on the agent's behalf, and in a place the agent never sees.

The costs are symmetrical. Instinct spends tokens and latency on reading. Town spends the possibility that the right tool was one of the 928 that never got offered.

There is a second-order difference that matters more over time. Instinct's rule puts the documentation inside the loop, which means the product's own help text is now load-bearing: if a description is wrong or stale, the agent is wrong in a way that looks like a model failure. Town's narrowing happens in code the agent never reads, so a bad narrowing rule produces an agent that confidently reports it cannot do something it can. Both failure modes are hard to see from a transcript, and they are different enough that the same error would be diagnosed in two different places.

Why the tool counts cannot be compared

A comparison article would put 571 against 1105 and conclude something. Neither number means what it looks like.

Instinct's three totals are all true at once, depending on whether you count families, actions, or the subset a given session can reach. Town's 1105 is a catalogue rather than an offer, and the number that governs behaviour is 177, or 17 if the work falls to that background writer.

The honest comparison is not a count at all. It is that one product's limit is enforced at read time and the other's is enforced at selection time, and only one of those is visible to the agent doing the work.

This is the clearest case on the bench of a number that gets quoted constantly and carries almost no information. A tool count tells you how much the vendor has built. It does not tell you how much is reachable in a given conversation, which is the only version of the question that affects what happens to your errand.

A leased browser against a Mac that can see your screen

Instinct orchestrates on a server with the work running in a sandbox, and browser tasks go to a leased cloud browser rather than the Chrome on your desk. Nothing of consequence runs locally, and the browser belongs to a third company.

Town handles sync and sessions server-side but ships a native Mac application that genuinely reaches into the system: messages, contacts, screen and audio.

This is the same axis that separates Pine from Town, and the conclusion is the same: the exposure differs more than the capability does. What makes the pairing with Instinct interesting is that it inverts the tool story. The product that shows its agent everything keeps it in a box. The product that shows its agent a sixth of the catalogue gives it the run of your laptop.

Memory you can edit, and a vault that never gives anything back

Town keeps a wiki, a memory store and a people model on the server, and the agent's personality is a document you can open and edit rather than a line buried in a prompt.

Instinct keeps three stores doing three different jobs: a read-only markdown set, an observation log, and a fill-only vault for secrets. That last one is the detail worth sitting with. It is built so that credentials can go in and never come out, which is a deliberate and good design, and it also means there is a part of what the product holds about you that neither you nor the agent can inspect.

The difference in practice is correction. Town has a document where its idea of you lives, and you can change it. Instinct has an observation log, and a read-only set you do not write, and a vault that answers nothing.

Neither arrangement is careless. A vault that cannot be read back is the right design for credentials and a worse design for everything else, and the question is which things ended up inside it. An editable personality document is the right design for a disposition and a riskier one for anything a confused user might overwrite. What a comparison can say is that only one of these two products gives you a place to go when it has the wrong idea.

Where you reach them, and who else can start them

Instinct is messaging first. WhatsApp, SMS and iMessage are the default surface, plus an address on a mail domain it gives you. The web app is mostly an admin panel.

Town gives web and Mac equal weight, and routines can be started by a clock, an email or a calendar entry rather than by you typing.

Put those beside the execution paragraph and the asymmetry sharpens. Town can be woken by a calendar entry and then act through an application that reads your messages. Instinct can be reached from anywhere you already type, and what it reaches is a sandbox. Neither is the safer product in general. One concentrates risk in what it can touch, the other in how easily it can be addressed.

The same gap in both teardowns

Our page for each product carries a list of what we could not establish, and the overlap is almost exact.

For Instinct we could not get the body of the main instruction set, only its section headings, nor which host the main agent runs on, nor the production model id, nor four named policy sections we found pointers to but no text for. For Town we could not get the full instruction set, because the member-facing interface refuses it, nor where the personality document is spliced in, nor the execution bodies of four newer routines.

Both products let us see the shape of the instruction set and not its contents. That is not coincidence, it is the boundary of what external examination reaches, and it is worth stating in a comparison because it is the part where any writeup, including ours, stops knowing and starts inferring. It is also the reason the two lists above are the most useful paragraphs on either teardown page: they mark where our confidence ends.

Which one the choice comes down to

Two questions decide it, and neither is answered by a feature list.

Does the work depend on things the agent can only learn by looking at your computer? If yes, Town's foothold is the product and Instinct cannot substitute for it. If no, it is exposure carried for nothing.

Do you want to reach the agent from wherever you already are? Instinct is built on that premise, and the design follows from it: no new app to open, no window to check, everything surviving as a text message. Town expects you to come to it, on the web or on the Mac, and in exchange it is present on the machine in a way a message thread can never be.

What the figures do and do not settle

Our scorecard carries a completion rate for Instinct with its interval and task count. Town's is withheld, because the figure we were given does not reconcile with the only records we can check.

That withheld row is the honest state of this pairing. We can compare these two products on how they are built, which is what this page does, and we cannot yet compare them on how often they finish an errand. Treat anyone who does the second thing confidently as having skipped the question of where the number came from.

FAQ

What is the main difference between Instinct and Town? How each one decides what its agent may do. Instinct exposes 571 actions as a command line and requires the help text to be read before an action runs. Town catalogues 1105 tools and shows its main assistant 177 of them.

Which has access to more of my computer? Town. Its Mac application is native and reaches messages, contacts, screen and audio. Instinct runs work in a sandbox and sends browser tasks to a leased cloud browser.

Can I change what either one thinks about me? Town's personality is a document you can open and edit. Instinct has a read-only markdown set, an observation log, and a vault built to accept secrets and never return them.

How do I reach each one? Instinct lives in WhatsApp, SMS and iMessage, plus a mail address it provides; its web app is mostly an admin panel. Town treats web and Mac equally and can also be started by a clock, an email or a calendar entry.

Which one scores better? Instinct has a published completion rate on our scorecard with an interval and a task count. Town's figure is withheld because it does not reconcile with the records we can check, so the comparison cannot be made yet.

Is this based on testing or on their documentation? On our own teardowns of both, Instinct and Town, including the lists of what we could not establish for each.