Personal Agent Bench

Muse vs Town, and two things people call a machine

Muse vs Town looks at first like the one pairing on this bench where both products agree. Both of them put the agent on a computer rather than leaving it to call APIs from a server. Read the two teardowns next to each other and the agreement disappears, because they mean opposite things by it. Muse's machine is a Linux VM that Meta owns, created for you and sitting in their cloud. Town's machine is the Mac already on your desk. Everything else, what each can reach, what each remembers, what happens when it is wrong, follows from whose computer it is.

  • Both answer "a machine", and one of them means theirs while the other means yours.
  • Both keep memory you could in principle correct, which is rarer than it should be.
  • Their tool surfaces are built on the same insight and implement it in different places.
  • One of the two has no published figure on our scorecard, and that is deliberate.

The two runtimes, side by side

AxisMuseTown
Where it runsOne Linux VM per user. The request is assembled on that machine rather than fetched as a finished blob. Browser work spawns a second, separate machine.Server-side sync and session handling. The Mac app is native and genuinely reaches into the system: messages, contacts, screen and audio.
How tools are exposedTwo layers kept deliberately apart: a permission sheet of 246 methods carrying no parameters at all, and the real definitions loaded a namespace at a time when needed.A catalogue of 1105, narrowed in four stages before anything is offered. The main assistant sees 177 of them; one background writer is allowed 17.
Where memory livesPlain markdown in a home directory, embedded straight into the request. Changing one of those files inserts a short note rather than rebuilding the whole package.Server-side: a wiki, a memory store, and a people model. The personality is a document you can open and edit, not a line buried in the prompt.
What you get to approvePermissions and approvals are not scattered through the runtime. A single gatekeeper is the only path to them, along with anything leaving the machine, and the interface carries an approvals tab beside the activity feed. Credentials are held outside the agent's reach by the same arrangement. Deletion is unusually concrete for this list: workspace files can be removed like files, and a forget capability sits among its skills.Approval is configuration rather than a prompt. The run mode, the account scope, the tool list and the approval mode together decide whether a task may start at all and whether it may write to anything outside, and an individual tool can carry extra approval on top of that. Approval records are kept alongside the people the assistants deal with. Memory has its own add, update and delete tools rather than a setting.
How you reach itLaunched on iOS, Android and the web, and reachable in WhatsApp. A Mac app since 17 September can act in apps on your computer. We tested the web surface only.Web and Mac carry equal weight, and routines can be started by a clock, an email or a calendar entry rather than by you typing.
What we could not establishThe raw traffic between the runtime and the model: structure known, contents never seen. Also credentials, live third-party sign-in, and an interactive terminal the owners chose not to open.The full instruction set (the member-facing interface refuses it), where the personality document is spliced in, and the execution bodies of four newer routines.
Completed, whole exam19%13 to 25 · 90 tasksWithheld
The completion row is every task in the exam, not a head to head on one errand. Neither figure is a ranking and we never combine them. As of 2026-09-22.

Whose computer it is

Muse gives each user a Linux virtual machine. The request is assembled on that machine rather than arriving as a finished blob, and browser work spawns a second, separate machine. Both are Meta's.

Town handles sync and session work on a server, and ships a native Mac application that reaches into the system: messages, contacts, screen and audio.

The difference is not how much computer there is. It is who is exposed when something goes wrong. A mistake on Meta's VM happens inside an environment Meta built and can inspect, and whatever it damages is whatever was on that VM. A mistake through Town's Mac app happens to your actual messages and your actual contacts, and the only record is whatever your own machine kept.

Both arrangements have a defensible logic. Only one of them puts the consequences on hardware you own.

Context is the thing both are buying

It is worth being clear about why either product bothers, because running on a machine is expensive and neither would do it for fun.

An agent that can only call APIs sees what the API returns. An agent on a machine sees the environment: files that are already there, a browser session already signed in, an application already open. Town's version of that is your screen and your messages, which is as much context as anyone has figured out how to give an agent. Muse's version is a home directory it has been writing to, which is less context about your life and more context about its own previous work.

That is the real split. Town reaches for context that already exists in your world. Muse accumulates context inside a world it controls. Both are called memory in the marketing and they are not the same substance.

The consequence shows up on day one against month six. Town is useful immediately, because your screen is already full of context it can read. Muse starts emptier and should get better, because every run leaves something on the disk. A review written in the first week and a review written in the sixth month would disagree about these two products, and almost every comparison published anywhere is a first-week review.

Memory you can open, in two different senses

Muse keeps plain markdown in a home directory, embedded straight into the request, and editing one of those files inserts a short note instead of rebuilding the whole package.

Town keeps a wiki, a memory store and a people model on the server, and the agent's personality is a document you can open and edit rather than a line buried in a prompt.

Both of these are unusually honest designs, and this is the one axis where this pairing agrees more than it disagrees. Most products on this bench treat what the agent believes about you as vendor configuration. These two both have an object you can point at. The difference is where it lives and therefore who can reach it: Muse's files sit on the VM, Town's documents sit on a server you log into.

Two layers against four stages

Muse keeps a permission sheet of 246 methods carrying no parameters at all, and loads the real definitions a namespace at a time when they are needed.

Town catalogues 1105 tools and narrows them in four stages before anything is offered. The main assistant sees 177 of them, and one background writer is allowed 17.

These are the same insight implemented at different moments. Both products have concluded that a model handed everything chooses badly. Muse solves it by deferring the detail: you can see that a method exists without the context paying for its definition. Town solves it by deciding in advance which tools are on the table at all. One is lazy loading, the other is curation, and the practical difference is that Muse's agent knows what it is not being shown and Town's does not.

Who can start the work

Muse launched on iOS, Android and the web, is reachable in WhatsApp, and since 17 September has had a Mac app that can act in applications on your computer. We examined the web surface only, and say so wherever it matters.

Town gives web and Mac equal weight, and routines can be started by a clock, an email or a calendar entry rather than by you typing.

Town's triggers are the part to think about alongside its foothold. Something that is not you can start an agent that reaches your messages. That combination is the one worth deciding about deliberately, and it is not unique to Town, but Town is where we have confirmed both halves.

Muse's Mac app is the open question in this comparison. If it reaches into applications the way the description says, Muse has quietly acquired the property this page treats as Town's distinguishing feature. We have not examined it, so we are not saying it has.

What neither will show us

For Muse we never saw the raw traffic between the runtime and the model. The structure is known, the contents are not. Nor credentials, live third-party sign-in, or an interactive terminal that exists on the machine and that the owners chose not to open.

For Town the full instruction set is refused by the member-facing interface, and we could not establish where the personality document is spliced in, nor the execution bodies of four newer routines.

Note what the two lists have in common. In both cases the thing we could not reach is the text that governs behaviour, and in both cases we could map everything around it. That boundary is the same for every product on this bench and it is worth naming in a comparison, because the parts a writeup is most confident about are usually the parts that were easiest to see.

Which one the choice comes down to

Strip out the parts both products claim and one question is left: do you want the agent inside your life, or beside it.

Town is inside. Its value is that it sees what you see, and the price is that its mistakes land on your machine, in your message history, where you will find them later. If the work you have in mind is inseparable from what is already on your screen, nothing that runs elsewhere can substitute for that, and the exposure is what you are paying for rather than an accident.

Muse is beside. Its machine is real but it is somewhere else, it accumulates its own working context rather than reading yours, and when it is wrong the wrongness is contained somewhere you do not have to clean up. The trade is that it only knows what has been told to it or written down on its own disk.

The second question is who you would rather hold the state. Both products keep something you can correct, which puts them ahead of most of this list, but one keeps it in a home directory on a vendor machine and the other on a server you sign into. Neither is on hardware you control, and a comparison that implies otherwise because one of them has a Mac app is reading the Mac app wrong: the Mac app is reach, not storage.

Why one of them has no figure

Muse has a published completion rate on our scorecard, with its interval and the number of tasks behind it. Town's row is withheld.

It is withheld because the figure we were given does not reconcile with the only records we can check: five batches for Town survive in the bench repository's history, all five were refused for incomplete domain coverage, and they work out to roughly 13% across 30 scorable tasks. We are not printing a number our own evidence disagrees with, and we are not printing it with a hedge next to it either.

That is the honest state of this pairing. These two can be compared on how they are built, which is what this page is for. They cannot yet be compared on how often they finish an errand, and anyone doing that confidently has not asked where the number came from.

FAQ

What is the main difference between Muse and Town? Whose computer the agent runs on. Muse gets a Linux virtual machine that Meta provides, plus a second one for browser work. Town's native Mac application runs on the computer you already own and reaches your messages, contacts, screen and audio.

Can I correct what either one remembers? Both, which is unusual. Muse keeps plain markdown files in a home directory. Town has a wiki, a memory store, a people model, and a personality document you can open and edit.

Which one has more tools? Town catalogues 1105 and shows its main assistant 177. Muse lists 246 methods on a permission sheet without parameters and loads real definitions only when needed. The numbers are not comparable, because they are counting different things.

Does Muse run on my Mac? The version we examined does not. A Mac app released on 17 September is described as acting in applications on your computer, and we have not examined it, so nothing on this page depends on it.

Why is there no score for Town? The figure we were given does not reconcile with the records we can check, which work out to roughly 13% across 30 scorable tasks. We withheld the row rather than publish a number our own evidence disagrees with.

Is this from testing or from marketing? From our own teardowns of Muse and Town, including the list of what we could not establish for each. For Muse, the web surface only.