Where a personal agent actually runs, and why that is the first thing to ask
Where a personal agent runs is the property you cannot change after you have committed to it. You can switch models, revoke a connector, clear a memory. You cannot move the machine the thing executes on. It is also, unlike almost everything else about these products, answerable today by observation rather than by waiting for a score, which is why it is the first axis on our scorecard and why we opened four products to establish it.
- We took four apart and found four different answers, none of which matched the marketing.
- Nine more make public claims about where they run. We have not checked any of them, and this page keeps that line visible.
- The spread runs from a product that executes nothing on your machine to one that runs on a five dollar chip with no operating system.
- A leased machine is a third party sitting inside an authenticated session, and almost nobody names who the lessor is.
Why this one before the others
Pick an agent and you are choosing a place for code to run while holding your credentials. Everything people usually compare, the model, the integrations, the price, sits on top of that choice and can be changed later. The execution substrate cannot.
It also determines what a failure can reach. An agent confined to a disposable container that is destroyed after each task can get an errand wrong. An agent with a native foothold on your Mac, reaching messages, contacts, screen and audio, can get an errand wrong in a place where the mistake persists. Neither is better in the abstract. They are different exposures, and the product page rarely says which one you are buying.
The reason it is answerable now is that execution leaves traces. Processes, hostnames, session behaviour and the shape of what comes back all narrow it down, which is the whole method behind how we take them apart.
What we found in the four we opened
These are our own findings, from products we examined rather than read about.
Pine runs nothing of substance on your machine. Neither the web app nor the Mac app runs the model. Browser work is handed to a separate agent on your own computer, while the cloud route leases you a remote desktop that you watch over VNC. Two different execution places for two different kinds of work, in one product.
Instinct orchestrates on a server with the work itself running in a sandbox, and browser tasks go to a leased cloud browser rather than the Chrome on your desk. Nothing local, and the browser is somebody else's.
Muse gives every user a Linux virtual machine. The request is assembled on that machine rather than arriving as a finished blob, and browser work spawns a second, separate machine. One user, two machines, both Meta's.
Town handles sync and sessions server-side, but its Mac app is native and genuinely reaches into the system: messages, contacts, screen and audio. It is the only one of the four with a real foothold on the computer in front of you.
Four products, four architectures, and no two of them exposing you to the same thing.
What the others say, which is not the same as what we found
Nine more products on our landscape describe where they run. We have not opened any of them, so what follows is their claim or a publication's reporting, labelled as such throughout.
Pickle says it books, orders and pays on a computer of its own, asking before each one. MimiClaw says it runs on a five dollar chip with no operating system, no Node and no small server behind it. Halo says it runs entirely on the phone with no backend at all, connecting to mail, calendar, reminders and health locally and calling models through your own keys. OpenClaw and Hermes Agent both say you run them on your own hardware or a managed host.
For four more, the source is a comparison published by Every on 30 September rather than the vendor. It places OpenAI's Dot in OpenAI's cloud with optional access to a local computer, Gemini Spark in Google's cloud, Poke in Poke's cloud, and Grok Bot in Cursor's cloud, which is the one claim on this page that two sources disagree about.
From nothing on your machine to a five dollar chip
Lay the claims out and the range is wider than any feature comparison would suggest.
At one end, Pine executes nothing locally that matters. At the other, MimiClaw claims to run on hardware costing less than lunch, with no operating system underneath it. Between them sit a Linux VM per user, a native Mac app with system access, a phone with no server behind it, and several arrangements where the machine belongs to a third company.
That spread is the strongest argument we know of against treating these products as interchangeable. Two assistants with identical feature lists can differ on this axis more than two pieces of software usually differ at all. It is also why a single leaderboard number flattens something worth seeing: the number says how often the errand got done, and says nothing about where it got done.
A leased machine is a third party in the room
Three of the arrangements above involve a computer that belongs to neither you nor the product's maker.
Instinct's browser tasks run in a leased cloud browser. Pine's cloud route leases a remote desktop. Grok Bot, if Every's reporting is right, runs on Cursor's cloud. In each case something holding a logged-in session to your accounts is executing on infrastructure operated by a company you did not choose and are mostly not told about.
This is not an accusation. Leasing browser infrastructure is ordinary engineering and the alternatives are expensive. The point is narrower: the number of parties with access to a live session is a fact about the product that almost no product states, and it is the kind of fact that only shows up if someone looks.
Running it yourself moves the problem rather than removing it
The self-hosted answers, OpenClaw and Hermes Agent and the rest, look like the clean solution and partly are. Nothing leaves your hardware that you did not send. There is no lessor.
What does not go away is everything the agent can reach once it is running. An agent on your own machine with your own credentials can still send the wrong message, accept the wrong offer and spend the wrong money, and now there is no vendor-side log to reconstruct what happened from. Self-hosting changes who holds the risk. It does not reduce the surface, and in one respect it enlarges it: the public audit of twenty projects claiming to self-host one of these products found tools asking for privileged access and leaning on undocumented interfaces, which is a category of risk the hosted versions do not have.
What actually changes for you
Four practical consequences follow from the answer, and they are the reason to ask.
What a mistake can touch. A disposable container limits the blast radius in a way a native desktop app does not. How much you can reconstruct afterwards. A vendor-run machine usually has logs you can request; your own laptop usually does not, and neither does a leased browser you never knew existed. Whether it works when you are asleep. An agent that needs your computer awake is a different product from one that does not, whatever the marketing implies. And who else is in the session, which the leasing arrangements above make a real question rather than a theoretical one.
None of this tells you which product to pick. It tells you which question you are actually deciding.
Finding out for something not on this list
The method generalises, and most of it needs no special access.
Watch whether work continues after the laptop sleeps. Look at what the installed application asks the operating system for, since a native app's permission prompts describe its reach more honestly than its website does. Notice whether browser tasks behave like your browser, carrying your logins and extensions, or like a clean one somewhere else. Check whether anything appears in local processes at all while a task runs.
Those four observations will usually separate server-side from local, and will often reveal a leased machine, which is the thing most likely to be absent from the documentation. What they will not give you is the detail, and the detail is where the four teardowns above ended up.
FAQ
Why does it matter where a personal agent runs? Because it decides what a mistake can reach, what you can reconstruct afterwards, whether the agent works while your computer is off, and how many companies have access to a session holding your credentials. It is also the one property you cannot change later.
Which personal agents run on your own device? Among products we have examined, Town is the only one with a real native foothold, reaching messages, contacts, screen and audio on a Mac. Several others claim local or self-hosted execution, including Halo, MimiClaw, OpenClaw and Hermes Agent, and we have not verified those claims.
What is a leased browser, and why does it come up here? It is browser infrastructure rented from a third company and driven remotely. Two of the four products we opened use one for browser work. It matters because something is operating a logged-in session to your accounts on a machine belonging to a company you did not pick.
Is self-hosting safer? It removes the lessor and keeps data on your hardware. It does not reduce what the agent can do with your credentials, it removes the vendor-side record you would otherwise use to find out what happened, and the self-hosting ecosystem around at least one product contains projects asking for more access than the product does.
How did you establish this for the four? By examination rather than by reading documentation, using the approach in how we take them apart. Every claim about the other nine on this page is attributed to the vendor or to a publication, and is not a finding of ours.
Does a benchmark score tell you any of this? No. A completion rate says how often the errand finished. It cannot distinguish an agent that finished it inside a disposable container from one that finished it on your laptop, and those are different products to live with.