Every personal AI agent we know of, and how far we have got with each
Personal AI agents arrive faster than anybody can take them apart. This page is the list we keep so that the gap between what exists and what we have examined is visible rather than implied. Four of them have been through the exam. One more we have read rather than probed. The rest are recorded here with a link and nothing else, because a name with no evidence behind it is not a finding.
- The list separates what we have examined from what merely exists. Most entries are in the second group and say so.
- Being on this list is not an endorsement and the order is not a ranking. Entries are grouped by how far we have got, which is a fact about us rather than about them.
- Every entry carries a link you can open. Names we could not trace to an official site or a first-hand report are held back, and there are more of those than there are entries here.
- There is no score column. The four products with figures have them on the scorecard, with intervals and sample counts, and those figures do not transfer to anything else on this page.
The list
| Product | How you reach it | Where we have got to | Source |
|---|---|---|---|
| InstinctSpear Street Technology | A message thread | Scored. Through the exam, rate on the scorecard | Open |
| MuseMeta | Phone apps and web | Scored. Through the exam, rate on the scorecard | Open |
| Pine19pine | It telephones for you | Scored. Through the exam, rate on the scorecard | Open |
| Towntown.com | Web and a Mac app | Scored. Through the exam, rate on the scorecard | Open |
| OpenMuseCopilotKit | You run it yourself | Taken apart. Examined and written up, not scored | Open |
| Grok BotSpaceXAI | A message thread | Listed. Recorded that it exists, nothing more | Open |
| OllieOllie | A message thread | Listed. Recorded that it exists, nothing more | Open |
| OpenPokeCommunity implementation | You run it yourself | Listed. Recorded that it exists, nothing more | Open |
| PallyPally | A message thread | Listed. Recorded that it exists, nothing more | Open |
| PokeThe Interaction Company, acquired by Cognition in July 2026 | A message thread | Listed. Recorded that it exists, nothing more | Open |
| TomoTomo | A message thread | Listed. Recorded that it exists, nothing more | Open |
What the three states mean
Scored means the product has been through the exam and has a completion rate on the scorecard, with an interval and a task count beside it. Four products are in this state.
Taken apart means we have examined how it is built and published what we found, without it having been through the exam. A teardown answers a different question from a rate: one is about construction, the other about outcome.
Listed means we have recorded that it exists, what it is reached through, and where to read about it. Nothing more. It is the honest state for most of this category most of the time, and writing it down is the point of the page.
The states are about our coverage, not about product quality. A listed product is not worse than a scored one. It is one we have not got to.
Why most of them are only listed
Taking a product apart is slow. Getting an account, watching what it does, working out where the model runs and what it can reach, and then writing only what the evidence supports takes far longer than a new one takes to launch.
That asymmetry is permanent, and pretending otherwise would mean either thinning the work until it says nothing, or quietly dropping the products we have not reached. This page is the third option. It keeps the unexamined ones visible, with their state attached, so the shape of what we have not done stays legible.
What we record about each one
Four fields, deliberately few.
The name and who makes it, because ownership changes and several of these have already been acquired. The surface, meaning how you reach it: a message thread, a phone call, an app, a web page, or your own machine. That single fact predicts more about how a product behaves than any feature list, and it is the first axis of every teardown we publish. The state, from the three above. And a link to something you can open.
No pricing, no ratings, no feature grid. Pricing for the four examined products sits on their own pages with the date it was checked, because prices go stale faster than anything else on a site like this.
What the surfaces tell you
Sort this list by how you reach each product and the category splits into groups that behave differently, which is why that column is there rather than a feature count.
The largest group lives in a message thread. You text it the way you would text a person, and it answers in the same place. That choice removes the confirmation screen every other kind of software leans on: there is no plan to review, no checkbox before it acts, no window showing a half finished thought. Everything the agent wants to tell you has to survive as a text message. It also means the agent can reach you whenever it decides to, which is the same capability seen from the other side.
A smaller group is a web product with apps attached. There is a surface to put a decision on, which matters whenever the next step costs money or sends something irreversible, and a place for a notification to sit rather than having to interrupt you.
One of them telephones on your behalf. That is a narrower job and a harder one, because a phone call cannot be undone, paused or edited once the other party has heard it.
And a few you run yourself. Those trade convenience for the ability to read the source, which is a real trade rather than a marketing distinction: it is the difference between believing a privacy claim and checking one.
Names move, and so does ownership
Two entries on this list have already changed hands, and a third shares its name with an unrelated company.
That is ordinary for a category this young, and it is a practical problem rather than a trivial one. Search for one of these products and you can land on a different company with a similar name, a predecessor that was acquired, or a community project that reimplemented it. We have had to separate at least one pair by hand while assembling this page, and the notes on individual entries record which is which where it is not obvious.
It is also why the source column exists. A link you can open settles which product an entry means, in a way a name on its own does not.
The names we are not showing
A competitor publishes a longer list. We read it, and fifteen of the names on it are not here.
Those are names we could not trace to an official site or a first-hand report in the time we spent looking. They may well be real products. They are held back because listing a name nobody can open is the kind of filler that makes a list longer without making it more useful, and because the whole point of this page is that every line can be checked.
If one of them is yours, or you simply know where it lives, that is exactly the sort of correction worth sending, and it is the fastest way for an entry to appear here.
What it takes to move up a state
The three states are not a queue we work through in order, and nothing is promised about how long a product stays in one.
Moving from listed to taken apart needs access first. Several of these are invite only or waitlisted, so the wait is not ours to shorten. Once inside, the work is watching what the product actually does rather than what its documentation says: where the model runs, what it can reach, what it keeps between sessions, and which of those we can establish with evidence rather than infer. Every teardown we publish ends with a list of what we failed to establish, and that list is usually the longest section.
Moving from taken apart to scored needs something different again. The exam has to be able to drive the product without a person in the loop, which not every one of these allows, and the tasks have to run in worlds that can be reset and replayed so two attempts are comparable. A product that can only be operated by hand can be examined but not scored, and that is a property of our exam rather than a verdict on the product.
One consequence worth stating plainly: the four products with scores are not the four best, or the four biggest. They are the four that were reachable when the exam was built. Reading the scorecard as a ranking of the category would be reading it as something it cannot be, which is exactly why this page exists beside it.
How this list gets used
It is not only a page. The same file drives what the radar watches, so adding a product here puts it in front of the collector the next morning, and third-party findings about it become publishable rather than being rejected as off topic.
Before 2026-09-24 that collector watched four products and nothing else, which meant a published finding about any other personal agent was invisible to us twice over: never searched for, and refused at the gate if it somehow arrived. The list fixed both, which is a better reason for it to exist than being a page.
FAQ
How do you decide what counts as a personal AI agent? Something aimed at an individual rather than a team, that acts on their behalf rather than only answering questions, and that reaches real accounts or services to do it. Chat assistants that only produce text are out, and so are tools sold to companies for their staff.
Why is there no score next to most of these? Because they have not been through the exam. Four have, and their rates are on the scorecard with intervals and task counts. Putting a number next to the others would mean inventing one, and a benchmark that does that has stopped being a benchmark.
Does listing a product mean you recommend it? No. The list records that a product exists and how it is reached. Several entries are ones we know very little about, which is precisely what the listed state means.
How often does this change? The date at the top is when the list was last checked. New products appear in this category faster than they can be examined, so treat the list as current to that date rather than as complete.
One of these is mine and the details are wrong. Corrections are welcome, particularly about the surface and who owns it, since both change. The same applies to the names held back for lack of a traceable source.