Personal Agent Bench

Grok Bot, and the difference between a coworker and an errand runner

Grok Bot is X Corp's personal agent, released in August 2026. You give it a task by message, it signs into your accounts, and it works your tools and websites the way you would, continuing after you close your laptop. It is the fifth product in our exam. It has no published score yet, and unlike the four on the scorecard we have not taken it apart at the runtime level, so this page is an account of what is known about it and where each piece came from.

  • It is sold as a coworker, not an assistant. The tasks it names are sales outbound, inbox triage, recruiting and expenses, which is a different job from the errands this bench scores.
  • Its defining move is signing in once and then using your accounts unattended. That is the capability, and it is also the risk.
  • Three sources gave three different answers for who makes it. The App Store listing, which the vendor fills in itself, says X Corp.
  • We have not examined it. Nothing on this page is a measurement, and the page says which claims are the vendor's own.

What it actually is

The product description is unusually direct about the mechanism. You log it in once. After that it uses your apps and websites as you would, and the listing names the awkward ones specifically: vendor portals, ad managers, CRMs, inboxes and older operations tools. It says the work lands where a human colleague would put it, rather than in a separate automation system you have to supervise.

Two further things are claimed. The first is that several bots can run at once, in parallel, passing work between themselves in a shared thread, with one promoted to coordinate when you want a single update. The second is that you can show a bot how a job is done by having it watch you do it once.

It is on iOS and desktop, free to install, listed under Productivity, and requires iOS 18. At the time of writing it is on version 1.14.0, updated on 6 October 2026, first released on 11 August 2026. It carries close to nine thousand ratings at an average just under 4.9, which is a larger public footprint than most products in this category have.

Sold as a coworker, which is a different job

Read the tasks it advertises and the positioning is unmistakable. Sales outbound. Inbox triage. Recruiting. Expenses. Account follow-ups. Ad managers and CRMs.

That is work, not life. The products this bench scores are aimed at the things that pile up outside a job: the subscription that will not cancel, the flight that needs changing, the charge that should not be there. The overlap is real, because both end in an agent operating an account that belongs to you, but the failure that matters is not the same one. A bot that sends a clumsy outbound email costs you a prospect. An agent that accepts a retention offer while trying to cancel costs you money every month until you notice.

It is in our exam anyway, and should be. The question of whether an agent can be trusted with a logged-in session does not care what the session is for.

What we will not do is read a result from one into a claim about the other. If it scores well on errands, that is a statement about errands.

The capability and the risk are the same sentence

"Log Grok Bot in once. It uses your apps and websites just like you would."

Strip the marketing and that is a standing, unattended, authenticated session across whatever you connected. Everything useful about the product follows from it, and so does everything worth worrying about. There is no version of this where you get the first half without the second.

The listing does name a boundary: each agent finishes multi-step work end to end and "only comes back when your approval is needed." That is the right shape for a control, and it is the exact claim our exam is built to test, because the gap between a product that asks before an irreversible step and one that asks after is invisible from any description. Our own scored tasks include a booking that must stop before payment and wait, for precisely this reason.

We have not run it yet. The approval behaviour described above is the vendor's claim, not our finding.

Who makes it, and why that was not obvious

Three sources gave three answers, which is worth recording because it is the kind of small error that propagates.

Our own landscape had the vendor as SpaceXAI, taken from a news summary. Every's comparison, published 30 September, labels it Cursor and xAI, says it runs on Cursor cloud, and prices it from Cursor Pro at twenty dollars a month. The App Store listing, which a vendor fills in itself and which is therefore the closest thing to a primary source here, names the publisher as X Corp with a seller address of x.ai.

We corrected our entry to X Corp on the strength of that. The Cursor half is unresolved and left that way. One thread does run through it: a security audit of self-hosting projects, described below, found a tool in that ecosystem asking for a Cursor session. That is a hint about a relationship, not an answer, and x.ai returns an error to anything that is not a browser, so we could not check further.

What other people have found

Four public findings about Grok Bot have passed through our radar. None is ours, each carries its source, and they are a mixed bag in a way that is itself informative.

The most substantive came on 5 October. In a single overnight job, a six-bot crew completed 612 tasks, blocked 47 actions and rejected 89 handoffs. The roles were split on purpose: scouting and scoring fed a risk check, and only the execution bot could take a final action, and only after a timestamped pass. A separate bot carried sources, read times and evidence with every handoff, while a coordinator enforced the order without doing specialist work. Whatever else that is, it is somebody taking the irreversible-action problem seriously inside a multi-agent setup.

The second is a security audit, published 1 October, of a circulating list of twenty GitHub repositories advertised as ways to self-host Grok Bot. Every link led to a real repository, but the projects differed sharply in what they were and what access they demanded: rule documents, self-hosted alternatives, decompiled copies of a commercial app, and tools leaning on undocumented interfaces. A working link is not a recommendation, and the access a project asks for matters more than its place on a popular list.

The third, from 25 September, concerns Grok 4.7 rather than the agent product. The model got stuck in a loop and printed text presented as its system prompt, including instructions not to disclose internal reasoning and to give system messages priority over user messages. We list it because it is about the same family, and we separate it because a model behaviour is not an agent behaviour.

The fourth is thin and we are saying so: a post reporting that a domain was bought and pointed at Grok Bot, without naming the domain. We keep it because the radar records what was found rather than only what was impressive.

What a teardown would have to answer

The four products on our scorecard each have a page describing where they run, what tools they hold, what they remember and how they reach you. Grok Bot has none of that here, and these are the questions it would need to settle.

Where does the work happen. Every's table says Cursor cloud, which if accurate is a notable answer, because a machine that is neither yours nor the vendor's is a third party in your authenticated session. What tools does it actually hold, as opposed to the categories the listing names. What persists between tasks, given the claim that bots "remember preferences across handoffs" and get better over time, which is a memory claim with no stated shape. And what the approval gate is really wired to, since the only published description of it is one clause in a store listing.

Until those have answers from somewhere other than the vendor, this page stays what it is.

Why it is listed without a score

A product can sit in our exam for a while before it has a number, and the scorecard says so rather than leaving a gap. Grok Bot is marked pending: in the exam, no published figure, no teardown.

That is a deliberate state rather than an oversight. Publishing a rate before the runs are clean is how a benchmark acquires a number it then has to defend, and we have one product already, Town, whose figure we are withholding because it does not reconcile with the only records we can check. Carrying an honest blank is cheaper than carrying a wrong number.

When a figure does arrive it will appear with a range and a task count beside it, the same as every other row, and this page will say what changed. The method behind that is in how the exam works.

FAQ

What is Grok Bot? A personal agent from X Corp, released in August 2026. You message it a task, it signs into your accounts, and it operates your apps and websites on your behalf, continuing while you are away. It runs on iOS and desktop and is free to install.

Who makes it? The App Store listing names X Corp as the publisher, with a seller address at x.ai. Other published sources have said SpaceXAI and have said Cursor and xAI, and that disagreement is covered above.

Have you tested it? Not yet. It is in our exam with no published figure, and we have not taken it apart at the runtime level either. Everything on this page is attributed to the vendor's own listing, to a third party, or to a source conflict.

How is it different from the agents you have scored? The tasks it advertises are work: outbound, inbox triage, recruiting, expenses. The ones we score are personal errands. Both end with an agent acting inside an account that belongs to you, which is why it is in the exam, but the consequences of a mistake differ.

Does it ask before doing something irreversible? Its listing says an agent "only comes back when your approval is needed." That is the vendor's description, not a finding of ours, and establishing what that gate is actually attached to is one of the things a teardown would have to do.

Can I self-host it? There are repositories claiming to let you, and a published audit of twenty of them found real projects of sharply varying kinds, some asking for privileged access or leaning on undocumented interfaces. Treat a working link as the start of the question rather than the end.