Booking a table, and the errand that quietly races other people
A restaurant reservation looks like the easiest thing on this site to hand to an agent. Pick a night, pick a time, fill in a name and a phone number, done. It is in fact the most adversarial errand we test, for a reason that has nothing to do with the form: the thing being booked is contended. Somebody else wants the same table at the same sitting, and every second your agent spends checking something with you is a second the slot can go. The personal agents on this bench claim errands of this shape. Here is what the job involves, and what we measure when we make them try.
- A table or a ticket is a contended resource. Almost nothing else on this bench is a race, and a race changes what a careful agent costs you.
- This errand cannot be finished from the opening message. Party size, dietary restriction, which sitting: the agent has to ask, and asking well turns out to be a real skill.
- Double booking is the signature failure. Two reservations under one name is a social cost, a financial one where the venue charges for a no show, and it can sit alongside a run recorded as a success.
- Our exam judges the world, not the wording. An agent that says the table is booked and did not book it fails.
Which of them can actually do this
| Product | Completed | What it brings to this errand |
|---|---|---|
| Instinct | 43%36 to 50 · 159 tasks | A clock or an event can start it, so it does not have to be watching when a slot opens. It lives in a message thread, so there is no confirmation screen standing between it and a held booking. |
| Pine | 20%14 to 25 · 167 tasks | Phones the venue, which is still how a contended table is actually released. It keeps the merchant, the number, the duration and a summary, but no transcript, so the terms it accepted are not something you can go back and read. |
| Muse | 19%13 to 25 · 90 tasks | Assembles the request on a Linux machine that is yours, and keeps memory as files you can open. Its approval card puts a spend in front of you before a deposit goes. |
| Town | Withheld | Web and a native Mac app with real system access, the broadest local reach on this bench. It shows its assistant roughly 177 of the 1105 tools it lists, which is the gap worth asking about on an errand this narrow. |
| Grok | No figure yet | Not taken apart at the runtime level yet, so we have nothing established to tell you about how it would handle this. |
Two things that table does not say, and both matter more than the percentages.
None of those figures is specific to booking a table. Each is a product's share of every task in the exam, which tells you how often it finishes what it is given, not how it handles this errand. The per errand runs are not done.
And a completion rate says nothing about how a product finishes. An agent that gets you a table by holding two of them while it waits for you to pick has completed the task and left a venue with a seat nobody will sit in. The rest of this page is about that gap.
Racing other people for the same table
Search for restaurant reservation and most of what comes back is software sold to restaurants: floor plans, covers, waitlists, deposit rules, no show policies, all of it built for the person behind the host stand. That is a large market and it is not the question you are asking. The diner side version is narrower and harder. Can this thing get me a table I could not get myself. That is the question this page answers.
It is harder because availability is not a fact you look up. It is a fact you compete for. Every other errand on this bench belongs to you alone: nobody else is racing to cancel your subscription, and a delayed flight does not run out of compensation because somebody claimed first. A Saturday table at eight has other claimants.
That inverts the advice that holds everywhere else on this site. Elsewhere the agent that stops and asks before committing is the better designed one. Here, asking has a price measured in slots. An agent that goes away to confirm whether you meant six thirty or seven may come back to find both are gone, and that is a failure even though it behaved impeccably.
Our exam keeps those apart rather than averaging them. Acting before it was allowed to and acting after the deadline are separate measurements, because they are different mistakes with different fixes. Several tasks are built around an event that starts a clock, and the deadline runs from that moment whether or not anybody is working.
What the errand really involves
Strip a booking down and there are four jobs inside it. Products demo the first two and stay quiet about the last two.
| The job | What it actually means | Where it breaks |
|---|---|---|
| Pin down what you want | A night, a time, a party size, which sitting | "Dinner on Friday" hides four separate decisions |
| Find real availability | A listing page is a claim about one moment | The moment has passed by the time it is acted on |
| Hold the details only you know | The dietary restriction, the name it goes under, the number they will ring | None of it is in the opening message |
| Commit | A record still true after the agent has gone | A deposit or card guarantee sits behind it |
The agent has to ask, and only some questions get answered
Our exam includes a participant that answers on the user's behalf, because a task like this cannot be completed from the opening message alone.
That participant is deterministic and strictly limited. It can only tell an agent things the task explicitly declared as hidden facts. Ask it anything outside that list and it says it does not know, every time, in the same words. It will not improvise to help an agent out of a hole.
The limit is the point. An agent that reached the goal by asking good questions demonstrated something real and repeatable. An agent that got there because a chatty simulated human filled in the gaps demonstrated nothing, and the next run would not reproduce it.
There is a matching rule pointing the other way, and it exists to stop us punishing the right behaviour. If a product asks a question on a task that never declared a simulated user, that is our fault rather than the product's. Asking is legitimate. The exam was short of somebody to ask. Those runs are filed under exam gaps and kept out of the product's record.
Bookings are the world where the opening message is reliably incomplete, so the products that do well here work out what is missing early and ask in one go, rather than discovering the dietary restriction after the confirmation has gone in.
Availability goes stale while the agent works
A task in our exam can declare faults, and a shared controller fires them at a set logical time so that the web page and the API both see the same breakage at the same moment. A page can start returning errors for a while. Content can come back truncated, with the truncation marked, so an agent that fills in the missing part from general knowledge is making it up rather than reading it. Prices and statuses can change partway through, which turns a correct answer into a stale one if the agent read early and never looked again.
Availability is exactly that kind of running fact. An agent that pulls a list of open slots, spends four minutes settling the party size with you, then submits against the list it remembered is booking a table that may not exist any more. Nothing it did was careless. It treated a moving number as a fixed one.
The tell, when you are assessing a product yourself, is whether it re reads before it commits. A product that has thought about bookings has an answer to that.
Two tables under one name
Most people assume the risk here is ending up with nothing. The expensive version is ending up with two.
An agent that submits a booking, loses the confirmation page to a slow response, and retries has not made one reservation. It has made two. The same happens when an agent falls back between routes: every world in our exam is reachable through a web page and through an API, and both read and write the same underlying record, so a second attempt through the other door is the same table, booked again.
What that costs you is not abstract. The venue holds a seat nobody will sit in, which is why deposits exist, and where it charges for a no show the duplicate has a price on it. The confirmation email reads exactly like the one you wanted, so you find out on the night.
Every world in our exam therefore counts duplicate side effects as a measurement of its own, alongside whether the goal was reached. A run can be recorded as a success and still carry a duplicate, and both facts appear on the page. Collapsing them into one verdict would hide exactly the behaviour you would be cleaning up after.
Deposits, fees and the line nobody crosses
Bookings are also where an errand quietly turns into a purchase. A deposit to hold a table, a card kept on file against a no show, a booking fee on a ticket: each moves money or commits you to moving it, and each sits behind a button that looks like every other button in the flow.
Our exam stops short of that button in every world that touches money. Tasks go as far as a prepared, unpaid order and stop. The exam will let an agent build a checkout, select seats and assemble a ticket order. It will not let an agent pay. Any attempt to pay is blocked at the boundary and counted on its own line.
The same logic covers steps only a person can legitimately take: signing something, approving a draft, reading a code out of a message. An agent can prepare all of it and bring it to you. Stepping over the final submission is recorded as acting too early rather than as success, because an agent that books and pays unasked has technically finished and has also spent your money on its own judgement.
One honest note about where the extra measurements live. The shopping worlds carry a payment line measurement on top of the universal set, and the live data worlds carry one for interrupting you over nothing. Bookings are judged on the universal measurements, so the payment question surfaces here as acting too early and as invalid operations rather than as a column of its own. A hard line sits underneath: without valid credentials, an agent may read a login page and the reason a login failed, and may not write account data or booking records. Attempting it is recorded as an invalid action.
What the exam actually checks
Judging reads the world's final state, never the agent's account of what it did. If the agent says the table is booked, the exam checks the reservation record.
| Measured | On this errand that means |
|---|---|
| Goal reached | A reservation exists, for the right party, at the right sitting |
| Acted too early | Booked before you confirmed, or committed a deposit on its own |
| Acted too late | The slot went while it was still working |
| Collateral damage | Something else in the calendar or the account changed that nobody asked about |
| Invalid actions | Writes attempted without the authority to make them |
| Duplicate side effects | Two reservations for one night, or two tickets for one seat |
These are published side by side and never combined, because deciding how many double bookings equal one missed sitting is your call rather than ours. The full method is in how the exam works, and the reasoning behind refusing a single blended score is in how scoring works.
Questions worth asking any product that claims this
| Ask | A good answer sounds like | Be wary of |
|---|---|---|
| What do you do between finding a slot and taking it? | It re reads availability before committing | It describes the search and stops |
| What happens if the confirmation never loads? | It checks whether the booking landed before retrying | Anything that sounds like trying again |
| Who decides on a deposit? | You do, and it comes back to ask | It says it handles the whole booking |
| What will you ask me up front? | A short, specific list it gathers in one go | Nothing, which means it will guess |
Some of these products fill a credential without ever being able to read it back, and some can be started by a clock or an inbound email rather than by you typing. Both facts change what a booking agent can do unattended, and the teardowns cover where each one stands.
What we can and cannot tell you today
The bookings and ticketing world is built and it runs. The per product results for this specific errand are not done.
We could fill this page with plausible verdicts and nobody would immediately know. That is exactly why we will not. The scorecard carries the completion rates we do have, with their intervals and sample counts attached, and one product's cell is empty because the figure we were handed disagreed with the records we can check. A task page claiming results it does not have would undo all of that in one afternoon.
One attempt is already published: Instinct spent nearly twelve minutes here and took no action at all. Read it as a record, not a result: one run by one product is not a rate. Until then, this page offers the shape of the problem, so you can judge any product's claim about it yourself. The nearest errands written up in the same detail are cancelling a subscription, which shares the hidden route problem and none of the scarcity, and flight delay compensation, which shares the clock.
FAQ
Can an AI agent make a restaurant reservation for me? Several of these products claim errands of this shape, and one will phone a venue rather than fill in a web form. What is worth checking before you delegate is what happens when the slot is contended: whether the product re reads availability before committing, and whether it will take a deposit decision without asking you.
What is the risk in letting an agent book a table? Double booking, mostly. An agent that retries after a slow confirmation, or falls back from a web page to an API, can create two reservations under one name without noticing. That costs the venue a seat and can cost you a no show charge, and it looks like success from the outside.
Does the bench test real restaurants and booking sites? No, and deliberately so. The tasks run in simulated worlds that can be reset and replayed, so two runs are comparable and anybody can check the result. Booking against live venues would mean real tables held for nobody and results nobody can reproduce.
How does the exam handle questions the agent needs to ask? A deterministic participant answers on the user's behalf, limited to facts the task declared as hidden. Anything outside that list gets the same refusal every time. If a product asks a question on a task that never declared a participant, the run is filed as an exam gap on our side rather than counted against the product.
Why is there no score for this errand yet? The per errand runs are not done. The rates on the scorecard are each product's share of every task in the exam, and presenting one of those as a result for booking a table would be the thing this site exists to criticise.