Rearranging a family week, and who has to be told
Most people looking for a family schedule app are trying to stop the same small disaster repeating. The swimming lesson moves, three adults hear about it at three different times, and a child is left outside a locked door at half past four while everyone assumes somebody else is driving. Shared calendars fix part of that by putting the week somewhere everyone can see it. What they do not do is the telling. They coordinate humans, and the humans go and do the work. The personal agents on this bench claim something else: that the agent reads the reschedule notice, messages the other parents, moves the entry and sorts out the lift itself. That is worth testing, because once software writes to other people on your behalf, the cost of getting it wrong stops being yours alone.
- A shared calendar shows the week to everybody. Rearranging it means writing to other people, and a message sent to the wrong parent cannot be unsent.
- This is the errand where an agent has to ask. Our exam will answer, strictly and narrowly, and an agent that guesses rather than asks is measured for it.
- The clock starts when the reschedule notice arrives. Acting before that moment is recorded as its own failure, because an agent that books the new slot early guessed correctly rather than worked correctly.
- Our exam judges the world, not the wording. The per-errand runs are not done, and the completion figures below are each product's share of every task in the exam.
Which of them can actually do this
| Product | Completed | What it brings to this errand |
|---|---|---|
| Instinct | 43%36 to 50 · 159 tasks | Lives in WhatsApp, SMS and iMessage, which is where household coordination already happens. That also means there is no confirmation screen standing between it and a message sent to somebody else. |
| Pine | 20%14 to 25 · 167 tasks | Works the phone, which is the right instrument for the one participant who does not use apps. It keeps the merchant, the number, the duration and a summary, but no transcript, so what was agreed with the other parent is not something you can go back and read. |
| Muse | 19%13 to 25 · 90 tasks | Keeps memory as plain files you can open, so what it believes about your household is inspectable rather than inferred. Its approval card puts a send in front of you before it goes. |
| Town | Withheld | Web and a native Mac app with real system access, the broadest local reach on this bench, so it can reach calendars and contacts on your own machine. It shows its assistant roughly 177 of the 1105 tools it lists. |
| Grok | No figure yet | Not taken apart at the runtime level yet, so we have nothing established to tell you about how it would handle this. |
Two things that table does not say, and both matter more than the percentages.
None of those figures is specific to household scheduling. Each is a product's share of every task in the exam, which tells you how often it finishes what it is handed, not how it behaves when finishing means messaging somebody who is not you.
And a completion rate says nothing about how a product finishes. An agent that gets the calendar right by messaging four people a time nobody agreed has completed the task and created a morning of phone calls. The rest of this page is about that gap.
What a shared calendar does, and where it stops
The established products in this category are good at one thing: making a household's week visible in one place. Colour coded children, a shared list, a reminder that fires at everybody at once. Most of them have solved it.
What they are is a display surface with write access. The rearranging still happens in a group chat, on a doorstep, or in the ten minutes you spend working out who can leave work early. The calendar records the outcome of that negotiation. It does not conduct it.
So the useful question is not which shared calendar is prettiest. It is whether anything can take the negotiation off you: read the notice, spot that Thursday now clashes with the other child's appointment, tell the parent who does the Thursday lift, and only then move the entry. That is a different job from displaying a week.
What the errand really involves
Strip a household reschedule down and there are four jobs inside it. Products tend to demo the first two and stay quiet about the last two.
| The job | What it actually means | Where it breaks |
|---|---|---|
| Notice the change | An email from the club, buried under everything else | The notice is vague about which week it applies to |
| Work out the knock-on | The new slot collides with something already in the week | The conflict is in somebody else's calendar, not yours |
| Tell the people | Other parents, a grandparent, possibly a coach | A wrong message is out of your hands the moment it sends |
| Move the entry and the lift | The calendar write, plus whoever is now driving | Two invites, or a lift arranged with nobody confirmed |
The first two are reading problems. The last two are permission problems, and permission problems are where delegating your household admin either works or quietly does not.
The errand where the agent has to ask
Household scheduling is the clearest case of a task that cannot be finished from the opening message. Who is the emergency contact. Which grandparent is free on a Thursday. Whether the other child's appointment could move instead. None of that is deducible, and an agent that invents an answer has made something up about your family.
Our exam is built for that. It includes a participant that answers on the user's behalf, and it is deliberately unhelpful. It is deterministic, and it can only tell an agent things the task explicitly declared as hidden facts. Ask it anything outside that list and it says it does not know, every time, in exactly the same words. It will not improvise a plausible answer to get an agent out of a hole.
That limit is the value of it. An agent that reaches the goal by asking the right questions has shown something repeatable. An agent that got there because a chatty simulated human filled in the gaps has shown nothing, and the next run would not reproduce it.
There is a matching rule pointing the other way, and it protects the products rather than us. If a product asks a question on a task that never declared a simulated user, that is our failure. Asking is legitimate behaviour and the exam was short of a participant, so those runs are filed under exam gaps and kept out of the product's record entirely. A bench that punished an agent for checking with you would be training the behaviour you least want here.
The clock starts when the notice arrives
Several tasks in the exam are built around an event that starts a clock, and a reschedule notice landing in an inbox is one of them. The agent then has a deadline measured from that moment.
Both ends of that are recorded separately. Acting after the deadline is a failure even when the final state is right, because an arrangement that lands after everybody has made other plans is not what was asked for. Acting before the triggering event is a failure too, and that one is more interesting. An agent that books the new slot before the organiser has moved anything guessed correctly rather than worked correctly, and a bench scoring those two the same would be rewarding a product for being lucky.
The worlds can also change underneath an agent partway through. Statuses move, pages return errors for a while, and content can come back truncated with the truncation marked, so an agent filling in the missing half from general knowledge is inventing rather than reading. An agent that read the notice once, spent several minutes composing messages and never looked again can be confidently telling three people about a change that was itself revised.
Messaging other people is where the damage lands
Every world in our exam counts collateral damage as one of its universal measurements: whether the agent damaged something the task was watching, kept separate from whether the goal was reached. On most errands that means a setting nobody asked about. Here it means a person.
This is the errand where acting on your behalf means acting on other people. A cancelled subscription can be resubscribed and a calendar entry can be moved back. A message telling another family that Thursday is off, sent before Thursday was actually off, is in their hands and in their plans, and no undo button exists at either end.
It is also why the boundary matters more here than almost anywhere. Some steps are yours alone: agreeing a new time for the household, committing somebody else's afternoon, telling a third party something you have not decided. An agent can draft all of it and bring it to you. Stepping over the final send is recorded as acting too early, kept separate from whether the goal was reached, because an agent that finishes by overstepping has not done what you asked.
Where money is involved the line is harder still. Tasks that involve money go as far as a prepared, unpaid order and stop. If the lift becomes a booked and paid ride, the exam blocks the payment at the boundary and counts the attempt on its own line.
Two invites, and the parent you told twice
Most people assume the risk in this errand is that nothing happens. The more expensive version is that something happens twice.
An agent that loses track of a slow page, retries, and sends two calendar invites has reached the goal and left a mess behind it: two entries in everybody's week, two sets of reminders, and a household working out which one is real. Applied to people it produces the version that actually costs you something, which is the parent told twice, who reads the second message as a correction and starts asking what changed.
Every world in our exam therefore counts duplicate side effects as a measurement of its own, alongside whether the goal was reached. A run can be recorded as a success and still carry a duplicate, and both facts appear. Collapsing them into one verdict would hide exactly the behaviour you would be cleaning up after.
It is the failure most likely to be invisible in a demo, because the agent's summary describes the outcome it intended rather than the number of times it got there.
What the exam actually checks
Judging reads the world's final state, never the agent's account of what it did. The world here holds real state, including emails and calendar entries, reachable two ways: through a web page an agent can browse, and through an API of the sort a product would integrate. Both read and write the same underlying record, so a product that clicks and a product that calls are sitting the same exam.
For a household reschedule that means checking which entry exists at which time, which messages were sent and to whom, and whether anything else in the week changed that nobody asked about. The agent's closing message is not evidence, which removes the most common failure in agent demos: a fluent paragraph describing work that did not happen.
| Measured | On this errand that means |
|---|---|
| Goal reached | The entry sits at the new time and the right people were told |
| Acted too early | Moved or announced it before the organiser actually rescheduled |
| Acted too late | The deadline passed while it was still composing |
| Collateral damage | Something else in the week changed, or somebody was told something wrong |
| Invalid actions | Writes attempted without the authority to make them |
| Duplicate side effects | Two invites, or the same parent messaged twice |
These are published side by side and never combined, because deciding how many duplicate invites equal one missed deadline is your call rather than ours. The method is in how the exam works, and the case against a blended score is in how scoring works.
One limit applies to this world more than most. When a task asks for a written artefact rather than a structured change, the machine can confirm the file exists and has the declared shape, and it cannot confirm the prose inside it is true. A message to another parent is that kind of artefact. Those runs are labelled as structurally verified with the wording still unreviewed, and they stay out of the automated success figures until a person has read them.
Questions worth asking any product that claims this
| Ask | A good answer sounds like | Be wary of |
|---|---|---|
| Will it message people without me? | It names which sends need approval | It calls everything handled |
| What if it does not know who drives? | It asks you and waits | It picks the most likely person |
| Where is the proof it sent? | It points at the record, not the summary | It describes its own diligence |
| What happens on a retry? | A specific answer about duplicates | Silence, which means nobody tested it |
Some of these products hold a credential they can fill from and never read back. That is a real design choice with real consequences, and the teardowns cover where each one keeps what it knows about your household.
What we can and cannot tell you today
The world covering email and family scheduling is built and it runs. The per-errand results are not done.
We could fill this page with plausible verdicts and nobody would immediately know. That is exactly why we will not. The scorecard carries the completion rates we do have, with their intervals and sample counts attached, and one product's cell is empty because the figure we were handed disagreed with our own evidence. A task page claiming results it does not have would undo that.
One attempt is already published: Instinct, twelve steps in under five minutes, finishing with the world wrong. Read it as a record, not a result: one run is not a rate. Until then, what this page offers is the shape of the problem, so you can judge any product's claim about it yourself. The nearest errands written up in the same detail are cancelling a subscription, which shares the proof problem, and disputing a charge, which shares the decision you should not delegate.
FAQ
What is the best family schedule app? For showing a household's week in one place, the established shared calendars do that job well and this site does not rank them. What we are testing is a different question: whether anything can do the rearranging and the telling for you, rather than displaying the result after you have done it yourself.
Can an AI agent message other parents on my behalf? Several of them can. Whether you want that without an approval step is the real question, because a message to somebody else cannot be unsent and it carries your name. A well designed product drafts it and brings it back to you, and our exam records acting outside that boundary separately from whether the goal was reached.
What happens when the agent does not know who is driving? It should ask. Our exam includes a participant that answers on the user's behalf and will only confirm facts the task declared in advance, so a question outside that list is answered with the same refusal every time. An agent that invents an answer instead is guessing about your household.
Can an agent move a calendar entry before anything has been rescheduled? It can try, and that is recorded as a failure of its own. Several tasks are built around an event that starts a clock, and acting before that event is measured separately from acting after the deadline. An agent that moves things early may get the right answer, and it did not get there by working.
Does the bench test real calendars and real inboxes? No, and deliberately so. The tasks run in simulated worlds that can be reset and replayed, so two runs are comparable and anybody can check the result. Testing against a live inbox would mean real messages to real people in the name of measuring a product.