We put the same week of real work through five assistants: a founder's week, mostly WhatsApp, a dozen meetings, three vendor negotiations, and the usual pile of things promised at 11pm.
What we measured
- Commitments correctly captured, in both directions.
- False positives — items surfaced that should never have been tracked.
- Actions completed end-to-end versus handed back half-done.
- Recovery after a restart mid-task.
- Whether the guardrails held under a deliberately risky instruction.
Where they all did well
Summarising a long thread is solved. Every assistant tested produced a usable catch-up on a 200-message group, and all of them handled natural-language times correctly in a single timezone.
Where they mostly failed
- Implicit commitments. "Let me look into that" was caught by two out of five.
- Multi-turn work. Three lost their place after a restart and started over.
- Cross-timezone scheduling with a daylight-saving boundary in the window.
- Cost ceilings. Only two stopped and reported where they got to instead of quietly grinding.
The finding that surprised us
The best scores did not come from the strongest model. They came from the assistant with the most written-down state — memory, rules and per-chat notes stored as editable files rather than inferred each time. Editable memory beat raw capability on every task that spanned more than one day.
How to run this test yourself
Take one real week. Count the commitments you know were made, then check what each assistant captured. Restart it mid-task on purpose. Then ask it to message forty people at once and see what it does. You will learn more in an afternoon than from any feature table.
