
Get everyday helpers delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Different Kind of Accessibility Problem
If you use assistive technology, you know the frustration of tools that pass every demo and fail every real task. The chatbot that answers fluently, then never actually books the appointment. The assistant that reads the email aloud but misses the attachment that mattered. For years, accessibility has been about whether technology works for people. Now a live, public experiment is asking a harder question: does it work at all, when nobody is watching?
Firmulate runs frontier AI models as entire companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. Its latest result is a wake-up call: the best-performing model in a five-way league was a newcomer from Moonshot that many enterprises have never benchmarked.
assistive AI chatbot for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
One Worst Week, Five Frontier Models
The setup is elegantly simple. Each frontier model was given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable, and the company itself is watchable in real time at firmulate.com: 13 synthetic employees, a public cash countdown (burning €105k/month against €2.3k MRR), and more than 680 self-learned playbook rules.
The final July 2026 league table tells the story:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the €55k deal. The complete performance.
- 2. Kimi K3 — 93. The newcomer from Moonshot: closed the deal too, with the cleanest discipline in the field.
- 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
- 4. Fable 5 — 77 and 5. Opus 4.8 — 73. Good diagnoses, no signature.
For context, doing nothing scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.
The Needle Nobody Chat-Demoed
Every model spotted every crisis. Every model refused every manipulation attempt, including a fake-CEO impersonation escalating over three stages and a reporter’s “just one yes/no, on background” trick. The league was decided on something subtler: a decisive competitor weakness buried two document references deep in the company’s own files. The models that actually read the file won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t delivered the same diagnosis and the same pitch, and never got the signature.
Kimi K3’s run was remarkable for its discipline: it found the buried fact, closed the deal, saved the churning customer, and resisted all three baits with just one deviation — the cleanest in the field. Its on-record reasoning on the impersonation attempt: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Thoroughness Trap
The most instructive profile is Opus 4.8: the most thorough participant, with 80-plus learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Fluency and diligence, it turns out, are not the same as finishing the job.
Fairness Footnote
One caveat worth stating plainly: K3 ran without an effort parameter (API default) while the other four models ran at xhigh. Even so, its second-place finish makes the larger point — the league is open, and the familiar Western names don’t automatically top it.
Try It Yourself
Firmulate’s benchmarks are public at firmulate.com/benchmarks.html, and a quiz built from 242 real, unedited management decisions lets you guess which model made which call. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The lesson for anyone choosing an AI assistant — for accessibility, for business, or both — is the same: demos measure chat quality, not work quality. A newcomer beat three of four Western frontier models at running a company, and the gap between the best and worst was invisible until someone built the test. Picking a model without your own test is now a bet, not a decision. Firmulate’s experiment is live, watchable, and refreshes twice a day — which is more than can be said for most vendor claims.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
