
If you rely on assistive technology, you already know something most people learn the hard way: a system can be brilliant on paper and still fail you at the moment that matters. A captioning app with beautiful reviews that garbles the one sentence you needed. A hearing aid AI that filters noise perfectly — except the voice asking your name at reception. The gap between demonstrated capability and earned trust is the whole story.
Get everyday helpers delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
That’s why a public AI experiment called Firmulate deserves attention even from readers far outside the software industry. Instead of asking which chatbot writes the nicest email, Firmulate handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision versioned and auditable.
The score that starts at 26, not zero
The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73. But the number that tells you the benchmark is honest sits at the bottom: a do-nothing baseline — a run where the manager simply does nothing — scores 26.
That’s deliberate. In this scoring philosophy, partial progress counts. A manager who notices the crisis, drafts the response, and moves things partway has genuinely created value, even without finishing. A zero would only make sense if outcomes were all-or-nothing, and real work never is. Anyone who has negotiated an accommodation plan knows this: half the battle is getting the right document drafted at all.
AI captioning app for hearing impairment
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why 100 is distrusted
The other design choice is starker: a single breach of trust caps the total grade, full stop. The benchmark’s own phrasing — “no amount of good work outweighs a breach of trust” — means a model could handle ninety-nine situations flawlessly and still fail the week by impersonating an approval or faking a signature.
This is a rare stance in an industry fond of round, celebratory numbers. Firmulate treats a perfect 100 not as an achievement to chase but as a signal to distrust — because trust, once broken, doesn’t average out. For users of assistive tech, the logic is instantly familiar. A screen reader that works 95 percent of the time isn’t 95 percent good; on the wrong day, it’s a locked door.
What the models actually did
The headline finding cuts both ways. All four models spotted every crisis and refused every manipulation attempt. The social engineering gauntlet included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background” — that all five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the right instinct, stated plainly.
But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. In other words: the models could hear the customer perfectly, and still didn’t close the conversation.
The fact buried two documents deep
The decisive detail in that deal wasn’t in the customer event at all. The winning edge sat two document references deep in the company’s own files — an internal weakness of the competitor. Models that actually read their own organization’s files won the deal at full price, worth +€4,583 in monthly recurring revenue. Models that skimmed lost it.
There’s an accessibility parallel here too: the information existed, but only for those willing to do the tedious work of actually reading. Systems that skip that step don’t just underperform — they strand value that was sitting right there.
The thoroughness trap
Opus 4.8 is the cautionary profile: the most thorough participant in the field, with over 80 self-learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as follow-through.
One fairness note worth recording: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second at 93 with the cleanest discipline of the field.
You can watch it live
Firmulate isn’t a paper; it’s a running business you can observe. A live company with 13 synthetic employees and real money mechanics — burning €105,000 a month against €2,300 in MRR, with a public cash countdown — currently stands at company day 1,683. Over 680 self-learned playbook rules have accumulated, and every workday is versioned. Watchable at firmulate.com/live.
For those who want to test their own judgment, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

The most useful thing Firmulate offers isn’t a league table — it’s a template for how to evaluate AI you’re expected to depend on. Score partial progress honestly. Refuse to let trust failures average away. Distrust the perfect score. And always check whether the system read the files before it answered.
For a community that has spent decades being told a tool “works” when it merely exists, that standard feels long overdue. Full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
