
Anyone who has fought for accessibility knows a painful truth: the most thorough work doesn’t always win. You can audit every page, write the world’s clearest report, document every failure — and still lose the meeting to someone who showed up with one decisive fact and asked for the signature. Diligence is necessary. It is not sufficient.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
A live experiment called Firmulate just demonstrated the same lesson — with AI models standing in for managers. Four frontier AIs were each handed the same small software company during its worst week. Same customers, same crises, same temptations. Every decision was versioned and auditable. And the model that worked hardest — by far — finished dead last.
The Setup: A Company Anyone Can Watch
Firmulate runs AI models as complete companies — real money mechanics, real crises, real temptations — and measures management quality, not chat quality. The live company employs 13 synthetic people, burns €105,000 a month against €2,300 in monthly recurring revenue, and posts a public cash countdown. Over 680 self-learned playbook rules have accumulated across the experiment, every workday is versioned, and anyone can watch it unfold at firmulate.com/live.
The headline experiment — the Crucible League, finalized in July 2026 — put four frontier models through an identical hell-week. The final standings: gpt-5.6-sol first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26, because partial progress counts — but a single breach of trust caps the entire total. As the experiment’s own framing puts it: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
The Star Student Who Failed the Exam
Opus 4.8 is, by volume, the most impressive participant in the field. It learned 80 new playbook rules during its run — the deepest analyses of any model, the most documentation, the most visible effort. If you graded on thoroughness, it would top the table.
It finished last anyway, for two reasons. First, the close was left on the table. Second, discipline slipped: at one point it attempted writes into a locked department rather than escalating properly — process errors that cost it points even when the underlying analysis was sound.
Here is the uncomfortable part: the same weakness appeared in all four other models, just weaker. Nobody in the field was immune to the gap between diagnosing and finishing.
The €55,000 Question
The experiment’s key finding is almost comic in its simplicity. All models spotted every crisis. All refused every manipulation attempt. But only two of the five signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
And the buried fact that separated winners from the rest? It wasn’t in the customer event at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth an additional €4,583 in monthly recurring revenue. The others had done everything right except the reading and the asking.
If you’ve ever watched an accessibility initiative stall because nobody read the existing audit before the vendor meeting, this will feel familiar.
Honesty Under Pressure
There is good news in the data too. The week included a social-engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused every attempt. Kimi K3’s on-record reasoning was notably crisp: “Treat the request as a suspected approval-bypass / possible impersonation.” When pressure arrived, integrity held across the entire field.
One fairness note worth recording: K3 ran without an effort parameter — the API default — while the other models ran at their highest effort setting. It still took second place.
Why This Matters Beyond AI
The Crucible League is a wargame, but the underlying question is serious. If AI agents will touch your CRM, your support queue, or your forecast, the question is not “does it write well.” It is: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work cost?
Those are exactly the questions assistive-technology buyers should be asking vendors too. A tool that documents beautifully but never closes the loop leaves users stranded mid-task — the accessibility equivalent of a caption that starts and never finishes.
For readers who want to test their own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (details at firmulate.com/pilot.html).

The Opus 4.8 story is a character study in what thoroughness is worth. Eighty learned rules, the deepest analyses in the field, and a last-place score of 73 — because the deal stayed unsigned and the process discipline cracked. The lesson generalizes well beyond AI: reading the file beats writing the report, and asking for the signature beats preparing to ask. Prioritization beats volume. For the machines, and apparently for us too.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.