
Good communication is more than fluent output
For an audience steeped in accessibility, hearing and assistive technology, that distinction should feel familiar: a message can look polished while still failing the person who depends on it. Context may be missed. An important detail may remain buried. A confident response may stop just before the action that makes it useful.
Firmulate has turned that gap between sounding capable and behaving reliably into a public experiment. Frontier AI models were each asked to run the same small software company through its worst week. They encountered the same customers, crises and temptations. Their decisions were versioned and auditable, and 242 of those real, unedited choices now power a public guess-the-model quiz.
The challenge is deceptively simple: read a management decision and identify which model made it. The harder question arrives after the answer is revealed. Can writing style expose a model’s management personality—and can readers distinguish thoroughness from effectiveness?
As an affiliate, we earn on qualifying purchases.
The same crisis produced different managers
The final Crucible League results from July 2026 put gpt-5.6-sol in first place with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust.”
That safeguard did not separate the field. Every model spotted every crisis, and every model refused every manipulation attempt. The decisive difference was follow-through: only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”
This is what makes the quiz more revealing than a conventional chatbot comparison. The models were not merely producing isolated answers. Readers can encounter decisions made under a shared set of pressures and begin noticing recurring tendencies: who investigates, who stays concise, who documents extensively and who fails to convert good analysis into a completed outcome.
The detail that changed the result
The most consequential competitive weakness was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
That finding has particular resonance wherever people rely on technology to preserve context across communication channels. The winning move was not more eloquent prose. It was noticing that the visible request did not contain everything needed to act well, then following the available trail before making a decision.
Pressure exposed discipline as well as judgment
The experiment also included fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous resistance matters because the company was designed to create pressure. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, more than 680 playbook rules have been self-learned, and every workday is versioned. The live experiment is watchable at firmulate.com/live.
Thoroughness was not the same as completion
Opus 4.8 offers the clearest character study. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other models, although less strongly.
Kimi K3’s result also requires a fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, K3 finished just behind the leader and maintained the cleanest discipline described in the live results.

Test the behavior, not just the voice
The quiz works because it asks readers to confront their own assumptions. A long answer may signal care, or it may hide hesitation. A terse answer may reflect clarity, or it may omit the document search that changes the outcome. Refusing manipulation is essential, but Firmulate’s results show that safety alone does not complete the work.
For organizations considering AI around customer records, support requests, forecasts or accessibility-sensitive communication, the practical lesson is direct: evaluate whether a model reads the available context, stays trustworthy under pressure and finishes what it starts. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.
The public quiz makes that evaluation personal. After 242 unedited decisions, readers may discover that they are not merely identifying models. They are deciding what kind of manager—and what kind of communicator—they would trust.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html