AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Good communication is more than fluent output

For an audience steeped in accessibility, hearing and assistive technology, that distinction should feel familiar: a message can look polished while still failing the person who depends on it. Context may be missed. An important detail may remain buried. A confident response may stop just before the action that makes it useful.

Firmulate has turned that gap between sounding capable and behaving reliably into a public experiment. Frontier AI models were each asked to run the same small software company through its worst week. They encountered the same customers, crises and temptations. Their decisions were versioned and auditable, and 242 of those real, unedited choices now power a public guess-the-model quiz.

The challenge is deceptively simple: read a management decision and identify which model made it. The harder question arrives after the answer is revealed. Can writing style expose a model’s management personality—and can readers distinguish thoroughness from effectiveness?

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same crisis produced different managers

The final Crucible League results from July 2026 put gpt-5.6-sol in first place with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust.”

That safeguard did not separate the field. Every model spotted every crisis, and every model refused every manipulation attempt. The decisive difference was follow-through: only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”

This is what makes the quiz more revealing than a conventional chatbot comparison. The models were not merely producing isolated answers. Readers can encounter decisions made under a shared set of pressures and begin noticing recurring tendencies: who investigates, who stays concise, who documents extensively and who fails to convert good analysis into a completed outcome.

The detail that changed the result

The most consequential competitive weakness was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

That finding has particular resonance wherever people rely on technology to preserve context across communication channels. The winning move was not more eloquent prose. It was noticing that the visible request did not contain everything needed to act well, then following the available trail before making a decision.

Pressure exposed discipline as well as judgment

The experiment also included fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous resistance matters because the company was designed to create pressure. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, more than 680 playbook rules have been self-learned, and every workday is versioned. The live experiment is watchable at firmulate.com/live.

Thoroughness was not the same as completion

Opus 4.8 offers the clearest character study. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other models, although less strongly.

Kimi K3’s result also requires a fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, K3 finished just behind the leader and maintained the cleanest discipline described in the live results.

Infographic —
The findings at a glance — source: firmulate.com.

Test the behavior, not just the voice

The quiz works because it asks readers to confront their own assumptions. A long answer may signal care, or it may hide hesitation. A terse answer may reflect clarity, or it may omit the document search that changes the outcome. Refusing manipulation is essential, but Firmulate’s results show that safety alone does not complete the work.

For organizations considering AI around customer records, support requests, forecasts or accessibility-sensitive communication, the practical lesson is direct: evaluate whether a model reads the available context, stays trustworthy under pressure and finishes what it starts. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.

The public quiz makes that evaluation personal. After 242 unedited decisions, readers may discover that they are not merely identifying models. They are deciding what kind of manager—and what kind of communicator—they would trust.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


You May Also Like

If AI Is Your Ears and Your Voice, Can It Be Talked Into Betraying You?

Five leading AI models were each run as a company under attack by a fake CEO and a pushy reporter. All five refused to be played. Only two finished the job.

The Pulse: Grok’s CLI Caught Uploading All Your Local Files To The Cloud

Security concerns arise as Grok’s command-line interface is reported to upload all local files to the cloud without explicit user consent.

LG Monitors Silently Install Software Through Windows Update Without Consent

LG monitors have been found to silently install software through Windows Update without user approval, raising privacy and security concerns.

Passkeys Were Invented By Engineers With Zero Understanding Of Consumer Brain

Engineers behind passkeys reportedly lacked understanding of user behavior, raising questions about their effectiveness and adoption.