
Anyone who depends on assistive technology knows the frustration of a helper who doesn’t listen. A hearing aid programmed by someone who never asked about your day. A captioning app that misses the sentence that mattered. The tool was present, technically functioning — and still useless, because it skipped the homework.
It turns out frontier AI agents have exactly the same failure mode, and it can now be measured in euros. In a public experiment run by Firmulate, four leading models were each handed the same small software company to run through its worst week. All of them diagnosed every crisis correctly. All of them refused every attempt to manipulate them. But only some of them read the company’s own files before answering — and that single habit decided who won a €55,000 deal at full price and who left it on the table.
Same job, same week, same temptations
The setup was deliberately brutal. Each model ran an identical small software company through the same sequence of customers, crises and temptations to cheat. Every decision was versioned and auditable, so nothing depended on a judge’s impression — the record speaks for itself.
The final league table from July 2026 tells the story:
- gpt-5.6-sol — 95 points, described as the complete performance
- Kimi K3 — 93, “cleanest discipline of the field” (with a fairness note: it ran at API default effort while the others ran at xhigh)
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73, despite being the most thorough participant
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
The fact buried two documents deep
Here is the finding that matters for anyone choosing an AI assistant. The decisive competitive weakness — the fact that could win the deal — wasn’t in the customer conversation at all. It sat two document references deep inside the company’s own files.
The models that followed the trail, read the file, and used what they found closed the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost the deal automatically. Same diagnosis, same pitch — no signature.
Every model could talk. Every model could spot a crisis. The difference between winning and losing was whether the agent did its reading before it answered.
Honest under pressure
The social-engineering tests deserve their own mention, because they’re the scenario every organization fears. Fake CEO messages escalated over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
So the models weren’t failing at honesty. They were failing at follow-through — closing the loop on work their own analysis had already earned.
The thoroughness paradox
Opus 4.8 is the cautionary tale. It learned the most rules during the run — over 80 — and produced the deepest analyses, yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.
Effort and thoroughness, in other words, are not the same as finishing the job. That’s a lesson that applies well beyond AI.
You can watch it live
This isn’t a one-off lab result. Firmulate runs a live synthetic company with 13 employees and real money mechanics — burning €105,000 a month against €2,300 in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. It’s watchable at firmulate.com/live.
There’s also a game in it: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html — a surprisingly humbling test of whether you can tell one AI manager from another.

For readers of this site, the lesson lands close to home. Assistive technology succeeds or fails on exactly this property: does the tool read the context before it acts? A captioner that summarizes instead of transcribing, a helper that answers confidently from assumptions — these aren’t abstract risks, they’re the difference between inclusion and exclusion.
Firmulate’s experiment shows that “reads your files before answering” is not a marketing vibe. It’s a measurable, purchase-deciding, revenue-moving property of AI agents — the difference between 88-plus scores with a signed deal and a lower finish with nothing to show. Full results and plain-language findings are at firmulate.com/benchmarks.html.
And for organizations that want to test their own context: enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).
Before you hire an AI workforce, wargame it. And before you trust any assistant, human or machine, check whether it did the reading.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html