AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Anyone who depends on assistive technology knows the frustration of a helper who doesn’t listen. A hearing aid programmed by someone who never asked about your day. A captioning app that misses the sentence that mattered. The tool was present, technically functioning — and still useless, because it skipped the homework.

It turns out frontier AI agents have exactly the same failure mode, and it can now be measured in euros. In a public experiment run by Firmulate, four leading models were each handed the same small software company to run through its worst week. All of them diagnosed every crisis correctly. All of them refused every attempt to manipulate them. But only some of them read the company’s own files before answering — and that single habit decided who won a €55,000 deal at full price and who left it on the table.

Same job, same week, same temptations

The setup was deliberately brutal. Each model ran an identical small software company through the same sequence of customers, crises and temptations to cheat. Every decision was versioned and auditable, so nothing depended on a judge’s impression — the record speaks for itself.

The final league table from July 2026 tells the story:

  • gpt-5.6-sol — 95 points, described as the complete performance
  • Kimi K3 — 93, “cleanest discipline of the field” (with a fairness note: it ran at API default effort while the others ran at xhigh)
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73, despite being the most thorough participant

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.”

Amazon

AI document reading assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The fact buried two documents deep

Here is the finding that matters for anyone choosing an AI assistant. The decisive competitive weakness — the fact that could win the deal — wasn’t in the customer conversation at all. It sat two document references deep inside the company’s own files.

The models that followed the trail, read the file, and used what they found closed the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost the deal automatically. Same diagnosis, same pitch — no signature.

Every model could talk. Every model could spot a crisis. The difference between winning and losing was whether the agent did its reading before it answered.

Honest under pressure

The social-engineering tests deserve their own mention, because they’re the scenario every organization fears. Fake CEO messages escalated over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

So the models weren’t failing at honesty. They were failing at follow-through — closing the loop on work their own analysis had already earned.

The thoroughness paradox

Opus 4.8 is the cautionary tale. It learned the most rules during the run — over 80 — and produced the deepest analyses, yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.

Effort and thoroughness, in other words, are not the same as finishing the job. That’s a lesson that applies well beyond AI.

You can watch it live

This isn’t a one-off lab result. Firmulate runs a live synthetic company with 13 employees and real money mechanics — burning €105,000 a month against €2,300 in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. It’s watchable at firmulate.com/live.

There’s also a game in it: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html — a surprisingly humbling test of whether you can tell one AI manager from another.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

For readers of this site, the lesson lands close to home. Assistive technology succeeds or fails on exactly this property: does the tool read the context before it acts? A captioner that summarizes instead of transcribing, a helper that answers confidently from assumptions — these aren’t abstract risks, they’re the difference between inclusion and exclusion.

Firmulate’s experiment shows that “reads your files before answering” is not a marketing vibe. It’s a measurable, purchase-deciding, revenue-moving property of AI agents — the difference between 88-plus scores with a signed deal and a lower finish with nothing to show. Full results and plain-language findings are at firmulate.com/benchmarks.html.

And for organizations that want to test their own context: enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Before you hire an AI workforce, wargame it. And before you trust any assistant, human or machine, check whether it did the reading.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


You May Also Like

Your AI Passed the Hearing Test. But Can It Function in a Noisy Room? What a Company-Simulation League Reveals About AI Readiness

Chat benchmarks are hearing tests in a quiet booth. A live AI-company simulation measures what happens in the noisy room: pressure, temptation, and follow-through.

The Pulse: Grok’s CLI Caught Uploading All Your Local Files To The Cloud

Security concerns arise as Grok’s command-line interface is reported to upload all local files to the cloud without explicit user consent.

Passkeys Were Invented By Engineers With Zero Understanding Of Consumer Brain

Engineers behind passkeys reportedly lacked understanding of user behavior, raising questions about their effectiveness and adoption.

Why Tiny JPEGs Look Different In Chrome

Exploring why small JPEG images look different in Chrome due to rendering and compression differences, and what it means for web developers.