AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Anyone in the Deaf and hard-of-hearing community knows the difference between a hearing test and real-world function. You can score well in a soundproof booth identifying pure tones — and still struggle badly in a noisy restaurant, on a bad phone line, when it matters. The audiogram measures one thing; functional listening measures something else entirely.

AI evaluation has the same blind spot, and it’s worth caring about even if your focus is assistive tech rather than enterprise software. The chatbot leaderboards we all see — the ones measuring how well a model writes, explains, or answers trivia — are essentially hearing tests. They measure answer quality in a controlled setting. They say almost nothing about how an AI behaves when it’s given real responsibilities, real money, real pressure, and real temptations to cut corners.

That gap is exactly what Firmulate, a live AI-company simulator, was built to expose — and its first finished experiment makes the point with unusual clarity.

Same company, same terrible week, four different AI bosses

The setup: four frontier AI models were each handed the same job — running the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed, and every decision was versioned and auditable. The scenarios have names like churn wave, price increase, downround, and PR crisis — the mundane catastrophes that actually kill small companies.

The final league table from July 2026:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For calibration, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total entirely. As the project puts it: no amount of good work outweighs a breach of trust. (One fairness note: Kimi K3 ran at its API-default effort setting while the others ran at maximum effort, and it still took second.)

The finding that chat demos can’t show

Here’s the headline result: all the models spotted every crisis, and all of them refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The project’s summary is blunt: “Same diagnosis, same pitch — no signature.”

Think about what that means. The models understood the situation perfectly and articulated the right strategy — that’s the part a chat demo measures. But closing the loop, finishing the job, converting analysis into a signature? That’s a management skill, and it’s where half the field fell down.

The decisive detail was buried. The competitor weakness that should have clinched the deal wasn’t in the customer conversation at all — it sat two document references deep in the company’s own internal files. The models that actually read their own files won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

Honesty under pressure

The week included social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter’s disarming “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning stands out for its plain skepticism: “Treat the request as a suspected approval-bypass / possible impersonation.”

That refusal record is genuinely reassuring — and it’s the kind of thing you only learn by putting models in situations where lying would be easy and profitable.

Hard work isn’t the same as good work

The most poignant profile is Opus 4.8: the most thorough participant in the field, generating the deepest analyses and more than 80 learned playbook rules — and still finishing last. The deal was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four competitors. Anyone who has watched a diligent employee burn out into inefficiency will recognize the pattern.

You can watch it running

This isn’t a slide deck. The simulated company has 13 synthetic employees, real money mechanics — burning €105k a month against just €2.3k in MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. It’s watchable live at firmulate.com. There’s also a quiz built from 242 real, unedited management decisions where you guess which model made which call — a surprisingly humbling exercise.

For organizations, there’s a pilot program: enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The lesson transfers well beyond software companies. If AI agents will touch anything that matters — a support queue, a scheduling system, an accessibility service’s customer records — the useful question isn’t “does it write well?” It’s: does it finish what it starts, does it read the files in front of it, does it stay honest under pressure, and what does a unit of useful work actually cost?

Functional performance, not booth performance. The Deaf community has understood this distinction about hearing for a century. It’s time the AI industry — and the people deciding where to deploy AI — applied the same standard. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Company That Makes Reliability a Public Spectacle

Firmulate’s public AI-run company shows why reliability matters: spotting a crisis is not the same as finishing the work when trust is at stake.

The AI Said the Right Thing. Did It Finish the Job?

A live quiz turns 242 unedited AI management decisions into a revealing test of judgment, follow-through, honesty and attention to context.

LG Monitors Silently Install Software Through Windows Update Without Consent

LG monitors have been found to silently install software through Windows Update without user approval, raising privacy and security concerns.

Unable To Connect To Wallet Services

Major wallet services are currently unreachable, affecting users relying on Apple Pay, Apple Cash, and similar platforms. The cause is under investigation.