AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday helpers delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A Different Kind of Accessibility Problem

If you use assistive technology, you know the frustration of tools that pass every demo and fail every real task. The chatbot that answers fluently, then never actually books the appointment. The assistant that reads the email aloud but misses the attachment that mattered. For years, accessibility has been about whether technology works for people. Now a live, public experiment is asking a harder question: does it work at all, when nobody is watching?

Firmulate runs frontier AI models as entire companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. Its latest result is a wake-up call: the best-performing model in a five-way league was a newcomer from Moonshot that many enterprises have never benchmarked.

Amazon

assistive AI chatbot for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Worst Week, Five Frontier Models

The setup is elegantly simple. Each frontier model was given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable, and the company itself is watchable in real time at firmulate.com: 13 synthetic employees, a public cash countdown (burning €105k/month against €2.3k MRR), and more than 680 self-learned playbook rules.

The final July 2026 league table tells the story:

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the €55k deal. The complete performance.
  • 2. Kimi K3 — 93. The newcomer from Moonshot: closed the deal too, with the cleanest discipline in the field.
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77 and 5. Opus 4.8 — 73. Good diagnoses, no signature.

For context, doing nothing scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.

The Needle Nobody Chat-Demoed

Every model spotted every crisis. Every model refused every manipulation attempt, including a fake-CEO impersonation escalating over three stages and a reporter’s “just one yes/no, on background” trick. The league was decided on something subtler: a decisive competitor weakness buried two document references deep in the company’s own files. The models that actually read the file won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t delivered the same diagnosis and the same pitch, and never got the signature.

Kimi K3’s run was remarkable for its discipline: it found the buried fact, closed the deal, saved the churning customer, and resisted all three baits with just one deviation — the cleanest in the field. Its on-record reasoning on the impersonation attempt: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Thoroughness Trap

The most instructive profile is Opus 4.8: the most thorough participant, with 80-plus learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Fluency and diligence, it turns out, are not the same as finishing the job.

Fairness Footnote

One caveat worth stating plainly: K3 ran without an effort parameter (API default) while the other four models ran at xhigh. Even so, its second-place finish makes the larger point — the league is open, and the familiar Western names don’t automatically top it.

Try It Yourself

Firmulate’s benchmarks are public at firmulate.com/benchmarks.html, and a quiz built from 242 real, unedited management decisions lets you guess which model made which call. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The lesson for anyone choosing an AI assistant — for accessibility, for business, or both — is the same: demos measure chat quality, not work quality. A newcomer beat three of four Western frontier models at running a company, and the gap between the best and worst was invisible until someone built the test. Picking a model without your own test is now a bet, not a decision. Firmulate’s experiment is live, watchable, and refreshes twice a day — which is more than can be said for most vendor claims.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Said the Right Thing. Did It Finish the Job?

A live quiz turns 242 unedited AI management decisions into a revealing test of judgment, follow-through, honesty and attention to context.

The Pulse: Grok’s CLI Caught Uploading All Your Local Files To The Cloud

Security concerns arise as Grok’s command-line interface is reported to upload all local files to the cloud without explicit user consent.

European “Age Verification” “App” Forcing Everyone To Use Android Or iOS

A new European age verification app mandates users to access via Android or iOS devices, raising privacy and accessibility concerns.

Show HN: Davit, A Apple Containers UI

Developer shares Davit, a UI for Apple Containers, on Show HN, with source code available for public use. The project aims to simplify Apple Containers interface design.