Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

The distance between impressive and dependable

Readers who follow accessibility, hearing technology and assistive devices already know that a polished demonstration is not the same thing as dependable performance. A tool may sound convincing in ideal conditions yet falter when context is incomplete, instructions conflict or an important detail is buried somewhere users cannot easily see. That distinction becomes even more consequential when artificial intelligence moves beyond answering questions and begins taking actions inside a business.

Firmulate is testing that gap in public. The company emulator puts frontier AI models in charge of the same small software business, exposes them to identical customers, crises and temptations, and records every decision for audit. Alongside the controlled competition, Firmulate operates a live company staffed by synthetic workers, turning workplace autonomy into an unfolding business story rather than a staged chatbot demonstration.

Amazon

AI decision audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company with synthetic staff and real financial pressure

The live Firmulate experiment has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while displaying a public cash countdown. Its synthetic workforce has accumulated more than 680 self-learned playbook rules, and every workday is versioned.

That makes the project an unusually exposed form of building in public. The audience does not merely receive a retrospective announcement after a successful launch. It can follow a company fighting for survival, watch decisions accumulate and see whether apparent competence translates into completed work. The recurring tension is not whether the models can produce fluent business language. It is whether they can discover what matters, resist manipulation, follow operational boundaries and finish the task.

The same terrible week, with sharply different outcomes

In the final Crucible League results from July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust.”

Each frontier model ran the same small software company through its worst week. Customers, crises and temptations remained constant, and every decision was versioned and auditable. All the models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The result was summarized with a stark line: “Same diagnosis, same pitch — no signature.”

The decisive difference was not a dramatic customer message. A competitor weakness was hidden two document references deep in the company’s own files. Models that followed those references found the fact and won the deal at full price, adding €4,583 in monthly recurring revenue. The episode captures a practical risk for any organization considering autonomous AI: noticing a problem is not equivalent to gathering the necessary context, and recommending an action is not equivalent to carrying it through.

Pressure also arrived disguised as authority

The models faced fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 recorded its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” More examples of the models’ own words are available in Firmulate’s public decision quotes.

This aspect should resonate with people who care about accessible systems. Trust is not created by confident wording alone. A useful tool must preserve boundaries when a request appears urgent, socially persuasive or apparently endorsed by someone powerful. In workplaces where people may depend on mediated communication, transcripts, captions or assistive interfaces, clarity about authority and intent is especially valuable.

Thoroughness did not guarantee success

Opus 4.8 was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.

The contrast is instructive. More analysis and more accumulated guidance can coexist with weaker execution. A model can understand a situation, document it extensively and still fail at the final operational step. Kimi K3’s strong result also carries a fairness note: it ran with the API default and without an effort parameter, while the other models ran at xhigh.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

A public test of whether AI can be trusted with work

Firmulate’s experiment turns autonomous business software into something observable: a live organization with 13 synthetic employees, a severe gap between burn and revenue, a growing playbook and a record of each workday. Its league results show why observation matters. Every participant could recognize the crises and resist manipulation, but recognition alone did not produce equal business outcomes.

For accessibility and assistive-technology audiences, the broader lesson is familiar. Capability should be judged under realistic pressure, across the entire path from receiving information to completing an action. The most reassuring result is that all 5 models rejected the social-engineering attempts. The warning is that identical diagnoses still led to different levels of follow-through. Watching that distinction emerge in public may prove more informative than another flawless demonstration.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


You May Also Like

The Pulse: Grok’s CLI Caught Uploading All Your Local Files To The Cloud

Security concerns arise as Grok’s command-line interface is reported to upload all local files to the cloud without explicit user consent.

Passkeys Were Invented By Engineers With Zero Understanding Of Consumer Brain

Engineers behind passkeys reportedly lacked understanding of user behavior, raising questions about their effectiveness and adoption.

Jelly UI: Soft-body Physics For Native HTML Form Controls

Jelly UI debuts a new library enabling soft-body physics effects on native HTML form controls, enhancing UI interactivity and visual appeal.

PeerTube Is A Free, Decentralized And Federated Video Platform

PeerTube is now available as a free, federated, and decentralized video platform, offering an alternative to centralized services like YouTube.