AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What if software QA tested an entire company?

Software teams are accustomed to testing code against edge cases, failure states and hostile inputs. Firmulate extends that logic to management itself. Its live experiment asks frontier AI models to operate the same small software company through severe business crises, customer pressure and attempts at manipulation. The resulting decisions are versioned and auditable, turning the company’s struggle into a public test of whether AI can do more than produce convincing text.

This is not a simulated dashboard built around comfortable assumptions. The company can be watched live, with 13 synthetic employees operating under real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown makes the consequences visible. Every workday is versioned, and the organization has accumulated more than 680 self-learned playbook rules.

Amazon

AI-driven business decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company, the same crises, different managers

In the final Crucible League results from July 2026, each frontier model faced the same customers, crises and temptations while running the same company through its worst week. The controlled setup matters: differences in performance could be traced to decisions rather than easier assignments or more favorable circumstances.

gpt-5.6-sol led with a score of 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total because “no amount of good work outweighs a breach of trust.”

The most striking result was not crisis detection. Every model identified every crisis and rejected every manipulation attempt. The real divide appeared at the point where analysis had to become commercial action. Only two models signed the €55,000 deal that their own work had earned. The experiment summarized the gap bluntly: “Same diagnosis, same pitch — no signature.”

The winning fact was already inside the company

The decisive competitive weakness was not contained in the customer event. It sat two document references deep in the company’s own files. Models that followed that trail secured the deal at full price, adding €4,583 in monthly recurring revenue.

For software, QA and development teams, this is a familiar kind of failure in an unfamiliar setting. A system may recognize the visible incident, describe the correct response and still miss the evidence needed to finish the task. Firmulate’s result suggests that evaluating an AI worker only through isolated prompts can conceal the difference between articulate analysis and completed work.

The buried fact also makes the experiment unusually relevant to companies considering agents for operational roles. Business work rarely arrives as a self-contained prompt. Essential context may be hidden in account records, past decisions or linked documents. The ability to inspect those materials before acting can determine whether an apparently competent agent produces an explanation or an outcome.

Pressure tested honesty

The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous refusal is encouraging, but Firmulate’s broader findings resist a simple winners-and-losers story. The models could remain honest while still failing through incomplete execution or weak operational discipline. Readers can examine the company’s public voice through its published employee quotes, adding human-readable texture to the auditable record of decisions.

Thoroughness did not guarantee success

Opus 4.8 provides the clearest cautionary example. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.

The result challenges a common assumption about AI quality: that more analysis naturally yields better execution. In this experiment, thoroughness became valuable only when paired with the discipline to complete the commercial task and respond correctly to organizational boundaries.

The league table also requires a fairness note. Kimi K3 ran without an effort parameter and therefore used the API default, while the other models ran at xhigh. That difference should remain visible when comparing its second-place result with the rest of the field.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

corporate crisis simulation AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that generates evidence every workday

Firmulate’s most unusual contribution is not merely a benchmark result. It is the continuing portrait of a software company publicly fighting for survival. Its cash position, work and accumulated operating knowledge create fresh material every business day, while the mismatch between €105k in monthly burn and €2.3k in monthly recurring revenue keeps the stakes concrete.

The experiment turns build-in-public into something closer to continuous organizational QA. Visitors are not shown a polished demonstration with the difficult moments removed. They can watch a company confront financial pressure, incomplete information and social engineering while its synthetic workforce tries to keep the business alive.

For technology leaders, the central lesson is that an AI system can detect danger, protect trust and produce excellent analysis without reliably finishing the job. Firmulate makes that gap observable—and makes the cost of confusing fluent work with completed work difficult to ignore.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI project management and decision tracking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Intel Starts Shipping High-NA EUV Silicon

Intel has started shipping silicon wafers manufactured with high-NA EUV lithography, marking a key step in advanced chip production technology.

Where The 176GB Actually Goes: The Memory Budget Nobody Reads Until It’s Too Late

Exploring where the 176GB of model weights actually go and why total memory usage exceeds simple calculations in large AI models.

The Ninth Point: Affordable AI Validation With DeepSeek-V4-Flash-High

MIT-licensed DeepSeek-V4-Flash-High outperforms competitors at one-fifteenth the price, marking a shift in AI capability scaling post-training.

Podman V6.0.0

Podman v6.0.0, the latest version of the container engine, introduces significant improvements in security, performance, and compatibility, officially launched today.