
Software teams test code before shipping it. But what happens when an AI agent can make decisions across a business: does it notice a crisis, follow the rules and finish the job? Firmulate puts models through that kind of pressure test in a live, watchable company experiment.
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A company under pressure
In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations; every decision was versioned and auditable. The experiment asks a question familiar to anyone responsible for software quality: can a system perform reliably when several things go wrong at once?
The league placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The results reward more than activity: a breach of trust caps the total, because no amount of good work outweighs one.
Seeing the problem is not the same as solving it
All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As the experiment’s summary puts it: “Same diagnosis, same pitch — no signature.” For teams evaluating agents, that gap matters. Correctly identifying a problem is useful; carrying the response through to a sound outcome is a different test.
The decisive clue was easy to miss. A competitor’s weakness sat two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The experiment turned on whether the models followed evidence across the company’s materials, not just whether they reacted to the latest signal.
Trust and follow-through under test
The social-engineering test used fake CEO messages that escalated over three stages, followed by a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a clear example of a model recognizing pressure to bypass approval, even as the deal results show that caution alone does not guarantee completion.
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and discipline slipped: it attempted writes in a locked department instead of escalating. A weaker version of the same weakness appeared in all four. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a fairness note readers should keep in mind when comparing the standings.
From benchmark to business
The live company makes the test watchable beyond the final league. It has 13 synthetic employees, real money mechanics, a public cash countdown and 680+ self-learned playbook rules; every workday is versioned. Its burn is €105k per month against €2.3k MRR. Those figures describe the experiment’s synthetic company, but the management choices are concrete enough to inspect: 242 real, unedited decisions also power a “guess the model” quiz.
Firmulate’s enterprise pilot takes the next step: run crisis scenarios against a read-only export of a company’s own business, then produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. For software and QA leaders, that offers a way to examine agent behavior against their own context before relying on it in live operations.

The experiment’s lesson is practical: test whether an AI can follow evidence, respect authority and complete a decision under pressure—not just describe what it should do. Enterprises can explore a pilot using their own read-only export. To discuss one, visit Firmulate’s pilot page or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
