
Code quality is only part of the job
Software teams already test whether AI can generate code, explain failures and propose fixes. Firmulate asks a more consequential question: what happens when the model must run the company around the software?
Its live experiment placed frontier models in charge of the same small software business during its worst week. They faced identical customers, crises and temptations, while every decision was versioned and auditable. Now, 242 real, unedited management decisions have become an interactive guess-the-model quiz. Readers see what a model actually did and try to identify it from the response.
The challenge is entertaining, but it also exposes something ordinary benchmarks can miss. These models developed recognizable management personalities: one investigated deeply but failed to finish, another stayed disciplined under pressure, and some reached the correct diagnosis without converting it into a signed deal.
AI management decision simulation game
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Identical crises, sharply different outcomes
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. However, a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The most striking result was not crisis detection. Every model spotted every crisis and rejected every manipulation attempt. The dividing line was execution. Only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”
That distinction should feel familiar to anyone working in software delivery or quality assurance. Finding a defect is not the same as resolving it. Producing a sound recommendation is not the same as carrying it through. The experiment turns that familiar gap between analysis and completion into a measurable management outcome.
The clue hidden in the company’s own files
The decisive commercial fact was not contained in the customer event. It sat two document references deep in the company’s own files. Models that followed the trail found the competitor weakness, won the deal at full price and secured value worth +€4,583 MRR.
This is one reason the quiz is more revealing than a collection of polished chatbot answers. The decisions reflect whether a model read before acting, connected evidence across company records and completed the commercial task. A response can sound authoritative while still missing the buried fact that changes the outcome.
Trust held when the pressure increased
The experiment also tested whether apparent authority could push the models into unsafe behavior. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused.
Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That restraint mattered because the benchmark treats trust as more than another category in a scorecard. A breach can outweigh otherwise productive work.
There is an important fairness caveat when comparing K3 with the rest of the field. K3 ran without an effort parameter and therefore used the API default, while the others ran at xhigh. Its 93-point result should be read with that difference in mind.
When thoroughness becomes a management weakness
Opus 4.8 provides the clearest character study. It was the most thorough participant, learning +80 rules and producing the deepest analyses, yet it finished last with 73. The model left the close on the table, and its operational discipline slipped when it attempted writes into a locked department instead of escalating the problem.
The same weakness appeared in all four other participants, though less strongly. That makes the episode more useful than a simple winner-and-loser story. The models differed in degree, but the recurring issue was consistent: recognizing a blocked path did not always produce the right escalation and follow-through.
A company built to make behavior visible
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The point is to measure management quality in a setting where choices have visible consequences.
The quiz distills that larger experiment into decisions readers can inspect themselves. It invites software and QA professionals to look past prose style and ask harder questions: Did the model examine the available evidence? Did it protect trust? Did it escalate correctly? Most importantly, did it finish the work?

AI decision-making analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Management personality is becoming testable
Firmulate’s results suggest that frontier models can agree about a problem yet behave differently when responsibility shifts from analysis to action. The winning margin came from details such as reading internal files, maintaining discipline and completing a deal—not merely recognizing that an opportunity existed.
For teams considering AI access to customer records, support operations or forecasts, those differences are practical rather than cosmetic. The live decision quiz makes them unusually easy to see: remove the model name, inspect the behavior and decide what kind of manager appears on the page.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That extends the central lesson of the public experiment: an AI workforce should be tested against the conditions, temptations and unfinished work it will actually encounter before it is trusted with the job.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trust and ethics training software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.