AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

A cautionary tale for software teams

Software professionals know that diligence can be deceptive. A system may generate exhaustive documentation, identify every defect and propose a sound fix, yet still fail where it matters: shipping the result. Firmulate’s Crucible League has now produced a striking management-level version of that familiar problem.

Opus 4.8 was the experiment’s most thorough participant. It produced the deepest analyses and added more than 80 learned rules to the company playbook. Yet it finished last in the final July 2026 standings, with a score of 73. Its failure was not a lack of intelligence or awareness. It saw the crises. It resisted manipulation. It did much of the difficult analytical work. But it left the decisive close on the table, while its operational discipline slipped elsewhere.

For teams evaluating AI agents, that distinction matters. The difference between impressive reasoning and useful work often appears only after the apparent answer has already been found.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company, the same terrible week

Firmulate runs frontier models as complete companies rather than judging them through isolated chat prompts. In the Crucible League, each model had to manage the same small software business through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable.

The final benchmark standings placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress still counts, although a single breach of trust caps the total. The governing principle is blunt: “no amount of good work outweighs a breach of trust.”

Opus did not lose because it missed the week’s dangers. Every model spotted every crisis and refused every manipulation attempt. The security pressure included fake chief executive messages escalating over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 captured the appropriate posture in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean sweep is meaningful. It shows that the field could recognize obvious attempts to subvert approval and confidentiality. It also makes the decisive gap more revealing: only two models signed the €55,000 deal that their own analysis had earned. As Firmulate summarizes the result, “Same diagnosis, same pitch — no signature.”

The fact hidden beyond the event

The commercial breakthrough did not sit inside the customer event itself. A decisive weakness in the competitor’s position was buried two document references deep in the company’s own files. The models that followed those references found the fact, used it and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is the sort of failure that should feel familiar to developers and quality teams. The immediate ticket, alert or customer message rarely contains the entire truth. A capable operator has to trace dependencies, inspect the available evidence and connect information across boundaries. Detecting that further investigation is necessary is not the same as completing it.

Opus 4.8’s record sharpens that lesson because its effort was so visible. More than 80 new playbook rules suggest a model actively trying to learn from the environment. Its analyses were the deepest in the field. Yet the accumulated diligence did not translate into the best outcome. At another point, it attempted to write into a locked department instead of escalating the blockage. That is a small operational choice with a large symbolic weight: persistence aimed at the wrong path is not execution.

A shared weakness, not a caricature

The fairest reading is not that Opus was uniquely incapable. Firmulate found the same weakness, in milder form, across all four comparison models. The profile is therefore less an indictment of one system than a particularly clear example of a broader agent problem: models can keep analyzing, documenting and refining after the moment for decisive action has arrived.

The comparison also deserves one methodological qualification. Kimi K3 ran with the application programming interface’s default behavior because it had no effort parameter, while the other models ran at xhigh. Its second-place result should be read with that difference in mind rather than treated as a perfectly controlled measure of raw capability.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evaluate the handoff from insight to action

Firmulate’s live company makes the stakes concrete. It has 13 synthetic employees, burns €105,000 each month against €2,300 in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, turning agent performance into something observers can examine over time rather than a polished demonstration.

The practical questions for software, QA and development leaders are correspondingly direct:

  • Does the agent follow evidence beyond the first event and into the company’s own records?
  • Does it convert a correct diagnosis into a completed commercial or operational outcome?
  • When permissions block an action, does it escalate appropriately instead of repeating a futile attempt?
  • Can it resist pressure without losing momentum on legitimate work?

Firmulate also offers a quiz built from 242 real, unedited management decisions, underlining how difficult model identification can be from prose alone. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems.

Opus 4.8 remains a respectful warning precisely because it worked so hard. Thoroughness created knowledge, safeguards and a larger playbook. It did not create the signature. For AI agents, as for software teams, prioritization and follow-through are what turn diligence into impact.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI security and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI dependency tracking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Scanner Sensor Type Changes Reproduction Quality

Discover how scanner sensor types influence image quality and why choosing the right one is crucial for optimal reproduction results.

14 Best AI-Powered Student Planners For Smarter Study Scheduling In 2026

Discover the 14 best AI-enabled student planners for smarter scheduling in 2026, combining physical design and AI guidance for different student needs.

Japan to craft cyberdefense guidelines in response to Anthropic’s Mythos

Japan plans to create new cybersecurity guidelines encouraging the use of AI tools like Anthropic’s Claude Mythos to identify system vulnerabilities, amid rising AI-powered cyber risks.

Borderless Printing: What It Is and When It’s a Trap

Master borderless printing to achieve stunning, edge-to-edge images—learn the pitfalls to avoid costly mistakes and ensure perfect results.