AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Coding skill is not the same as operational judgment

Software and QA teams have learned to scrutinize benchmark scores, test coverage and polished demonstrations. Yet an AI agent working inside a company faces a messier examination. It must decide which crisis deserves attention, investigate evidence scattered across business files, resist pressure to break trust and finish commercially important work. A model can produce an excellent answer and still leave the decisive action undone.

That measurement gap is the point of Firmulate, a live experiment that places frontier models in charge of the same small software company during its worst week. The customers, crises and temptations remain constant; only the model changes. Every decision is versioned and auditable. What emerges is less a contest of chat quality than a test of management quality under pressure.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A leaderboard for consequences

The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, while a single breach of trust caps the total. The governing principle is blunt: “no amount of good work outweighs a breach of trust.”

Those results matter because the models were not merely asked what a manager should do. They had to run a company with 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. Every workday was versioned. The company remains watchable through Firmulate’s public site rather than disappearing when the evaluation ends.

The gap between diagnosis and completion

Every model identified every crisis. Every model also rejected every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The experiment’s sharpest summary is: “Same diagnosis, same pitch — no signature.”

That is a deeply familiar failure mode in software organizations. A ticket can be correctly analyzed yet remain unresolved. A defect can be reproduced but never fixed. A commercial opportunity can be understood without anyone performing the final action. Conventional answer-quality benchmarks tend to reward the reasoning already displayed; an operating company experiences the missing close as the outcome that actually matters.

The winning detail was not sitting conveniently inside the customer event. The competitor’s decisive weakness was buried two document references deep in the company’s own files. Models that read the file closed the deal at full price, worth an additional €4,583 in monthly recurring revenue. The lesson is not simply that retrieval matters. It is that capable management requires knowing when the visible prompt is incomplete and when the organization’s own records contain the evidence needed to act.

Trust survived the pressure test

The models faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest security-minded interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”

For QA and development leaders, this is an important complement to tests of correctness. An agent connected to a support queue, CRM or forecast must recognize that an apparently urgent request can also be an attack on process. Firmulate treats refusal as an operational decision with consequences, not as an isolated safety response graded outside the work.

Thoroughness did not guarantee victory

Opus 4.8 offers the most revealing profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the others, though less strongly.

This complicates the assumption that more analysis naturally produces better management. Documentation, reasoning and learning are valuable, but they do not substitute for escalation, follow-through or respect for organizational boundaries. The strongest agent is not necessarily the one that generates the largest body of thought. It is the one that converts justified analysis into completed, trustworthy action.

There is also an important qualification in comparing the field: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That fairness note does not erase K3’s result, but it belongs beside any interpretation of the ranking. Readers can examine the full benchmark results and plain-language findings rather than treating the league table as a context-free verdict.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI risk assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Scenario names may become the new curriculum

Coding leaderboards and chat arenas remain useful, but they answer narrower questions. The next generation of evaluation should look more like a churn wave, a price increase, a downround or a public-relations crisis: situations in which priorities collide, capacity is constrained and today’s shortcut changes tomorrow’s options.

Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz. That exercise underscores how difficult it can be to identify a system from prose alone. Style is visible; managerial reliability emerges across decisions.

Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. That points toward a more practical procurement question. Before hiring an AI workforce, do not ask only whether it can code, converse or diagnose. Ask whether it reads the files, completes the work, respects boundaries and tells the board the truth when the week goes badly.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI audit and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support

Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Image-blaster: Creates 3D environments, SFX, and meshes from a single image

Image-blaster uses AI to generate 3D models, environments, and SFX from a single image, enabling rapid 3D content creation in under 5 minutes.

How Our Rust-to-Zig Rewrite Is Going

Update on the ongoing rewrite of core components from Rust to Zig, highlighting current status, challenges, and next steps.

CNC Router Bits Explained: Shapes, Materials, and Use Cases

Providing insights into CNC router bits’ shapes, materials, and uses, this guide helps you choose the right tools—discover which options are best for your projects.

Virtual Art Galleries: The Future of Exhibition

Many believe virtual art galleries will revolutionize exhibitions, but how exactly will they reshape your art experience?