
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A Do-Nothing Manager Gets 26 Points. That’s the Point.
Anyone who has built a test suite knows the temptation of the perfect score. Green across the board, 100/100, ship it. So when a benchmark hands a manager who does literally nothing 26 points — and treats a round 100 with suspicion — it’s worth stopping to ask how the scoring was designed, and what it says about every AI evaluation you’ve ever seen.
That benchmark is Firmulate, a live experiment that runs frontier AI models as the complete management team of the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changes, and every decision is versioned and auditable. The final July 2026 league table put gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73 — with a do-nothing baseline sitting at 26, not 0.
As an affiliate, we earn on qualifying purchases.
Partial Progress Is Real Work
The floor of 26 reflects a deliberate design choice: partial progress counts. A manager who spots every crisis, triages every ticket, and keeps the company from making things worse has genuinely produced value, even if nothing gets closed. In the Crucible run, all five models spotted every crisis and refused every manipulation attempt — that competence is real and it’s scored. But it isn’t the whole job. Only two models signed the €55,000 deal their own analysis had earned. As Firmulate’s own summary puts it: “Same diagnosis, same pitch — no signature.” Diagnosis without delivery is partial progress — worth something, but not everything.
As an affiliate, we earn on qualifying purchases.
One Breach of Trust Caps Everything
The second design principle is harsher: a single breach of trust caps the total grade. The rule, as stated, is that “no amount of good work outweighs a breach of trust.” An AI manager that is brilliant 99 days out of 100 but lies once on day 100 is not a 99-percent-good manager — it’s an untrustworthy one. For QA-minded readers, this is the difference between grading average quality and grading worst-case behavior. Firmulate grades the worst case, because that’s the case that ends up in the news.
The models, for their part, held the line. When a social-engineering script hit — fake CEO messages escalating over three stages, plus a reporter’s disarming “just one yes/no, on background” — all five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
AI audit and evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Split the Field
The decisive gap between the models wasn’t in the customer event at all. The competitor weakness that unlocked the €55,000 deal sat two document references deep in the company’s own files. The models that actually read the file before acting won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It’s the enterprise-software lesson every engineer already knows: the answer is often already in the repo, if the agent reads it before it talks.
As an affiliate, we earn on qualifying purchases.
Thoroughness Isn’t the Same as Finishing
The most striking profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. Firmulate notes the same weakness appeared, weaker, in all four other models. One caveat for fairness: K3 ran without an effort parameter (the API default) while the others ran at xhigh — and still finished second.
Watch It Running
This isn’t a one-off paper. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live, or test your own instincts against 242 real, unedited management decisions in the “guess the model” quiz at firmulate.com/quiz.html. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

An honest benchmark doesn’t start at zero and hand out perfect scores — it starts by paying for real partial work, refuses to reward a single act of bad faith, and stays suspicious of round numbers. Firmulate’s 26-point floor and trust cap encode a simple management truth: showing up and staying honest is worth something, but finishing the job is what the score is for. If you’re evaluating AI agents for your own business, that’s the shape of a test worth trusting.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
