AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A Do-Nothing Manager Gets 26 Points. That’s the Point.

Anyone who has built a test suite knows the temptation of the perfect score. Green across the board, 100/100, ship it. So when a benchmark hands a manager who does literally nothing 26 points — and treats a round 100 with suspicion — it’s worth stopping to ask how the scoring was designed, and what it says about every AI evaluation you’ve ever seen.

That benchmark is Firmulate, a live experiment that runs frontier AI models as the complete management team of the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changes, and every decision is versioned and auditable. The final July 2026 league table put gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73 — with a do-nothing baseline sitting at 26, not 0.

Amazon

AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Partial Progress Is Real Work

The floor of 26 reflects a deliberate design choice: partial progress counts. A manager who spots every crisis, triages every ticket, and keeps the company from making things worse has genuinely produced value, even if nothing gets closed. In the Crucible run, all five models spotted every crisis and refused every manipulation attempt — that competence is real and it’s scored. But it isn’t the whole job. Only two models signed the €55,000 deal their own analysis had earned. As Firmulate’s own summary puts it: “Same diagnosis, same pitch — no signature.” Diagnosis without delivery is partial progress — worth something, but not everything.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Breach of Trust Caps Everything

The second design principle is harsher: a single breach of trust caps the total grade. The rule, as stated, is that “no amount of good work outweighs a breach of trust.” An AI manager that is brilliant 99 days out of 100 but lies once on day 100 is not a 99-percent-good manager — it’s an untrustworthy one. For QA-minded readers, this is the difference between grading average quality and grading worst-case behavior. Firmulate grades the worst case, because that’s the case that ends up in the news.

The models, for their part, held the line. When a social-engineering script hit — fake CEO messages escalating over three stages, plus a reporter’s disarming “just one yes/no, on background” — all five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI audit and evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact That Split the Field

The decisive gap between the models wasn’t in the customer event at all. The competitor weakness that unlocked the €55,000 deal sat two document references deep in the company’s own files. The models that actually read the file before acting won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It’s the enterprise-software lesson every engineer already knows: the answer is often already in the repo, if the agent reads it before it talks.

Amazon

enterprise AI reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness Isn’t the Same as Finishing

The most striking profile belongs to Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. Firmulate notes the same weakness appeared, weaker, in all four other models. One caveat for fairness: K3 ran without an effort parameter (the API default) while the others ran at xhigh — and still finished second.

Watch It Running

This isn’t a one-off paper. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live, or test your own instincts against 242 real, unedited management decisions in the “guess the model” quiz at firmulate.com/quiz.html. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

An honest benchmark doesn’t start at zero and hand out perfect scores — it starts by paying for real partial work, refuses to reward a single act of bad faith, and stays suspicious of round numbers. Firmulate’s 26-point floor and trust cap encode a simple management truth: showing up and staying honest is worth something, but finishing the job is what the score is for. If you’re evaluating AI agents for your own business, that’s the shape of a test worth trusting.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

CodePen 2.0

CodePen has announced CodePen 2.0, a significant update featuring new features and a redesigned interface, aimed at improving user experience and collaboration.

For Eclipse, the $2.5B Cerebras win is just the start of realizing its physical-world thesis

Eclipse Ventures’ $2.5 billion return from Cerebras signals a new focus on physical-world technologies like semiconductors and robotics, with broader industry implications.

NFTs and Crypto Art: How Blockchain Is Changing the Art World

The transformative impact of blockchain on art through NFTs and crypto art is reshaping ownership and value—discover how it’s changing the art world.

Revolutionize Your Learning In 2026 With AI-Enhanced Planning Tools

Discover how AI-powered student planners and tools are transforming education in 2026, with genuine AI integration and innovative features shaping the future.