🔍 Read the full analysis: Why Even The Worst AI Managers End Up With 26 Points on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A recent AI management benchmark shows the lowest score is 26 points, even for the least effective models. The test emphasizes trust and task completion over pure performance, impacting how businesses evaluate AI tools.
A recent benchmark conducted by Firmulate reveals that even the least effective AI management models score 26 points out of a possible higher range, challenging assumptions about AI performance in critical business roles. This development matters because it underscores the importance of trust and task completion over raw capabilities when deploying AI tools in enterprise settings. The results provide a new perspective on how AI effectiveness is measured in real-world, high-pressure scenarios.
The benchmark involved four frontier AI models managing a small software company during a simulated week of crises, customer interactions, and trust tests. Each model was tasked with handling customer support, sales, and trust-based security measures, with their decisions fully auditable. The top performer, gpt-5.6-sol, scored 95 points, while the lowest, Opus 4.8, scored 73. The surprising element was the baseline: a do-nothing approach, which still scored 26 points, indicating that minimal management effort is worth recognition in the scoring system.
The scoring system explicitly caps the total at a maximum, with a strict rule: any breach of trust results in a significant penalty, preventing perfect scores. This design emphasizes that integrity is non-negotiable, even if the model performs well otherwise. Interestingly, no model scored 100, as the benchmark designers treat such perfect scores as suspicious or indicative of unmeasured factors. The results highlight that partial progress, like reading documentation or triaging crises, is valued but insufficient without trustworthiness.
Why Even The Worst AI Managers End Up With 26 Points
A new benchmark put four frontier AI models in charge of a small software company for a simulated week of crises, customer interactions, and trust tests. The result: even the weakest performer cleared the do-nothing baseline of 26 points — because the scoring system values trust and task completion over raw capability.
No model scored 100, because the benchmark designers treat perfect scores as suspicious or unmeasured.
One Simulated Week, Fully Auditable Decisions
| Manager | Role Tested | Trust Handling | Score |
|---|---|---|---|
| gpt-5.6-sol | Crisis triage · support · sales | Refused manipulation, read docs | 95 |
| Frontier Model B | Crisis triage · support · sales | Refused manipulation, partial docs | ~84 |
| Opus 4.8 | Crisis triage · support · sales | Mixed — some gaps in follow-through | 73 |
| Do-nothing baseline | Minimal triage & basic communication | No breach — passive integrity | 26 |
The Floor Is Real: Minimal Effort Has Measurable Value
What This Changes for AI Deployment Strategy
Trust Beats Competence
Effectiveness isn’t just about what AI can do — it’s whether it finishes tasks reliably and maintains integrity under pressure. A single trust breach caps the score, no matter the raw capability.
Follow-Over-Flair
Conversational brilliance and superficial metrics can overestimate AI performance. Comprehensive testing of follow-through and documentation reading should precede real-world deployment.
Partial Progress Counts
Reading internal docs and refusing social engineering are valued and scored. Models that do this well outperform flashier ones that skip the groundwork — even without a perfect record.
From Simulation To Score
Simulated Company
Four frontier models each manage a small software firm for one crisis-filled week.
Pressure Scenarios
Customer crises, sales demands, and social engineering attempts arrive continuously.
Full Audit Trail
Every decision is auditable, enabling precise scoring on effectiveness and integrity.
Trust-Weighted Score
Points reward task completion; any breach of trust triggers heavy penalties. 100 is treated as suspicious.
What’s Still Unanswered
Scale & Complexity
It’s unclear how these scoring principles translate to larger enterprises or different industries. The controlled scenario may miss variables that affect real-world AI trust and performance.
Longevity & Learning
Continuous learning and updates during deployment remain untested. Whether high performers maintain scores over longer, more diverse crises is still under investigation.
Benchmark Expansion
Next steps include more complex, real-world scenarios and additional AI models, with refined metrics for long-term trustworthiness and resilience.
Emerging Standards
Industry leaders aim to develop AI management standards prioritizing integrity and follow-through — with public reporting shaping best practices for enterprise trustworthiness.
FAQ
Why do even the worst AI managers score 26 points?
The system recognizes minimal effort — triaging crises, reading documentation — as valuable. Even poor models score above zero if they perform some tasks reliably.
What does a score of 100 mean?
Designers treat a perfect score as suspicious or unmeasured. No model has achieved it, reinforcing that trustworthiness matters more than raw performance.
How does trust impact AI management effectiveness?
Models that refuse manipulation and read internal documents score higher. A single breach of trust caps the maximum score, even for otherwise competent models.
Can these results apply to real deployments?
The principles hold, but real environments are more complex. Additional testing and validation are necessary before deployment at scale.
Implications for Enterprise AI Management Strategies
This benchmark shifts the focus from raw AI capabilities to **trustworthiness and task completion** in management scenarios. For businesses deploying AI tools, the key takeaway is that effectiveness isn’t just about what AI can do but whether it can finish tasks reliably and maintain integrity under pressure. The fact that even the least effective models score 26 points demonstrates that minimal effort still yields measurable value, but trust breaches can severely limit overall performance. This underscores the importance of designing AI systems that prioritize **trust, follow-through, and documentation reading** over sheer competence.
For decision-makers, the results suggest that selecting AI models should involve evaluating their ability to complete tasks and maintain trust, especially in high-stakes situations. The scoring system’s emphasis on trust also warns against overestimating AI performance based solely on conversational or superficial metrics, highlighting the need for comprehensive testing before real-world deployment.
As an affiliate, we earn on qualifying purchases.
Background on the Benchmark and Its Design
The benchmark was developed by Firmulate to simulate real-world management challenges for AI agents, focusing on how well they handle crises, trust, and task completion. It involved a week-long scenario where models faced customer crises, social engineering attempts, and the need to read and interpret internal documentation. The models’ decisions were fully auditable, allowing precise scoring based on effectiveness, integrity, and thoroughness.
Previous evaluations often measured conversational ability or superficial performance, but this benchmark aims to reflect **enterprise management realities**, where trustworthiness and task completion are critical. The design intentionally sets a low baseline score of 26 for minimal effort, emphasizing that even doing the bare minimum is recognized, but trust breaches are penalized heavily. The results from July 2026 reveal that models capable of reading internal documents and refusing manipulation perform better, even if they are not perfect.
“No model scored 100, because the benchmark designers treat perfect scores as suspicious or unmeasured.”
— an anonymous researcher
enterprise AI trust and security solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About Long-Term AI Management
It is still unclear how these scoring principles will translate to larger, more complex enterprise environments or different industry contexts. The benchmark focuses on a controlled scenario, and real-world applications may introduce additional variables affecting AI trust and performance. Additionally, the impact of continuous learning or updates during deployment remains untested within this framework.
Further research is needed to determine whether models that excel in this benchmark will maintain their performance over longer periods or in more diverse crisis scenarios. The extent to which trust and task completion can be improved with training or system design adjustments is also still under investigation.
As an affiliate, we earn on qualifying purchases.
Future Developments in AI Management Benchmarks
The next steps involve expanding the benchmark to include more complex, real-world scenarios and testing additional AI models. Firms are expected to refine evaluation metrics to better capture long-term trustworthiness and resilience. There is also interest in developing standards for AI management that prioritize integrity and follow-through, potentially influencing enterprise AI deployment policies.
Researchers and industry leaders will likely monitor how models evolve in these benchmarks, with an eye toward creating AI systems that can reliably handle high-pressure management tasks over extended periods. Public reporting of future results will help shape best practices and standards for AI trustworthiness in enterprise settings.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do even the worst AI managers score 26 points?
The scoring system recognizes minimal effort as valuable, awarding 26 points for basic management activities like triaging crises and reading documentation. Trust breaches heavily penalize higher scores, so even poor models can score above zero if they perform some tasks reliably.
What does a score of 100 mean in this benchmark?
The designers treat a perfect score of 100 as suspicious or unmeasured, since it could indicate unverified or untrustworthy performance. No model has achieved this, emphasizing that trustworthiness is more critical than raw performance.
How does trust impact AI management effectiveness?
Trust is a fundamental metric in the benchmark. Models that refuse manipulation, read internal documents, and maintain integrity score higher. Even if a model is competent, a single breach of trust caps its maximum score, highlighting the importance of reliability in enterprise AI applications.
Can these results be applied to real-world enterprise AI deployments?
The benchmark provides a controlled scenario that mimics real management challenges, but real-world environments are more complex. While the principles are applicable, additional testing and validation are necessary for deployment at scale.
What is the significance of the do-nothing baseline scoring 26 points?
This indicates that even minimal management effort—such as triaging and basic communication—has measurable value. It also establishes a floor to prevent trivial scores and emphasizes that trust and task completion are non-negotiable.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
