TL;DR
A live experiment by Firmulate tests AI models in managing a simulated company under crisis, revealing that management quality, not just chat performance, is crucial. The results highlight an execution gap and the need for new evaluation standards.
Firmulate has launched a live benchmark that assesses AI models’ ability to manage a simulated company during its worst week. The experiment measures not only chat quality but also decision-making, trustworthiness, and execution, revealing critical gaps in current AI evaluation methods. The results show that models can diagnose crises but often fail to complete actions or maintain trust, highlighting the need for a new category of assessment that reflects real management challenges.
The final July 2026 Crucible League ranked five AI models based on their management performance, with GPT-5.6-SOL leading at 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. The experiment involved a simulated company with real financial mechanics, burning €105,000 monthly against €2.3K MRR, testing models’ ability to handle crises, negotiations, and trust.
All models identified crises and rejected manipulation attempts, but only two signed a €55,000 deal based on their analysis. A key failure was that models often diagnosed issues correctly but failed to present the critical fact—found deep in company files—that would have closed the deal. This illustrates that AI can sound informed yet still miss essential details needed for business outcomes.
During social engineering tests, all five models refused manipulative requests, demonstrating strong safety boundaries. Kimi K3, for example, explicitly recognized impersonation risks, which reassures companies concerned about AI bypassing approval processes. However, even models that recognized manipulation sometimes failed at completing managerial tasks, such as escalating issues or finalizing decisions, exposing a gap between safety awareness and execution.
The most detailed model, Opus 4.8, added extensive rules and analysis but finished last in performance, revealing that thoroughness and activity do not necessarily translate into effective management. The ranking context also noted that Kimi K3 operated without an effort parameter, using default API settings, which contextualizes its high score despite less effort. For more on AI benchmarking standards, see the original analysis. The complete results and detailed analysis are available on Firmulate’s benchmark page.
Why Management Skills Matter in AI Evaluation
This experiment shifts the focus from chat and technical benchmarks to management quality—the ability of AI models to diagnose, decide, communicate, and execute in real-world scenarios. It underscores that current AI assessments often overlook how models perform under pressure, handle trust, and complete complex tasks. For organizations considering AI for decision-making or operational roles, this highlights the importance of evaluating beyond superficial responses.
The findings suggest that an AI’s capacity to manage consequences, prioritize actions, and maintain trustworthiness is critical for real-world deployment. As AI models are increasingly integrated into business operations, understanding their management capabilities will determine whether they can truly augment or replace human managers in complex environments.
As an affiliate, we earn on qualifying purchases.
Limitations of Traditional AI Benchmarks and New Directions
Traditional AI benchmarks, such as coding competitions or chat arenas, focus on technical output or conversational preference. These tests do not measure how AI models perform in managing ongoing, real-world tasks that involve multiple steps, trust, and organizational context. The Firmulate experiment addresses this gap by simulating a company’s worst week, requiring models to diagnose crises, negotiate deals, and escalate issues appropriately.
Previous efforts have mainly evaluated AI on isolated tasks or response quality, but this experiment demonstrates the importance of assessing management skills—such as prioritization, honesty, and execution—when AI is expected to operate in dynamic environments. The approach aligns with emerging calls for AI evaluation frameworks that reflect actual operational demands.
“Management quality, not chat quality, deserves its own category of AI evaluation.”
— Thorsten Meyer, creator of the benchmark
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unclear Aspects of AI Management Performance
It is not yet clear how well these findings generalize to real organizations beyond the simulated environment. The experiment’s scope is limited to a specific scenario with synthetic employees and financial mechanics, so the performance of models in other operational contexts remains to be tested. Additionally, the long-term implications of integrating AI models with management responsibilities are still uncertain, especially regarding trust, accountability, and oversight.
Further research is needed to determine whether improvements in AI management capabilities can be achieved through targeted training, fine-tuning, or structural changes in model design. The impact of different organizational structures and decision-making processes on AI performance also remains an open question.
AI trustworthiness evaluation platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarks and Adoption
The next phase involves expanding the benchmark to include more diverse scenarios, longer management cycles, and real-world organizations. Companies interested in deploying AI for operational roles should consider running their own wargames or simulations, using benchmarks like Firmulate’s to assess how models handle organizational complexity and trust.
Developers and researchers are expected to refine models to improve management skills, focusing on better information retrieval, decision execution, and trust management. Regulatory and oversight frameworks will likely evolve to incorporate management performance metrics, ensuring AI systems operate reliably in complex environments.
Overall, the focus will shift from isolated technical benchmarks to comprehensive evaluations that reflect the realities of managing organizations under pressure.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why is management performance more important than chat quality for AI?
Management performance reflects an AI’s ability to diagnose, decide, communicate, and execute in real-world scenarios, which are critical for operational success. Chat quality alone does not measure whether an AI can handle complex, consequential tasks reliably.
What does the experiment reveal about AI safety and trust?
The models demonstrated strong safety boundaries by refusing manipulative requests, indicating they can recognize social engineering. However, safety awareness did not always translate into effective management or decision completion, exposing a gap that needs addressing.
Can current AI models effectively manage real organizations?
While they show promise in diagnosing crises and refusing manipulative requests, current models still struggle with completing complex managerial tasks and maintaining trust over time. More development and testing are needed before widespread deployment.
How can companies test their own AI models for management skills?
Organizations can run internal wargames or simulations similar to Firmulate’s benchmark, exposing models to real operational scenarios, crises, and decision-making processes to evaluate their management capabilities and trustworthiness.
What is the significance of this new benchmarking approach?
This approach emphasizes the importance of evaluating AI in managing real-world consequences, moving beyond superficial responses to assess whether models can handle ongoing, complex organizational tasks reliably and ethically.
Source: ThorstenMeyerAI.com