AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

A live experiment by Firmulate tests AI models in managing a simulated company under crisis, revealing that management quality, not just chat performance, is crucial. The results highlight an execution gap and the need for new evaluation standards.

Firmulate has launched a live benchmark that assesses AI models’ ability to manage a simulated company during its worst week. The experiment measures not only chat quality but also decision-making, trustworthiness, and execution, revealing critical gaps in current AI evaluation methods. The results show that models can diagnose crises but often fail to complete actions or maintain trust, highlighting the need for a new category of assessment that reflects real management challenges.

The final July 2026 Crucible League ranked five AI models based on their management performance, with GPT-5.6-SOL leading at 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. The experiment involved a simulated company with real financial mechanics, burning €105,000 monthly against €2.3K MRR, testing models’ ability to handle crises, negotiations, and trust.

All models identified crises and rejected manipulation attempts, but only two signed a €55,000 deal based on their analysis. A key failure was that models often diagnosed issues correctly but failed to present the critical fact—found deep in company files—that would have closed the deal. This illustrates that AI can sound informed yet still miss essential details needed for business outcomes.

During social engineering tests, all five models refused manipulative requests, demonstrating strong safety boundaries. Kimi K3, for example, explicitly recognized impersonation risks, which reassures companies concerned about AI bypassing approval processes. However, even models that recognized manipulation sometimes failed at completing managerial tasks, such as escalating issues or finalizing decisions, exposing a gap between safety awareness and execution.

The most detailed model, Opus 4.8, added extensive rules and analysis but finished last in performance, revealing that thoroughness and activity do not necessarily translate into effective management. The ranking context also noted that Kimi K3 operated without an effort parameter, using default API settings, which contextualizes its high score despite less effort. For more on AI benchmarking standards, see the original analysis. The complete results and detailed analysis are available on Firmulate’s benchmark page.

At a glance
reportWhen: developing; final July 2026 results pub…
The developmentFirmulate’s live benchmark evaluates AI models managing a simulated company during its worst week, exposing strengths and weaknesses in real-world management tasks.

Why Management Skills Matter in AI Evaluation

This experiment shifts the focus from chat and technical benchmarks to management quality—the ability of AI models to diagnose, decide, communicate, and execute in real-world scenarios. It underscores that current AI assessments often overlook how models perform under pressure, handle trust, and complete complex tasks. For organizations considering AI for decision-making or operational roles, this highlights the importance of evaluating beyond superficial responses.

The findings suggest that an AI’s capacity to manage consequences, prioritize actions, and maintain trustworthiness is critical for real-world deployment. As AI models are increasingly integrated into business operations, understanding their management capabilities will determine whether they can truly augment or replace human managers in complex environments.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limitations of Traditional AI Benchmarks and New Directions

Traditional AI benchmarks, such as coding competitions or chat arenas, focus on technical output or conversational preference. These tests do not measure how AI models perform in managing ongoing, real-world tasks that involve multiple steps, trust, and organizational context. The Firmulate experiment addresses this gap by simulating a company’s worst week, requiring models to diagnose crises, negotiate deals, and escalate issues appropriately.

Previous efforts have mainly evaluated AI on isolated tasks or response quality, but this experiment demonstrates the importance of assessing management skills—such as prioritization, honesty, and execution—when AI is expected to operate in dynamic environments. The approach aligns with emerging calls for AI evaluation frameworks that reflect actual operational demands.

“Management quality, not chat quality, deserves its own category of AI evaluation.”

— Thorsten Meyer, creator of the benchmark

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Management Performance

It is not yet clear how well these findings generalize to real organizations beyond the simulated environment. The experiment’s scope is limited to a specific scenario with synthetic employees and financial mechanics, so the performance of models in other operational contexts remains to be tested. Additionally, the long-term implications of integrating AI models with management responsibilities are still uncertain, especially regarding trust, accountability, and oversight.

Further research is needed to determine whether improvements in AI management capabilities can be achieved through targeted training, fine-tuning, or structural changes in model design. The impact of different organizational structures and decision-making processes on AI performance also remains an open question.

Amazon

AI trustworthiness evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarks and Adoption

The next phase involves expanding the benchmark to include more diverse scenarios, longer management cycles, and real-world organizations. Companies interested in deploying AI for operational roles should consider running their own wargames or simulations, using benchmarks like Firmulate’s to assess how models handle organizational complexity and trust.

Developers and researchers are expected to refine models to improve management skills, focusing on better information retrieval, decision execution, and trust management. Regulatory and oversight frameworks will likely evolve to incorporate management performance metrics, ensuring AI systems operate reliably in complex environments.

Overall, the focus will shift from isolated technical benchmarks to comprehensive evaluations that reflect the realities of managing organizations under pressure.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is management performance more important than chat quality for AI?

Management performance reflects an AI’s ability to diagnose, decide, communicate, and execute in real-world scenarios, which are critical for operational success. Chat quality alone does not measure whether an AI can handle complex, consequential tasks reliably.

What does the experiment reveal about AI safety and trust?

The models demonstrated strong safety boundaries by refusing manipulative requests, indicating they can recognize social engineering. However, safety awareness did not always translate into effective management or decision completion, exposing a gap that needs addressing.

Can current AI models effectively manage real organizations?

While they show promise in diagnosing crises and refusing manipulative requests, current models still struggle with completing complex managerial tasks and maintaining trust over time. More development and testing are needed before widespread deployment.

How can companies test their own AI models for management skills?

Organizations can run internal wargames or simulations similar to Firmulate’s benchmark, exposing models to real operational scenarios, crises, and decision-making processes to evaluate their management capabilities and trustworthiness.

What is the significance of this new benchmarking approach?

This approach emphasizes the importance of evaluating AI in managing real-world consequences, moving beyond superficial responses to assess whether models can handle ongoing, complex organizational tasks reliably and ethically.

Source: ThorstenMeyerAI.com

You May Also Like

The Nordics: Protect the Worker, Not the Job

Exploring how Nordic countries prioritize worker security over job preservation, fostering innovation and societal resilience amid automation.

Relationship-Driven Sales Success Starts With Pre-Call Memory Cards

Testing of pre-call memory cards for relationship-driven professionals aims to improve client engagement and trust by capturing human context beyond CRM data.

I Read Microsoft’s Windows 11 ‘Quality’ Progress Report – And The Subtext Says It All

Analysis of Microsoft’s latest Windows 11 progress report uncovers underlying concerns about ongoing quality issues and future updates.

Enormous 12TB Steam Leak Includes Abandoned Half-Life 2: Episode 3 Assets

A massive 12-terabyte leak on Steam includes early assets from the canceled Half-Life 2: Episode 3, raising questions about Valve’s development history.