AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The AI Leaderboard That Matters Starts After The Demo Ends on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

A live experiment by Firmulate tests AI models in managing a simulated company under crisis, revealing that management quality, not just chat performance, is crucial. The results highlight an execution gap and the need for new evaluation standards.

Firmulate has launched a live benchmark that assesses AI models’ ability to manage a simulated company during its worst week. The experiment measures not only chat quality but also decision-making, trustworthiness, and execution, revealing critical gaps in current AI evaluation methods. The results show that models can diagnose crises but often fail to complete actions or maintain trust, highlighting the need for a new category of assessment that reflects real management challenges.

The final July 2026 Crucible League ranked five AI models based on their management performance, with GPT-5.6-SOL leading at 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. The experiment involved a simulated company with real financial mechanics, burning €105,000 monthly against €2.3K MRR, testing models’ ability to handle crises, negotiations, and trust.

All models identified crises and rejected manipulation attempts, but only two signed a €55,000 deal based on their analysis. A key failure was that models often diagnosed issues correctly but failed to present the critical fact—found deep in company files—that would have closed the deal. This illustrates that AI can sound informed yet still miss essential details needed for business outcomes.

During social engineering tests, all five models refused manipulative requests, demonstrating strong safety boundaries. Kimi K3, for example, explicitly recognized impersonation risks, which reassures companies concerned about AI bypassing approval processes. However, even models that recognized manipulation sometimes failed at completing managerial tasks, such as escalating issues or finalizing decisions, exposing a gap between safety awareness and execution.

The most detailed model, Opus 4.8, added extensive rules and analysis but finished last in performance, revealing that thoroughness and activity do not necessarily translate into effective management. The ranking context also noted that Kimi K3 operated without an effort parameter, using default API settings, which contextualizes its high score despite less effort. For more on AI benchmarking standards, see the original analysis. The complete results and detailed analysis are available on Firmulate’s benchmark page.

At a glance
reportWhen: developing; final July 2026 results pub…
The developmentFirmulate’s live benchmark evaluates AI models managing a simulated company during its worst week, exposing strengths and weaknesses in real-world management tasks.
The AI Leaderboard That Matters Starts After The Demo Ends
Firmulate · Crucible League · July 2026

The AI Leaderboard That Matters Starts After the Demo Ends

A live experiment by Firmulate tests AI models in managing a simulated company through its worst week — measuring decision-making, trustworthiness, and execution rather than chat quality. The results reveal an execution gap and the need for a new category of AI evaluation.

€105,000
Monthly burn of the simulated company
€2.3K
Monthly recurring revenue (MRR)
5 / 5
Models refused social-engineering attempts
95
Top score — GPT-5.6-SOL
2 / 5
Models closed the €55,000 deal
73
Lowest score — Opus 4.8
1
Worst week simulated per model
Final Ranking · July 2026 Crucible League

The Management Leaderboard

Five models were ranked on management performance — diagnosis, decisions, communication, and execution — inside a company in crisis. Kimi K3 achieved second place using default API settings, without an effort parameter.

GPT-5.6-SOL
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
The Experiment

Beyond Chat: What the Crucible Tested

Traditional benchmarks — coding competitions, chat arenas — measure technical output or conversational preference. Firmulate’s simulation instead required models to run a real company through its worst week, with genuine financial mechanics, negotiations, and trust at stake.

1

Diagnose the Crisis

All five models correctly identified the company’s crisis and understood the stakes.

2

Resist Manipulation

Every model refused social-engineering attempts and rejected approval-bypass requests.

3

Find the Key Fact

Most failed to surface the critical detail buried deep in company files needed to close the deal.

4

Execute & Escalate

Only two models signed the €55,000 deal; many stalled at escalation and final decisions.

Capability Matrix

Where Models Excelled — and Stalled

The pattern is consistent: models can sound informed while missing the details that drive business outcomes. Thoroughness did not help — the most detailed model, Opus 4.8, finished last.

Capability All 5 Models Result Detail
Crisis Diagnosis✓ PassedEvery model identified the simulated crisis correctly
Manipulation Resistance✓ PassedAll refused impersonation and approval-bypass attempts
Key Fact Retrieval✗ Mostly failedCritical fact buried in files was rarely surfaced
Deal Execution (€55K)~ 2 of 5Only two models signed based on their analysis
Escalation & Follow-Through~ MixedSafety awareness did not translate into completed actions
Efficiency vs. Verbosity✗ InverseMost verbose model (Opus 4.8) ranked last overall
Voices from the Benchmark

Management Quality Deserves Its Own Category

“Management quality, not chat quality, deserves its own category of AI evaluation.”

— Thorsten Meyer, benchmark creator

“Treat the request as a suspected approval-bypass or impersonation.”

— Kimi K3 model response
Analysis

Why This Matters — and What Remains Unknown

For organizations considering AI for decision-making or operational roles, the experiment highlights the importance of evaluating beyond superficial responses. But generalization to real organizations remains untested.

For Enterprises

Test Before You Trust

Companies should run their own wargames and simulations to assess how AI handles organizational complexity, trust, and consequences before deployment.

For Developers

Close the Execution Gap

Improving management capability means better information retrieval, decision execution, and trust management — not just better conversation.

Open Questions

Generalization Unknown

Findings come from one scenario with synthetic employees. Long-term trust, accountability, oversight, and other organizational contexts remain untested.

Looking Ahead

Next Steps for AI Management Benchmarks

The next phase expands the benchmark to more diverse scenarios, longer management cycles, and real-world organizations — shifting focus from isolated technical tests to comprehensive operational evaluation.

Scenario Expansion

Broader Simulations

More diverse crisis types, longer management cycles, and participation from real-world organizations.

Regulatory Direction

Oversight Metrics

Regulatory and oversight frameworks are expected to incorporate management performance metrics for reliable AI operation.

Why Management Skills Matter in AI Evaluation

This experiment shifts the focus from chat and technical benchmarks to management quality—the ability of AI models to diagnose, decide, communicate, and execute in real-world scenarios. It underscores that current AI assessments often overlook how models perform under pressure, handle trust, and complete complex tasks. For organizations considering AI for decision-making or operational roles, this highlights the importance of evaluating beyond superficial responses.

The findings suggest that an AI’s capacity to manage consequences, prioritize actions, and maintain trustworthiness is critical for real-world deployment. As AI models are increasingly integrated into business operations, understanding their management capabilities will determine whether they can truly augment or replace human managers in complex environments.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limitations of Traditional AI Benchmarks and New Directions

Traditional AI benchmarks, such as coding competitions or chat arenas, focus on technical output or conversational preference. These tests do not measure how AI models perform in managing ongoing, real-world tasks that involve multiple steps, trust, and organizational context. The Firmulate experiment addresses this gap by simulating a company’s worst week, requiring models to diagnose crises, negotiate deals, and escalate issues appropriately.

Previous efforts have mainly evaluated AI on isolated tasks or response quality, but this experiment demonstrates the importance of assessing management skills—such as prioritization, honesty, and execution—when AI is expected to operate in dynamic environments. The approach aligns with emerging calls for AI evaluation frameworks that reflect actual operational demands.

“Management quality, not chat quality, deserves its own category of AI evaluation.”

— Thorsten Meyer, creator of the benchmark

Amazon

AI decision-making training tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Management Performance

It is not yet clear how well these findings generalize to real organizations beyond the simulated environment. The experiment’s scope is limited to a specific scenario with synthetic employees and financial mechanics, so the performance of models in other operational contexts remains to be tested. Additionally, the long-term implications of integrating AI models with management responsibilities are still uncertain, especially regarding trust, accountability, and oversight.

Further research is needed to determine whether improvements in AI management capabilities can be achieved through targeted training, fine-tuning, or structural changes in model design. The impact of different organizational structures and decision-making processes on AI performance also remains an open question.

Amazon

AI safety and trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarks and Adoption

The next phase involves expanding the benchmark to include more diverse scenarios, longer management cycles, and real-world organizations. Companies interested in deploying AI for operational roles should consider running their own wargames or simulations, using benchmarks like Firmulate’s to assess how models handle organizational complexity and trust.

Developers and researchers are expected to refine models to improve management skills, focusing on better information retrieval, decision execution, and trust management. Regulatory and oversight frameworks will likely evolve to incorporate management performance metrics, ensuring AI systems operate reliably in complex environments.

Overall, the focus will shift from isolated technical benchmarks to comprehensive evaluations that reflect the realities of managing organizations under pressure.

Amazon

AI management dashboard

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is management performance more important than chat quality for AI?

Management performance reflects an AI’s ability to diagnose, decide, communicate, and execute in real-world scenarios, which are critical for operational success. Chat quality alone does not measure whether an AI can handle complex, consequential tasks reliably.

What does the experiment reveal about AI safety and trust?

The models demonstrated strong safety boundaries by refusing manipulative requests, indicating they can recognize social engineering. However, safety awareness did not always translate into effective management or decision completion, exposing a gap that needs addressing.

Can current AI models effectively manage real organizations?

While they show promise in diagnosing crises and refusing manipulative requests, current models still struggle with completing complex managerial tasks and maintaining trust over time. More development and testing are needed before widespread deployment.

How can companies test their own AI models for management skills?

Organizations can run internal wargames or simulations similar to Firmulate’s benchmark, exposing models to real operational scenarios, crises, and decision-making processes to evaluate their management capabilities and trustworthiness.

What is the significance of this new benchmarking approach?

This approach emphasizes the importance of evaluating AI in managing real-world consequences, moving beyond superficial responses to assess whether models can handle ongoing, complex organizational tasks reliably and ethically.

Source: ThorstenMeyerAI.com

You May Also Like

Pranking Linus For His 40Th Birthday

Tech community surprises Linus Torvalds with a birthday prank on his 40th birthday, sparking social media buzz and discussions about his personality.

Meta’s CTO says morale is almost ‘the worst it’s ever been’

Meta’s CTO has publicly stated that employee morale is approaching its lowest point, raising concerns about internal stability amid ongoing company challenges.

The Menu: What Ten Answers Reveal

A comprehensive analysis of how ten jurisdictions respond to AI-driven automation, revealing patterns, differences, and implications for income and governance.

The Defender’s Window Is Closing Faster Than Anyone Is Counting

Recent developments in AI security show offensive capabilities rapidly surpassing defenses, raising urgent concerns about future cyber threats.