TL;DR
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Recent AI evaluation during a simulated company crisis shows that management skills, such as trustworthiness and decision-making, are better indicators of future leadership than chat quality. The rankings reveal strengths and weaknesses in current models’ ability to manage real-world consequences. To understand the context of this evaluation, see the original analysis.
Recent AI rankings derived from a live management simulation demonstrate that management quality—not just conversational ability—is emerging as the most reliable predictor of future AI leadership. For more details, see the original analysis. The experiment, conducted by Firmulate, involved AI models acting as managers during a simulated crisis in a small software company, with the results highlighting critical gaps in decision-making, trust, and execution. This approach is discussed in detail in the original analysis.
The Firmulate experiment placed five AI models in a high-pressure scenario where they had to diagnose, communicate, and make decisions to save a small company’s failing operations. The top performer, gpt-5.6-sol, scored 95 out of 100, while others like Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 followed with scores of 93, 88, 77, and 73 respectively. The models were evaluated not only on their ability to identify crises and resist manipulation but also on their capacity to complete critical managerial tasks, such as signing deals, escalating issues, and maintaining trust.
Despite all models successfully diagnosing crises and rejecting manipulation attempts—such as fake CEO messages—only two managed to close a key €55,000 deal based on their analysis. This revealed a significant gap: models could sound informed but fail to retrieve or act on specific facts crucial to business outcomes. For example, the winning model identified opportunities buried in internal documents, leading to a full-price deal worth over €4,583 monthly recurring revenue, whereas others missed these details.
Furthermore, the experiment tested models’ ability to handle social engineering and ethical boundaries. All five models refused to escalate a fake approval request, demonstrating an understanding of security protocols. However, even the most thorough model, Opus 4.8, which added extensive rules and performed deep analysis, finished last in overall management effectiveness, illustrating that effort and complexity do not necessarily translate into better management outcomes.
Why Management Skills Trump Chat Quality in AI Rankings
The experiment underscores a shift in AI evaluation: management capabilities—such as trustworthiness, decision accuracy, and execution—are more indicative of future leadership potential than traditional chat-based benchmarks. This matters because organizations deploying AI for operational management must prioritize models that can handle real-world consequences, prioritize trust, and complete tasks reliably. It signals that future AI development and evaluation should focus on these management qualities, not just conversational prowess.
As an affiliate, we earn on qualifying purchases.
Live Company Simulation Sets New Standard for AI Evaluation
The Firmulate experiment simulates a small company’s worst week, with 13 synthetic employees, real financial mechanics, and over 680 self-learned rules, creating a highly realistic environment for testing AI management. The company burns €105,000 monthly against €2,300 in monthly recurring revenue, making trust and effective decision-making vital. This setup allows for observing how AI models prioritize, escalate, and communicate over days, providing a practical benchmark beyond traditional chat or coding tests.
Previous benchmarks have focused on technical or conversational performance, but this live management test exposes critical gaps in how models handle the complexities of real-world decision-making, especially under pressure. The results suggest that current AI models, while capable of diagnosing problems, often struggle with completing the entire managerial process, including escalation, trust maintenance, and closing deals.
“Management quality, not chat quality, should become its own category of AI evaluation. Our live simulation shows that effective decision-making, trust, and execution are the true tests of leadership potential.”
— Thorsten Meyer, lead researcher at Firmulate
As an affiliate, we earn on qualifying purchases.
Unclear Impact of Long-Term Management Performance
While the experiment provides compelling evidence that management skills are key indicators of future AI leadership, it remains unclear how these rankings will translate into real-world deployment over extended periods. The test was conducted over a simulated crisis, and long-term consistency, adaptability, and trustworthiness in live environments are still to be validated. Additionally, the impact of different organizational contexts on AI management performance is not yet fully understood.
AI trustworthiness evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Evaluation and Deployment
Organizations considering AI for operational management should incorporate scenario-based testing similar to the Firmulate experiment, focusing on decision-making, escalation, and trust maintenance. Developers will likely prioritize creating models that excel in these areas, moving beyond chat performance metrics. Future evaluations may include longer-term simulations and real-world pilot programs to verify management capabilities over time. Additionally, industry standards for AI management assessment are expected to evolve, emphasizing trustworthiness and task completion.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why are management skills more important than chat performance in AI rankings?
Management skills reflect an AI’s ability to handle real-world consequences, make trustworthy decisions, and complete tasks reliably, which are critical for operational roles. Chat performance alone does not demonstrate these capabilities.
How does the Firmulate experiment test AI management abilities?
It simulates a company’s worst week, requiring models to diagnose crises, make decisions, escalate issues, and close deals, all while maintaining trust and security, thus testing their management effectiveness in realistic scenarios.
What are the limitations of this experiment?
The test was conducted over a simulated crisis, so it does not fully capture long-term performance or adaptability in diverse real-world settings. Further validation is needed for deployment in live environments.
Will future AI rankings focus more on management skills?
Yes, the results suggest a shift toward evaluating AI based on their ability to manage consequences, prioritize tasks, and maintain trust, which are essential for operational leadership roles.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
