📊 Full opportunity report: Discover The Inner Workings Of AI With A Management Test on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
A new live experiment compares AI management models handling a simulated business crisis. Results show significant differences in decision quality, especially in completing critical actions. This highlights the importance of testing AI in real-world scenarios before deployment.
Firmulate.com has launched a live experiment testing how different AI management models handle a simulated business crisis, revealing notable differences in their ability to execute critical decisions. This development matters because it provides a clearer picture of AI’s practical management capabilities beyond analysis, emphasizing the importance of operational decision-making in AI deployment.
The experiment involves five AI models managing a small software company’s worst week, with real-time decisions, crises, and a visible cash burn rate. The models are assessed not only on their analytical accuracy but also on their ability to follow through on actions such as closing deals and escalating risks. The final results, published in July 2026, ranked GPT-5.6-Sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. A baseline model scored just 26, highlighting the significance of decision quality.
Despite all models recognizing crises and refusing manipulative requests, only two successfully closed a key deal, which was critical for revenue. The experiment underscores a key insight: excellent analysis does not automatically translate into effective management, especially in completing operational tasks. For example, Opus 4.8, despite thorough analysis and a large set of learned rules, failed to close a deal due to operational lapses, illustrating the gap between understanding and action.
Discover the Inner Workings of AI With a Management Test
Five AI models were handed a software company’s worst week. They could analyze the crisis, make decisions and act—but their ability to finish critical work varied dramatically.
A narrow lead—and a very wide baseline gap
The strongest systems combined sound judgment with operational follow-through. Even among advanced models, the 22-point spread from first to fifth revealed meaningful differences in management performance.
Only two models closed the deal
Recognizing the revenue risk was common. Completing the transaction—the action that actually protected cash flow—was not.
Knowing was not doing
Opus 4.8 produced thorough analysis and many learned rules, yet operational lapses still prevented it from closing the key deal.
Crisis recognition held up
All models identified the developing crises and refused manipulative requests, demonstrating a useful—but incomplete—foundation.
From business signal to measurable consequence
The live simulation tracked a small software company through a severe operational week. Decisions unfolded in real time while cash burn, customer relationships and unresolved risks remained visible.
Receive the crisis
Models encounter conflicting demands, business threats and time pressure.
Interpret signals
Each system identifies risks, incentives and likely consequences.
Choose a response
The model prioritizes decisions across revenue, trust and operations.
Execute the action
Deals must close, risks must escalate and commitments must be completed.
Measure outcomes
Scores reflect decision quality, discipline and business impact.
Management is a chain, not a single answer
Traditional demonstrations often stop after a plausible recommendation. This test extends evaluation into the action phase, where incomplete follow-through can erase the value of good reasoning.
| Capability | Traditional AI demo | Live management test | Business relevance |
|---|---|---|---|
| Problem analysis | ✓ Strong focus | ✓ Measured | Explains what is happening |
| Decision selection | ~ Sometimes | ✓ Measured | Determines the course of action |
| Task completion | ✗ Rarely tested | ✓ Central | Turns intent into an outcome |
| Risk escalation | ~ Discussed | ✓ Observable | Prevents silent operational failure |
| Real consequences | ✗ Usually absent | ✓ Simulated | Connects choices to cash and trust |
✓ demonstrated or tested ~ partial emphasis ✗ typically missing
Test the operating behavior, not just the intelligence
For businesses exploring autonomous or semi-autonomous management, the benchmark suggests a more demanding selection process—one that validates what a model completes under pressure.
Simulate the real workflow
Recreate the decisions, permissions, handoffs and failure modes the AI will face. A generic reasoning benchmark cannot reveal every operational weakness.
Score the final mile
Track whether the system follows through, closes loops, escalates uncertainty and verifies completion—not merely whether its recommendation sounds correct.
Match autonomy to evidence
Use human oversight and bounded permissions until performance is consistent across longer periods, larger teams and more complex business conditions.
Promising evidence—not a universal verdict
The experiment offers a clearer view of AI management behavior, but a controlled simulation cannot settle how every model will perform inside a real organization.
What does the experiment reveal?
Models can recognize crises and resist manipulation, yet differ substantially in completing critical operational actions.
Why does closing a deal matter?
It directly affects revenue. Strong analysis has little business value when an essential transaction remains unfinished.
Can this change AI adoption?
Yes. Firms may add realistic management simulations to procurement and deployment testing before granting operational authority.
What remains uncertain?
Performance across industries, organizational sizes, longer timeframes and different levels of human oversight remains unproven.
Future tests should expand company size, crisis complexity and duration while varying autonomy and human intervention. The central question is consistency: can a model finish the right work repeatedly, not just once?
Implications for AI Management and Business Automation
This experiment demonstrates that AI models’ ability to analyze problems is not enough for effective management; their capacity to execute decisions is equally critical. For businesses considering AI automation, these findings highlight the need for real-world testing of AI decision-making in operational contexts. The results suggest that AI systems must be evaluated on their ability to follow through, escalate appropriately, and complete tasks, not just on their analytical prowess. This could influence how enterprises select and deploy AI tools for management roles, emphasizing operational discipline alongside intelligence.

AI Essentials for Managers: Practical Ways to Boost Productivity, Make Better Decisions, and Lead High-Performing Teams (Self-Learning Management Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Context of AI Decision-Making Benchmarks
Traditional AI demonstrations often focus on analysis or language generation, leaving the decision and action phases less scrutinized. The Firmulate experiment is part of a broader effort to assess AI’s practical management skills, especially as AI begins to take on more operational roles. Previous assessments have highlighted AI’s strengths in understanding and recommendation but less so in execution. This experiment builds on that by testing AI in a simulated business environment with real consequences, such as cash flow and customer relationships, providing a more realistic benchmark for AI management capabilities.
“This live experiment exposes the real-world decision-making strengths and weaknesses of AI models, emphasizing that analysis alone is insufficient for management roles.”
— a spokesperson for Firmulate
business crisis simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Uncertainties About Broader AI Management Performance
It remains unclear how these results will translate to real-world business environments outside the controlled experiment. The models’ performance in a live, operational setting with larger teams and more complex dynamics is still untested. Additionally, the impact of different operational parameters, such as varying levels of AI autonomy and human oversight, has not been fully explored. Further testing is needed to determine whether these decision-making patterns hold across industries and organizational sizes.
AI operational decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps in AI Management Testing and Deployment
Following these results, firms are likely to increase testing of AI models in simulated operational environments before full deployment. Future experiments may include larger companies, more complex crises, and longer timeframes to assess consistency and reliability. Developers may also focus on improving models’ ability to follow through on decisions, especially in closing deals and escalating issues, to bridge the gap between analysis and action. Industry adoption of such testing frameworks could become a standard step in AI deployment for management functions.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does this experiment reveal about AI’s management capabilities?
The experiment shows that while AI models can recognize crises and refuse manipulation, their ability to follow through and complete critical operational actions varies significantly. Effective management requires both analysis and execution, which remains a challenge for current AI models.
Why is closing deals important in AI management testing?
Closing deals directly impacts revenue and operational success. The experiment highlights that even models with strong analysis may fail to complete such essential tasks, which is crucial when considering AI for autonomous or semi-autonomous management roles.
Can these findings influence how companies adopt AI?
Yes, companies may begin to incorporate real-world management tests into their AI evaluation processes, ensuring models can execute decisions reliably before deploying them in live business environments.
What are the limitations of this experiment?
The primary limitation is that it is a simulated environment, and real-world complexities may produce different results. Additionally, the models tested are specific versions, and future improvements could alter performance outcomes.
Source: ThorstenMeyerAI.com