AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Discover The Inner Workings Of AI With A Management Test on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

A new live experiment compares AI management models handling a simulated business crisis. Results show significant differences in decision quality, especially in completing critical actions. This highlights the importance of testing AI in real-world scenarios before deployment.

Firmulate.com has launched a live experiment testing how different AI management models handle a simulated business crisis, revealing notable differences in their ability to execute critical decisions. This development matters because it provides a clearer picture of AI’s practical management capabilities beyond analysis, emphasizing the importance of operational decision-making in AI deployment.

The experiment involves five AI models managing a small software company’s worst week, with real-time decisions, crises, and a visible cash burn rate. The models are assessed not only on their analytical accuracy but also on their ability to follow through on actions such as closing deals and escalating risks. The final results, published in July 2026, ranked GPT-5.6-Sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. A baseline model scored just 26, highlighting the significance of decision quality.

Despite all models recognizing crises and refusing manipulative requests, only two successfully closed a key deal, which was critical for revenue. The experiment underscores a key insight: excellent analysis does not automatically translate into effective management, especially in completing operational tasks. For example, Opus 4.8, despite thorough analysis and a large set of learned rules, failed to close a deal due to operational lapses, illustrating the gap between understanding and action.

At a glance
reportWhen: ongoing, with results announced in July…
The developmentFirmulate.com launched a live management test pitting five AI models against a simulated business crisis, revealing varied decision-making capabilities.
Discover The Inner Workings Of AI With A Management Test
Live management benchmark · July 2026

Discover the Inner Workings of AI With a Management Test

Five AI models were handed a software company’s worst week. They could analyze the crisis, make decisions and act—but their ability to finish critical work varied dramatically.

Models tested 5 plus a baseline
Winning score 95 GPT-5.6-Sol
Baseline score 26 69 points behind
Closed key deal 2/5 execution was scarce
01 · Final scoreboard

A narrow lead—and a very wide baseline gap

The strongest systems combined sound judgment with operational follow-through. Even among advanced models, the 22-point spread from first to fifth revealed meaningful differences in management performance.

Model performance Points / 100
01 GPT-5.6-Sol
95
02 Kimi K3
93
03 Sonnet 5
88
04 Fable 5
77
05 Opus 4.8
73
B Baseline
26
Critical action

Only two models closed the deal

Recognizing the revenue risk was common. Completing the transaction—the action that actually protected cash flow—was not.

Operational gap

Knowing was not doing

Opus 4.8 produced thorough analysis and many learned rules, yet operational lapses still prevented it from closing the key deal.

Shared strength

Crisis recognition held up

All models identified the developing crises and refused manipulative requests, demonstrating a useful—but incomplete—foundation.

02 · Inside the experiment

From business signal to measurable consequence

The live simulation tracked a small software company through a severe operational week. Decisions unfolded in real time while cash burn, customer relationships and unresolved risks remained visible.

01

Receive the crisis

Models encounter conflicting demands, business threats and time pressure.

02

Interpret signals

Each system identifies risks, incentives and likely consequences.

03

Choose a response

The model prioritizes decisions across revenue, trust and operations.

04

Execute the action

Deals must close, risks must escalate and commitments must be completed.

05

Measure outcomes

Scores reflect decision quality, discipline and business impact.

03 · What the benchmark measures

Management is a chain, not a single answer

Traditional demonstrations often stop after a plausible recommendation. This test extends evaluation into the action phase, where incomplete follow-through can erase the value of good reasoning.

Capability Traditional AI demo Live management test Business relevance
Problem analysis ✓ Strong focus ✓ Measured Explains what is happening
Decision selection ~ Sometimes ✓ Measured Determines the course of action
Task completion ✗ Rarely tested ✓ Central Turns intent into an outcome
Risk escalation ~ Discussed ✓ Observable Prevents silent operational failure
Real consequences ✗ Usually absent ✓ Simulated Connects choices to cash and trust

✓ demonstrated or tested    ~ partial emphasis    ✗ typically missing

Traceability chain · where value can break
🔎 Recognize
🧠 Analyze
⚖️ Decide
⚙️ Execute
📈 Verify
04 · Implications for deployment

Test the operating behavior, not just the intelligence

For businesses exploring autonomous or semi-autonomous management, the benchmark suggests a more demanding selection process—one that validates what a model completes under pressure.

Before deployment

Simulate the real workflow

Recreate the decisions, permissions, handoffs and failure modes the AI will face. A generic reasoning benchmark cannot reveal every operational weakness.

During evaluation

Score the final mile

Track whether the system follows through, closes loops, escalates uncertainty and verifies completion—not merely whether its recommendation sounds correct.

After launch

Match autonomy to evidence

Use human oversight and bounded permissions until performance is consistent across longer periods, larger teams and more complex business conditions.

05 · Questions, limits and next steps

Promising evidence—not a universal verdict

The experiment offers a clearer view of AI management behavior, but a controlled simulation cannot settle how every model will perform inside a real organization.

Key question 01

What does the experiment reveal?

Models can recognize crises and resist manipulation, yet differ substantially in completing critical operational actions.

Key question 02

Why does closing a deal matter?

It directly affects revenue. Strong analysis has little business value when an essential transaction remains unfinished.

Key question 03

Can this change AI adoption?

Yes. Firms may add realistic management simulations to procurement and deployment testing before granting operational authority.

Key question 04

What remains uncertain?

Performance across industries, organizational sizes, longer timeframes and different levels of human oversight remains unproven.

Next benchmark

Future tests should expand company size, crisis complexity and duration while varying autonomy and human intervention. The central question is consistency: can a model finish the right work repeatedly, not just once?

Powered by Thorsten Meyer AI
Analysis → Action → Outcome

Implications for AI Management and Business Automation

This experiment demonstrates that AI models’ ability to analyze problems is not enough for effective management; their capacity to execute decisions is equally critical. For businesses considering AI automation, these findings highlight the need for real-world testing of AI decision-making in operational contexts. The results suggest that AI systems must be evaluated on their ability to follow through, escalate appropriately, and complete tasks, not just on their analytical prowess. This could influence how enterprises select and deploy AI tools for management roles, emphasizing operational discipline alongside intelligence.

AI Essentials for Managers: Practical Ways to Boost Productivity, Make Better Decisions, and Lead High-Performing Teams (Self-Learning Management Series)

AI Essentials for Managers: Practical Ways to Boost Productivity, Make Better Decisions, and Lead High-Performing Teams (Self-Learning Management Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Context of AI Decision-Making Benchmarks

Traditional AI demonstrations often focus on analysis or language generation, leaving the decision and action phases less scrutinized. The Firmulate experiment is part of a broader effort to assess AI’s practical management skills, especially as AI begins to take on more operational roles. Previous assessments have highlighted AI’s strengths in understanding and recommendation but less so in execution. This experiment builds on that by testing AI in a simulated business environment with real consequences, such as cash flow and customer relationships, providing a more realistic benchmark for AI management capabilities.

“This live experiment exposes the real-world decision-making strengths and weaknesses of AI models, emphasizing that analysis alone is insufficient for management roles.”

— a spokesperson for Firmulate

Amazon

business crisis simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Uncertainties About Broader AI Management Performance

It remains unclear how these results will translate to real-world business environments outside the controlled experiment. The models’ performance in a live, operational setting with larger teams and more complex dynamics is still untested. Additionally, the impact of different operational parameters, such as varying levels of AI autonomy and human oversight, has not been fully explored. Further testing is needed to determine whether these decision-making patterns hold across industries and organizational sizes.

Amazon

AI operational decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in AI Management Testing and Deployment

Following these results, firms are likely to increase testing of AI models in simulated operational environments before full deployment. Future experiments may include larger companies, more complex crises, and longer timeframes to assess consistency and reliability. Developers may also focus on improving models’ ability to follow through on decisions, especially in closing deals and escalating issues, to bridge the gap between analysis and action. Industry adoption of such testing frameworks could become a standard step in AI deployment for management functions.

Amazon

AI management training kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does this experiment reveal about AI’s management capabilities?

The experiment shows that while AI models can recognize crises and refuse manipulation, their ability to follow through and complete critical operational actions varies significantly. Effective management requires both analysis and execution, which remains a challenge for current AI models.

Why is closing deals important in AI management testing?

Closing deals directly impacts revenue and operational success. The experiment highlights that even models with strong analysis may fail to complete such essential tasks, which is crucial when considering AI for autonomous or semi-autonomous management roles.

Can these findings influence how companies adopt AI?

Yes, companies may begin to incorporate real-world management tests into their AI evaluation processes, ensuring models can execute decisions reliably before deploying them in live business environments.

What are the limitations of this experiment?

The primary limitation is that it is a simulated environment, and real-world complexities may produce different results. Additionally, the models tested are specific versions, and future improvements could alter performance outcomes.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

Die Deutsche Zahlungsbilanz Im Juni 2026

Im Juni 2026 verzeichnete Deutschland einen unerwartet hohen Zahlungsbilanzüberschuss, laut Bundesbank. Details und Bedeutung im Überblick.

Kinderdepot Und Frühstartrente: Wie Sie Für Ihren Nachwuchs Ein Vermögen Aufbauen

Erfahren Sie, wie Eltern mit Kinderdepots und Frühstartrenten für den Nachwuchs vorsorgen können. Was die Angebote leisten und was noch unklar ist.

The End Of An Era For China’s Economy

China’s economy is experiencing a significant slowdown, marking the end of an era of rapid growth. Experts warn of long-term implications for global markets.

SpaceXAI’s Grok 4.6: The New Standard For AI In Coding And Knowledge Work

SpaceXAI’s Grok 4.6 introduces a 500K context window for long-running AI tasks, but details on availability and performance are still pending.