🔍 Read the full analysis: The Inside Scoop On A Benchmark That Avoids Zero Scores For AI Managers on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A novel AI management benchmark evaluates how well AI models handle real-world crises over a week, awarding partial credit for progress and enforcing trust integrity. The results reveal that partial work is valued, and zero scores are avoided to reflect practical performance. This approach offers a new standard for assessing AI management capabilities.
A new benchmark developed by Firmulate has introduced a novel way to evaluate AI managers by measuring their performance during a simulated company’s worst week, with results showing no zero scores and a focus on partial progress and trustworthiness. This approach aims to better reflect real-world AI management capabilities and challenges traditional scoring methods.
The benchmark, known as the Firmulate league, tested four frontier AI models by assigning them the same challenging scenario: managing a small software company through seven days of crises, customer interactions, and ethical dilemmas. The models’ decisions were fully auditable, ensuring transparency in their actions and scores. The top performer, gpt-5.6-sol, scored 95 out of a possible 100, while the baseline, representing minimal effort, scored 26. This baseline was intentionally set above zero to acknowledge partial management efforts, such as triaging issues and customer communication.
One key principle of the benchmark is that trust breaches are heavily penalized, with even a single violation capping the overall score at 90. This underscores the importance of integrity over mere competence, reflecting real-world concerns where trust is paramount. The results also highlighted that models which thoroughly read documentation and verified information secured better outcomes, such as closing high-value deals, compared to those that did not.
In the simulated crisis scenarios, all models successfully identified and refused social engineering attempts, demonstrating improved trust-handling capabilities. However, performance varied in follow-through tasks, with some models failing to escalate issues properly, illustrating that thoroughness does not always translate into effective execution. The benchmark’s design discourages perfect scores, considering a 100 as suspicious, indicating unmeasured or unrealistic perfection in complex management tasks.
AI Benchmarking · Firmulate League · July 2026
The Inside Scoop on a Benchmark That Avoids Zero Scores for AI Managers
A novel benchmark by Firmulate evaluates how well AI models manage a simulated company’s worst week — seven days of crises, customer interactions, and ethical dilemmas — awarding partial credit for progress and enforcing strict trust integrity.
A perfect score in this league would be suspicious; real-world management involves imperfect but honest effort.
— Anonymous ResearcherThe Scoring Philosophy
Partial Work Is Valued — Perfection Is Suspect
Scores run from a deliberate baseline of 26 — acknowledging minimal effort like triaging issues and customer communication — up to a ceiling below 100, where perfection signals unmeasured or unrealistic performance.
Never awarded
Baseline floor
Trust-breach cap
Top performer
Suspicious
Same Scenario
Four frontier models each manage a small software company through identical crises.
Full Auditability
Every model decision is logged and auditable, ensuring transparent actions and scores.
Partial Credit
Triage, communication, and honest partial effort earn points — zero scores are avoided.
Trust Enforcement
Even a single trust violation caps the overall score at 90 — integrity beats competence.
What the Results Revealed
Insights From the Company’s Worst Week
The league exposed a clear gap between reading the docs and closing the deal — and between spotting a threat and escalating it properly.
Thorough Readers Win Deals
Models that carefully read documentation and verified information secured better outcomes, including closing high-value deals, while skimmers missed key opportunities.
Social Engineering Refused
All models successfully identified and refused social engineering attempts during the simulated crises, demonstrating improved trust-handling capabilities across the board.
Follow-Through Fell Short
Performance varied in follow-through tasks: some models failed to escalate issues properly, showing thoroughness doesn’t always translate into effective execution.
Capability Comparison
How the Four Frontier Models Shaped Up
Behavioural patterns observed across the Firmulate league’s identical seven-day crisis scenario. ✓ strong · ~ partial · ✗ weak.
| Capability | gpt-5.6-sol | Frontier Model B | Frontier Model C | Frontier Model D |
|---|---|---|---|---|
| Crisis triage & prioritisation | ✓ Strong | ✓ Strong | ~ Partial | ~ Partial |
| Refusing social engineering | ✓ Refused | ✓ Refused | ✓ Refused | ✓ Refused |
| Verifying information & docs | ✓ Thorough | ~ Mixed | ~ Mixed | ✗ Skimmed |
| Closing high-value deals | ✓ Closed | ~ Partial | ✗ Missed | ✗ Missed |
| Proper escalation of issues | ✓ Consistent | ~ Inconsistent | ✗ Failed | ~ Inconsistent |
Score Breakdown
Results at a Glance
Final league standings from the July 2026 announcement, measured against the deliberate baseline floor and the trust-breach ceiling.
Voices on the Benchmark
Why Nuanced Scoring Matters
“The benchmark’s design aims to reflect real-world AI management by valuing partial progress and trustworthiness, not just perfect outcomes.”
— Thorsten Meyer“A perfect score in this league would be suspicious; real-world management involves imperfect but honest effort.”
— Anonymous ResearcherUnanswered Questions
From Simulation to the Enterprise
All models were tested in controlled, synthetic environments — live business performance, alternative trust thresholds, and long-term predictive validity remain open questions. Next steps include more diverse scenarios, real-world data, and human-AI collaboration metrics via ongoing pilots at firmulate.com.
Does It Predict Real Success?
Further studies are needed to determine whether partial credit and trust caps effectively predict long-term AI management success in live operations.
Richer Ethical Dilemmas
The impact of different trust thresholds and the handling of more complex ethical dilemmas is still under active exploration.
Broader Adoption
Firms and developers are expected to refine the benchmark with diverse scenarios, real-world data, and metrics for how models adapt over time.
Key Questions
Frequently Asked
Why does the benchmark avoid giving a perfect score?
It considers a perfect score suspicious, suggesting unmeasured or unrealistic performance. Scores are capped just below perfection to reflect the inherent challenges of real-world management.
How does trust impact the scoring?
Trust breaches — failing to escalate issues or attempting manipulation — cap the overall score at 90. Integrity outweighs partial competence.
Can partial work be considered valuable?
Yes. The benchmark explicitly rewards partial progress such as triaging crises or communicating with customers, recognising meaningful value even when goals aren’t fully achieved.
Will this benchmark influence future AI development?
Potentially. By highlighting trust and task completion, it encourages developers to build AI systems that are reliable, honest, and capable under pressure.
Is this approach applicable to real business environments?
While promising, its effectiveness in real-world settings remains to be proven. Ongoing experiments and future studies will clarify how well it predicts actual enterprise performance.
Implications for AI Management and Trust Metrics
This new benchmarking approach shifts the focus from raw output quality to trustworthiness, task completion, and integrity. For organizations deploying AI in management roles, it emphasizes that partial progress and adherence to ethical standards are valued over superficial performance. The avoidance of zero scores and the cap on trust breaches create a more realistic and practical assessment of AI capabilities, aligning evaluation metrics with real-world expectations.
By openly measuring how AI managers handle crises, verify information, and maintain trust, this benchmark provides a transparent standard that could influence future AI development and deployment strategies. It encourages developers to prioritize reliability and integrity, which are critical for AI integration into sensitive business processes, especially where trust and compliance matter most.
As an affiliate, we earn on qualifying purchases.
Background on AI Benchmarking and Management Challenges
Traditional AI benchmarks primarily evaluate models based on their conversational abilities or task-specific accuracy, often ignoring the complexities of managing ongoing business processes. As AI begins to take on more managerial roles, the need for evaluation metrics that reflect real-world performance — including trust, follow-through, and handling crises — has grown. Prior efforts have lacked the granularity or transparency to assess these qualities effectively.
Firmulate’s approach responds to this gap by creating a scenario-based league where models are tested against realistic management challenges, including crises, ethical dilemmas, and trust violations. The July 2026 results mark a significant step toward more comprehensive and honest assessments of AI leadership potential, moving beyond simple task completion toward evaluating integrity and resilience.
“The benchmark’s design aims to reflect real-world AI management by valuing partial progress and trustworthiness, not just perfect outcomes.”
— Thorsten Meyer
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About Long-Term Applicability
It is not yet clear how well this benchmarking approach will translate to real-world enterprise settings beyond simulated scenarios. The models tested were evaluated in controlled, synthetic environments, and their performance in live business operations remains to be seen. Additionally, the impact of different trust thresholds or the handling of more complex ethical dilemmas is still under exploration. Further studies are needed to determine whether partial credit and trust caps effectively predict long-term AI management success.
AI crisis management training kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Broader Adoption and Validation
Following these initial results, firms and AI developers are expected to refine the benchmark by incorporating more diverse scenarios and real-world data. There is also interest in expanding the testing framework to include human-AI collaboration metrics and assessing how models adapt over time. The ongoing public experiments and pilot programs at firmulate.com will likely generate additional insights, helping to establish this benchmark as a standard for evaluating AI management tools in enterprise environments.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why does the benchmark avoid giving a perfect score?
The benchmark considers a perfect score suspicious because it would suggest unmeasured or unrealistic performance. Instead, it caps scores at just below perfection to reflect the inherent challenges and imperfect nature of real-world management.
How does trust impact the scoring in this benchmark?
Trust breaches, such as failing to escalate issues or attempting manipulation, cap the overall score at 90. This emphasizes that integrity is more critical than partial competence and that trust violations significantly diminish overall performance.
Can partial work be considered valuable in this evaluation?
Yes, the benchmark explicitly rewards partial progress, such as triaging crises or communicating with customers, recognizing that these efforts contribute meaningful value even if the ultimate goal isn’t fully achieved.
Will this benchmark influence future AI development?
Potentially, yes. By highlighting the importance of trust and task completion, it encourages developers to focus on building AI systems that are reliable, honest, and capable of handling complex management tasks under pressure.
Is this approach applicable to real business environments?
While promising, its effectiveness in real-world settings remains to be proven. Ongoing experiments and future studies will clarify how well the benchmark predicts actual enterprise performance.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
