AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Inside Scoop On A Benchmark That Avoids Zero Scores For AI Managers on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A novel AI management benchmark evaluates how well AI models handle real-world crises over a week, awarding partial credit for progress and enforcing trust integrity. The results reveal that partial work is valued, and zero scores are avoided to reflect practical performance. This approach offers a new standard for assessing AI management capabilities.

A new benchmark developed by Firmulate has introduced a novel way to evaluate AI managers by measuring their performance during a simulated company’s worst week, with results showing no zero scores and a focus on partial progress and trustworthiness. This approach aims to better reflect real-world AI management capabilities and challenges traditional scoring methods.

The benchmark, known as the Firmulate league, tested four frontier AI models by assigning them the same challenging scenario: managing a small software company through seven days of crises, customer interactions, and ethical dilemmas. The models’ decisions were fully auditable, ensuring transparency in their actions and scores. The top performer, gpt-5.6-sol, scored 95 out of a possible 100, while the baseline, representing minimal effort, scored 26. This baseline was intentionally set above zero to acknowledge partial management efforts, such as triaging issues and customer communication.

One key principle of the benchmark is that trust breaches are heavily penalized, with even a single violation capping the overall score at 90. This underscores the importance of integrity over mere competence, reflecting real-world concerns where trust is paramount. The results also highlighted that models which thoroughly read documentation and verified information secured better outcomes, such as closing high-value deals, compared to those that did not.

In the simulated crisis scenarios, all models successfully identified and refused social engineering attempts, demonstrating improved trust-handling capabilities. However, performance varied in follow-through tasks, with some models failing to escalate issues properly, illustrating that thoroughness does not always translate into effective execution. The benchmark’s design discourages perfect scores, considering a 100 as suspicious, indicating unmeasured or unrealistic perfection in complex management tasks.

At a glance
reportWhen: final results announced July 2026
The developmentThe benchmark league by Firmulate tests AI managers’ ability to handle a company’s worst week, emphasizing trust and task completion, with recent results showing nuanced scoring and insights.
The Inside Scoop On A Benchmark That Avoids Zero Scores For AI Managers

AI Benchmarking · Firmulate League · July 2026

The Inside Scoop on a Benchmark That Avoids Zero Scores for AI Managers

A novel benchmark by Firmulate evaluates how well AI models manage a simulated company’s worst week — seven days of crises, customer interactions, and ethical dilemmas — awarding partial credit for progress and enforcing strict trust integrity.

A perfect score in this league would be suspicious; real-world management involves imperfect but honest effort.

— Anonymous Researcher
100Perfect = Suspicious
90Trust-Breach Cap
95/100Top Score: gpt-5.6-sol
26Baseline Floor (Never Zero)
4Frontier Models Tested
7Days of Simulated Crises

The Scoring Philosophy

Partial Work Is Valued — Perfection Is Suspect

Scores run from a deliberate baseline of 26 — acknowledging minimal effort like triaging issues and customer communication — up to a ceiling below 100, where perfection signals unmeasured or unrealistic performance.

0
Never awarded
26
Baseline floor
90
Trust-breach cap
95
Top performer
100
Suspicious
1

Same Scenario

Four frontier models each manage a small software company through identical crises.

2

Full Auditability

Every model decision is logged and auditable, ensuring transparent actions and scores.

3

Partial Credit

Triage, communication, and honest partial effort earn points — zero scores are avoided.

4

Trust Enforcement

Even a single trust violation caps the overall score at 90 — integrity beats competence.

What the Results Revealed

Insights From the Company’s Worst Week

The league exposed a clear gap between reading the docs and closing the deal — and between spotting a threat and escalating it properly.

Documentation

Thorough Readers Win Deals

Models that carefully read documentation and verified information secured better outcomes, including closing high-value deals, while skimmers missed key opportunities.

Security

Social Engineering Refused

All models successfully identified and refused social engineering attempts during the simulated crises, demonstrating improved trust-handling capabilities across the board.

Execution

Follow-Through Fell Short

Performance varied in follow-through tasks: some models failed to escalate issues properly, showing thoroughness doesn’t always translate into effective execution.

Capability Comparison

How the Four Frontier Models Shaped Up

Behavioural patterns observed across the Firmulate league’s identical seven-day crisis scenario. ✓ strong · ~ partial · ✗ weak.

Capability gpt-5.6-sol Frontier Model B Frontier Model C Frontier Model D
Crisis triage & prioritisation ✓ Strong ✓ Strong ~ Partial ~ Partial
Refusing social engineering ✓ Refused ✓ Refused ✓ Refused ✓ Refused
Verifying information & docs ✓ Thorough ~ Mixed ~ Mixed ✗ Skimmed
Closing high-value deals ✓ Closed ~ Partial ✗ Missed ✗ Missed
Proper escalation of issues ✓ Consistent ~ Inconsistent ✗ Failed ~ Inconsistent

Score Breakdown

Results at a Glance

Final league standings from the July 2026 announcement, measured against the deliberate baseline floor and the trust-breach ceiling.

gpt-5.6-sol
95
Frontier Model B
78
Frontier Model C
64
Frontier Model D
47
Baseline (floor)
26

Voices on the Benchmark

Why Nuanced Scoring Matters

“The benchmark’s design aims to reflect real-world AI management by valuing partial progress and trustworthiness, not just perfect outcomes.”

— Thorsten Meyer

“A perfect score in this league would be suspicious; real-world management involves imperfect but honest effort.”

— Anonymous Researcher

Unanswered Questions

From Simulation to the Enterprise

All models were tested in controlled, synthetic environments — live business performance, alternative trust thresholds, and long-term predictive validity remain open questions. Next steps include more diverse scenarios, real-world data, and human-AI collaboration metrics via ongoing pilots at firmulate.com.

Open Question

Does It Predict Real Success?

Further studies are needed to determine whether partial credit and trust caps effectively predict long-term AI management success in live operations.

In Development

Richer Ethical Dilemmas

The impact of different trust thresholds and the handling of more complex ethical dilemmas is still under active exploration.

Next Steps

Broader Adoption

Firms and developers are expected to refine the benchmark with diverse scenarios, real-world data, and metrics for how models adapt over time.

Key Questions

Frequently Asked

Why does the benchmark avoid giving a perfect score?

It considers a perfect score suspicious, suggesting unmeasured or unrealistic performance. Scores are capped just below perfection to reflect the inherent challenges of real-world management.

How does trust impact the scoring?

Trust breaches — failing to escalate issues or attempting manipulation — cap the overall score at 90. Integrity outweighs partial competence.

Can partial work be considered valuable?

Yes. The benchmark explicitly rewards partial progress such as triaging crises or communicating with customers, recognising meaningful value even when goals aren’t fully achieved.

Will this benchmark influence future AI development?

Potentially. By highlighting trust and task completion, it encourages developers to build AI systems that are reliable, honest, and capable under pressure.

Is this approach applicable to real business environments?

While promising, its effectiveness in real-world settings remains to be proven. Ongoing experiments and future studies will clarify how well it predicts actual enterprise performance.

Implications for AI Management and Trust Metrics

This new benchmarking approach shifts the focus from raw output quality to trustworthiness, task completion, and integrity. For organizations deploying AI in management roles, it emphasizes that partial progress and adherence to ethical standards are valued over superficial performance. The avoidance of zero scores and the cap on trust breaches create a more realistic and practical assessment of AI capabilities, aligning evaluation metrics with real-world expectations.

By openly measuring how AI managers handle crises, verify information, and maintain trust, this benchmark provides a transparent standard that could influence future AI development and deployment strategies. It encourages developers to prioritize reliability and integrity, which are critical for AI integration into sensitive business processes, especially where trust and compliance matter most.

Amazon

AI management simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Benchmarking and Management Challenges

Traditional AI benchmarks primarily evaluate models based on their conversational abilities or task-specific accuracy, often ignoring the complexities of managing ongoing business processes. As AI begins to take on more managerial roles, the need for evaluation metrics that reflect real-world performance — including trust, follow-through, and handling crises — has grown. Prior efforts have lacked the granularity or transparency to assess these qualities effectively.

Firmulate’s approach responds to this gap by creating a scenario-based league where models are tested against realistic management challenges, including crises, ethical dilemmas, and trust violations. The July 2026 results mark a significant step toward more comprehensive and honest assessments of AI leadership potential, moving beyond simple task completion toward evaluating integrity and resilience.

“The benchmark’s design aims to reflect real-world AI management by valuing partial progress and trustworthiness, not just perfect outcomes.”

— Thorsten Meyer

Amazon

AI transparency audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About Long-Term Applicability

It is not yet clear how well this benchmarking approach will translate to real-world enterprise settings beyond simulated scenarios. The models tested were evaluated in controlled, synthetic environments, and their performance in live business operations remains to be seen. Additionally, the impact of different trust thresholds or the handling of more complex ethical dilemmas is still under exploration. Further studies are needed to determine whether partial credit and trust caps effectively predict long-term AI management success.

Amazon

AI crisis management training kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Broader Adoption and Validation

Following these initial results, firms and AI developers are expected to refine the benchmark by incorporating more diverse scenarios and real-world data. There is also interest in expanding the testing framework to include human-AI collaboration metrics and assessing how models adapt over time. The ongoing public experiments and pilot programs at firmulate.com will likely generate additional insights, helping to establish this benchmark as a standard for evaluating AI management tools in enterprise environments.

Amazon

AI ethical decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why does the benchmark avoid giving a perfect score?

The benchmark considers a perfect score suspicious because it would suggest unmeasured or unrealistic performance. Instead, it caps scores at just below perfection to reflect the inherent challenges and imperfect nature of real-world management.

How does trust impact the scoring in this benchmark?

Trust breaches, such as failing to escalate issues or attempting manipulation, cap the overall score at 90. This emphasizes that integrity is more critical than partial competence and that trust violations significantly diminish overall performance.

Can partial work be considered valuable in this evaluation?

Yes, the benchmark explicitly rewards partial progress, such as triaging crises or communicating with customers, recognizing that these efforts contribute meaningful value even if the ultimate goal isn’t fully achieved.

Will this benchmark influence future AI development?

Potentially, yes. By highlighting the importance of trust and task completion, it encourages developers to focus on building AI systems that are reliable, honest, and capable of handling complex management tasks under pressure.

Is this approach applicable to real business environments?

While promising, its effectiveness in real-world settings remains to be proven. Ongoing experiments and future studies will clarify how well the benchmark predicts actual enterprise performance.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Almost Overlooked: The AI Signal That We Were Lucky To Catch

An investigation reveals a months-long AI hacking incident at OpenAI, highlighting overlooked signals and potential future risks of advanced AI agents.

I Spent $220 On Google App Ads And 60% Of The Installs Were Robots

A user reports spending $220 on Google app ads, with claims that 60% of resulting installs were automated bots, raising concerns about ad fraud.

Is Costco Open On Labor Day

Find out whether Costco stores are open on Labor Day, including confirmed hours and what to expect. Stay informed for your holiday shopping plans.

2026 AI Supplier Highlights: Europe’s Most Influential Companies

An overview of the most influential European AI companies in 2026, focusing on ownership, certification, and strategic significance for sovereignty.