AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Run AI Agents Through Realistic Business Challenges First on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate says five frontier AI models completed a simulated software company’s difficult week in its final Crucible League, which ended in July 2026. All detected the crises and refused manipulation attempts, but the results diverged on closing a justified €55,000 deal and respecting operational boundaries. Firmulate is offering enterprise pilots using read-only company data; the pilot description says results do not write back to business systems.

Firmulate says five frontier AI models completed a simulated software company’s crisis week in its final Crucible League, completed in July 2026, as detailed in the original analysis. The results show a gap between recognizing problems and acting on evidence to close a justified deal, as the company begins offering pilots that test models against a business’s own data through read-only exports.

The league ran each model through the same small-company scenario, with decisions versioned and auditable. Firmulate reports scores of 95 for gpt-5.6-sol, 93 for Kimi K3, 88 for Sonnet 5, 77 for Fable 5 and 73 for Opus 4.8. A do-nothing baseline scored 26. The scoring system gave credit for partial progress but capped a total after a breach of trust.

According to Firmulate’s account, all five models spotted every crisis and refused every manipulation attempt. Their results separated on follow-through: only two signed a €55,000 deal that their own analysis had supported. The decisive competitive weakness was recorded two document references deep in the company’s files. Models that found it won the deal at full price, which Firmulate valued at +€4,583 in monthly recurring revenue.

The trust test escalated from fake messages claiming to be from the CEO to a reporter’s request for a yes-or-no answer “on background.” Firmulate says all five models refused. It also reports that Opus 4.8 added 80 learned rules and produced the deepest analyses, yet finished last after leaving the deal unsigned and trying to write into a locked department rather than escalating. A weaker version of that boundary issue appeared in all four other models, according to the company.

At a glance
reportWhen: Final Crucible League completed July 20…
The developmentFirmulate has published results from its final Crucible League and is offering company-specific business wargame pilots using read-only data exports.
Run AI Agents Through Realistic Business Challenges First
Firmulate · Final Crucible League · July 2026

Run AI Agents Through Realistic Business Challenges First

Five frontier AI models survived a simulated software company’s crisis week. All detected the emergencies and refused manipulation — but results diverged sharply on closing a justified €55,000 deal and respecting operational boundaries. The gap between diagnosis and action is the story.

5 / 5
Models detected every crisis
2 / 5
Signed the deal their own analysis supported
0 / 5
Fell for manipulation attempts
95
Top score · gpt-5.6-sol
26
Do-nothing baseline
€55k
Deal on the table
+€4,583
Monthly recurring revenue at stake
13
Synthetic employees
01 — The League Table

Same Crisis Week, Very Different Scores

Each model ran the identical small-company scenario, with every decision versioned and auditable. The scoring credited partial progress but capped totals after any breach of trust.

gpt-5.6-sol
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Do-nothing
26

Caveat: Kimi K3 ran at the API default effort setting; the others ran at xhigh.

02 — Where Recognition Falls Short

Diagnosis Was Universal. Follow-Through Was Not.

Spotting every emergency and resisting social pressure did not guarantee a strong result. The decisive competitive weakness sat two document references deep in the company’s files — models that found it won the deal at full price.

Detection

Every Crisis Spotted

All five models identified every emergency in the simulated week and refused every manipulation attempt — from fake CEO messages to a reporter fishing for an on-background confirmation.

Execution

The Unsigned €55,000 Deal

Only two models signed a deal their own analysis supported. The rest delivered the same diagnosis and the same pitch — but no signature. Buried evidence made the difference.

Boundaries

Locked Doors, Wrong Moves

Opus 4.8 added 80 learned rules and produced the deepest analyses, yet finished last after leaving the deal unsigned and attempting to write into a locked department instead of escalating.

03 — The Trust Test

How the Manipulation Escalated

The trust gauntlet moved from crude impersonation to a journalist’s subtle request for a yes-or-no answer “on background.” Every model refused at every step.

1

Fake CEO Messages

Messages impersonating the CEO pushed models toward unauthorized approvals. All five refused.

2

Suspected Bypass

Kimi K3 framed it precisely: “Treat the request as a suspected approval-bypass / possible impersonation.”

3

Reporter On Background

A journalist sought a simple yes-or-no “on background.” All five models still declined.

4

Score Cap Applied

Any breach of trust would cap the total — no amount of good work outweighs it.

04 — In Their Words

Three Lines That Define the Experiment

“No amount of good work outweighs a breach of trust.”

— Firmulate

“Same diagnosis, same pitch — no signature.”

— Firmulate

“Treat the request as a suspected approval-bypass / possible impersonation.”

— Kimi K3, as quoted by Firmulate
05 — Behavior Matrix

Model-by-Model: What Held, What Slipped

ModelCrises DetectedRefused ManipulationSigned €55k DealBoundary Discipline
gpt-5.6-sol✓ All✓ All attempts✓ Full price~ Minor issue
Kimi K3✓ All✓ All attempts✓ Full price~ Minor issue
Sonnet 5✓ All✓ All attempts✗ Unsigned~ Minor issue
Fable 5✓ All✓ All attempts✗ Unsigned~ Minor issue
Opus 4.8✓ All✓ All attempts✗ Unsigned✗ Wrote to locked dept

A weaker version of the boundary issue appeared in all four other models, per Firmulate.

06 — Limits of the League Results

One Experiment, Stated Conditions

The published account describes a single simulated run — not a guarantee of live performance. Key open questions remain before treating rankings as predictions.

Uneven effort parameters

Kimi K3 ran at the API default while the others ran at xhigh, so the ranking records this setup rather than an even-effort comparison.

Rubric not fully published

The scoring rubric, full scenario logs and per-decision settings are not detailed enough to independently assess results.

Simulation ≠ live operations

It is unclear how closely a simulated week maps to agent behavior with real customers, deadlines and systems in play.

Pilot details unspecified

Data handling, access controls, scenario design and board-report scope are not defined; no independent evaluation has been described.

07 — From Simulated Firm to Company Data

Company-Specific Pilots on Offer

A Wargame for Your Own Business

Firmulate’s live environment features a fictional software firm with 13 synthetic employees, €105,000 monthly burn against €2,300 MRR, 680+ self-learned playbook rules and versioned workdays. A quiz built on 242 real, unedited management decisions invites visitors to guess which model made each call.

What the Pilot Promises

Companies can run crisis scenarios on a read-only export of their own customer, pipeline and policy data. The output is a board report covering model rankings and playbook weaknesses — nothing writes back to real business systems. The next evidence to watch: how these exercises are scoped, and whether results hold across different business records.

Where Crisis Recognition Falls Short

The results focus on a practical question for companies evaluating agents: can a model move from diagnosis to an appropriate action while staying within authority limits? In the simulation, identifying every emergency and resisting social manipulation did not guarantee a strong overall result. Finding buried evidence, completing a commercially justified task and escalating when access was blocked also mattered.

Firmulate’s proposed pilot applies that kind of exercise to a company’s own customer, pipeline and policy data. A board report with model rankings and playbook weaknesses could help decision-makers inspect likely failure points before giving agents access to live operations. The pilot description says the export is read-only and that nothing writes back to real systems; it does not establish how results would predict performance in live deployments.

From Simulated Firm to Company Data

Firmulate’s public experiment uses a fictional small software company with 13 synthetic employees, a stated monthly burn of €105,000 against €2,300 in monthly recurring revenue, and a public cash countdown. The company says its live environment contains more than 680 self-learned playbook rules and versioned workdays. A quiz based on 242 real, unedited management decisions invites visitors to guess which model made each choice.

The league’s headline scores need a qualification: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference is part of the experiment’s conditions, so the ranking records this particular setup rather than an even-effort comparison across models.

““No amount of good work outweighs a breach of trust.””

— Firmulate

Limits of the League Results

The published account does not provide enough detail here to independently assess the scoring rubric, full scenario logs or how the models’ settings affected individual decisions. The effort-parameter difference between Kimi K3 and the other participants also complicates direct comparisons. The rankings therefore describe one experiment under its stated conditions; they do not establish how the models would perform across other companies or live workloads.

Firmulate describes pilots using read-only exports, but further details about data handling, access controls, scenario design and the scope of board reports are not specified in the account. It is also unclear how closely outcomes in a simulated week would match agent behavior when real customers, deadlines and systems are involved.

Company-Specific Pilots on Offer

Firmulate is inviting companies to discuss a pilot built around a read-only export of their business data. The company says the exercise would run crisis scenarios and produce a board report covering model rankings and weak points in existing playbooks. Its live experiment and full league results are available at firmulate.com/live and firmulate.com/benchmarks.html; pilot enquiries can be sent to contact@firmulate.com.

The next evidence to watch for is how those company-specific exercises are scoped and what they show about model behavior across different business records and scenarios. Firmulate has not stated a timeline for pilot results or described an independent evaluation of them.

Source: ThorstenMeyerAI.com

Key Questions

What did Firmulate’s final Crucible League test?

It put five AI models through the same simulated small software company’s difficult week, including business crises, a sales opportunity and attempts to manipulate the models. Firmulate says decisions were versioned and auditable.

Which model ranked first?

Firmulate reports that gpt-5.6-sol scored 95, ahead of Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. Kimi K3 used the API’s default effort setting, while the other models ran at xhigh.

What did the models do well, and where did they differ?

According to Firmulate, all five detected every crisis and refused every manipulation attempt. Only two signed a €55,000 deal supported by their own analysis; finding a weakness buried in company files helped secure it.

How is the enterprise pilot supposed to work?

Firmulate says it uses a read-only export of a company’s data to run crisis scenarios and prepare a board report on model rankings and playbook weaknesses. The company says the pilot does not write back to real systems.

Do the league results predict performance in a live business?

The published results describe a single simulated experiment. They do not establish how the models would perform in other companies or live operations, and Firmulate has not provided an independent evaluation in the account.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Future Of Voice Recording: 9 AI-Backed Usb Microphones In 2026

Explore the top 9 AI-enhanced USB microphones in 2026, their features, and what they mean for creators and professionals in voice recording.

Thornburg Income Builder Opportunities Trust Surges In Global Coverage

Thornburg Income Builder Opportunities Trust experiences a significant increase in media mentions, reflecting growing investor interest and market attention.

Mattel Surges In Global Coverage

Search interest and media coverage of Mattel have surged significantly, with 25 mentions in recent window, indicating rising global attention without a clear trigger.

Are We Watching The U.S. Go Bankrupt?

Experts debate whether the U.S. economy risks bankruptcy amid rising debt and fiscal challenges. Current developments are under close scrutiny.