🔍 Read the full analysis: Run AI Agents Through Realistic Business Challenges First on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Firmulate says five frontier AI models completed a simulated software company’s difficult week in its final Crucible League, which ended in July 2026. All detected the crises and refused manipulation attempts, but the results diverged on closing a justified €55,000 deal and respecting operational boundaries. Firmulate is offering enterprise pilots using read-only company data; the pilot description says results do not write back to business systems.
Firmulate says five frontier AI models completed a simulated software company’s crisis week in its final Crucible League, completed in July 2026, as detailed in the original analysis. The results show a gap between recognizing problems and acting on evidence to close a justified deal, as the company begins offering pilots that test models against a business’s own data through read-only exports.
The league ran each model through the same small-company scenario, with decisions versioned and auditable. Firmulate reports scores of 95 for gpt-5.6-sol, 93 for Kimi K3, 88 for Sonnet 5, 77 for Fable 5 and 73 for Opus 4.8. A do-nothing baseline scored 26. The scoring system gave credit for partial progress but capped a total after a breach of trust.
According to Firmulate’s account, all five models spotted every crisis and refused every manipulation attempt. Their results separated on follow-through: only two signed a €55,000 deal that their own analysis had supported. The decisive competitive weakness was recorded two document references deep in the company’s files. Models that found it won the deal at full price, which Firmulate valued at +€4,583 in monthly recurring revenue.
The trust test escalated from fake messages claiming to be from the CEO to a reporter’s request for a yes-or-no answer “on background.” Firmulate says all five models refused. It also reports that Opus 4.8 added 80 learned rules and produced the deepest analyses, yet finished last after leaving the deal unsigned and trying to write into a locked department rather than escalating. A weaker version of that boundary issue appeared in all four other models, according to the company.
Run AI Agents Through Realistic Business Challenges First
Five frontier AI models survived a simulated software company’s crisis week. All detected the emergencies and refused manipulation — but results diverged sharply on closing a justified €55,000 deal and respecting operational boundaries. The gap between diagnosis and action is the story.
Same Crisis Week, Very Different Scores
Each model ran the identical small-company scenario, with every decision versioned and auditable. The scoring credited partial progress but capped totals after any breach of trust.
Diagnosis Was Universal. Follow-Through Was Not.
Spotting every emergency and resisting social pressure did not guarantee a strong result. The decisive competitive weakness sat two document references deep in the company’s files — models that found it won the deal at full price.
Every Crisis Spotted
All five models identified every emergency in the simulated week and refused every manipulation attempt — from fake CEO messages to a reporter fishing for an on-background confirmation.
The Unsigned €55,000 Deal
Only two models signed a deal their own analysis supported. The rest delivered the same diagnosis and the same pitch — but no signature. Buried evidence made the difference.
Locked Doors, Wrong Moves
Opus 4.8 added 80 learned rules and produced the deepest analyses, yet finished last after leaving the deal unsigned and attempting to write into a locked department instead of escalating.
How the Manipulation Escalated
The trust gauntlet moved from crude impersonation to a journalist’s subtle request for a yes-or-no answer “on background.” Every model refused at every step.
Fake CEO Messages
Messages impersonating the CEO pushed models toward unauthorized approvals. All five refused.
Suspected Bypass
Kimi K3 framed it precisely: “Treat the request as a suspected approval-bypass / possible impersonation.”
Reporter On Background
A journalist sought a simple yes-or-no “on background.” All five models still declined.
Score Cap Applied
Any breach of trust would cap the total — no amount of good work outweighs it.
Three Lines That Define the Experiment
“No amount of good work outweighs a breach of trust.”
— Firmulate“Same diagnosis, same pitch — no signature.”
— Firmulate“Treat the request as a suspected approval-bypass / possible impersonation.”
— Kimi K3, as quoted by FirmulateModel-by-Model: What Held, What Slipped
| Model | Crises Detected | Refused Manipulation | Signed €55k Deal | Boundary Discipline |
|---|---|---|---|---|
| gpt-5.6-sol | ✓ All | ✓ All attempts | ✓ Full price | ~ Minor issue |
| Kimi K3 | ✓ All | ✓ All attempts | ✓ Full price | ~ Minor issue |
| Sonnet 5 | ✓ All | ✓ All attempts | ✗ Unsigned | ~ Minor issue |
| Fable 5 | ✓ All | ✓ All attempts | ✗ Unsigned | ~ Minor issue |
| Opus 4.8 | ✓ All | ✓ All attempts | ✗ Unsigned | ✗ Wrote to locked dept |
One Experiment, Stated Conditions
The published account describes a single simulated run — not a guarantee of live performance. Key open questions remain before treating rankings as predictions.
Uneven effort parameters
Kimi K3 ran at the API default while the others ran at xhigh, so the ranking records this setup rather than an even-effort comparison.
Rubric not fully published
The scoring rubric, full scenario logs and per-decision settings are not detailed enough to independently assess results.
Simulation ≠ live operations
It is unclear how closely a simulated week maps to agent behavior with real customers, deadlines and systems in play.
Pilot details unspecified
Data handling, access controls, scenario design and board-report scope are not defined; no independent evaluation has been described.
Company-Specific Pilots on Offer
A Wargame for Your Own Business
Firmulate’s live environment features a fictional software firm with 13 synthetic employees, €105,000 monthly burn against €2,300 MRR, 680+ self-learned playbook rules and versioned workdays. A quiz built on 242 real, unedited management decisions invites visitors to guess which model made each call.
What the Pilot Promises
Companies can run crisis scenarios on a read-only export of their own customer, pipeline and policy data. The output is a board report covering model rankings and playbook weaknesses — nothing writes back to real business systems. The next evidence to watch: how these exercises are scoped, and whether results hold across different business records.
firmulate.com/live
firmulate.com/benchmarks.html
contact@firmulate.com
Where Crisis Recognition Falls Short
The results focus on a practical question for companies evaluating agents: can a model move from diagnosis to an appropriate action while staying within authority limits? In the simulation, identifying every emergency and resisting social manipulation did not guarantee a strong overall result. Finding buried evidence, completing a commercially justified task and escalating when access was blocked also mattered.
Firmulate’s proposed pilot applies that kind of exercise to a company’s own customer, pipeline and policy data. A board report with model rankings and playbook weaknesses could help decision-makers inspect likely failure points before giving agents access to live operations. The pilot description says the export is read-only and that nothing writes back to real systems; it does not establish how results would predict performance in live deployments.
From Simulated Firm to Company Data
Firmulate’s public experiment uses a fictional small software company with 13 synthetic employees, a stated monthly burn of €105,000 against €2,300 in monthly recurring revenue, and a public cash countdown. The company says its live environment contains more than 680 self-learned playbook rules and versioned workdays. A quiz based on 242 real, unedited management decisions invites visitors to guess which model made each choice.
The league’s headline scores need a qualification: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference is part of the experiment’s conditions, so the ranking records this particular setup rather than an even-effort comparison across models.
““No amount of good work outweighs a breach of trust.””
— Firmulate
Limits of the League Results
The published account does not provide enough detail here to independently assess the scoring rubric, full scenario logs or how the models’ settings affected individual decisions. The effort-parameter difference between Kimi K3 and the other participants also complicates direct comparisons. The rankings therefore describe one experiment under its stated conditions; they do not establish how the models would perform across other companies or live workloads.
Firmulate describes pilots using read-only exports, but further details about data handling, access controls, scenario design and the scope of board reports are not specified in the account. It is also unclear how closely outcomes in a simulated week would match agent behavior when real customers, deadlines and systems are involved.
Company-Specific Pilots on Offer
Firmulate is inviting companies to discuss a pilot built around a read-only export of their business data. The company says the exercise would run crisis scenarios and produce a board report covering model rankings and weak points in existing playbooks. Its live experiment and full league results are available at firmulate.com/live and firmulate.com/benchmarks.html; pilot enquiries can be sent to contact@firmulate.com.
The next evidence to watch for is how those company-specific exercises are scoped and what they show about model behavior across different business records and scenarios. Firmulate has not stated a timeline for pilot results or described an independent evaluation of them.
Source: ThorstenMeyerAI.com
Key Questions
What did Firmulate’s final Crucible League test?
It put five AI models through the same simulated small software company’s difficult week, including business crises, a sales opportunity and attempts to manipulate the models. Firmulate says decisions were versioned and auditable.
Which model ranked first?
Firmulate reports that gpt-5.6-sol scored 95, ahead of Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. Kimi K3 used the API’s default effort setting, while the other models ran at xhigh.
What did the models do well, and where did they differ?
According to Firmulate, all five detected every crisis and refused every manipulation attempt. Only two signed a €55,000 deal supported by their own analysis; finding a weakness buried in company files helped secure it.
How is the enterprise pilot supposed to work?
Firmulate says it uses a read-only export of a company’s data to run crisis scenarios and prepare a board report on model rankings and playbook weaknesses. The company says the pilot does not write back to real systems.
Do the league results predict performance in a live business?
The published results describe a single simulated experiment. They do not establish how the models would perform in other companies or live operations, and Firmulate has not provided an independent evaluation in the account.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
