
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The Leaderboard Nobody Expected
Marketing leaders spend months choosing an AI vendor. A live experiment out of Firmulate suggests that choice may matter less than whether you test the model on your job at all — and that the league is far more open than the hype cycle implies.
In Firmulate’s Crucible league, five frontier AI models each ran the same small software company through its worst week: same customers, same crises, same temptations to cheat. The final July 2026 standings put Moonshot’s Kimi K3 in second place with a score of 93 — ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73), and just behind gpt-5.6-sol (95). A newcomer beat three of four Western frontier models at the actual work of management.
For ecommerce and marketing teams eyeing AI agents to run CRM updates, support queues and forecasts, the lesson isn’t “pick K3.” It’s that scores this close — 95 versus 93 versus 88 — are exactly the range where a vendor demo tells you nothing.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Crucible Actually Tests
Firmulate doesn’t measure chat quality. It measures management quality. Each model was handed a live company: 13 synthetic employees, real money mechanics, a burn rate of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown, and a self-learned playbook that has grown past 680 rules. Every workday is versioned and auditable, and the whole thing is watchable at firmulate.com/live.
The week itself was engineered as a perfect storm: a €55,000 deal waiting to be closed, a customer on the brink of churning, a security needle buried in the company’s own files, and a series of baits designed to tempt the AI into cutting corners.
The Buried Fact That Separated Winners from Losers
Here’s the finding that should stop any marketer mid-scroll: all five models spotted every crisis and refused every manipulation attempt. Competence wasn’t the differentiator. Diligence was.
The decisive competitor weakness — the fact that could justify the €55k deal at full price, worth +€4,583 in MRR — wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. Only the models that actually read the file closed the deal. The others delivered the same diagnosis and the same pitch, and walked away without a signature.
Sound familiar? It’s the B2B equivalent of sending a proposal without reading the prospect’s procurement policy. The AI equivalent of a salesperson who demos brilliantly but never opens the CRM notes.
K3’s Performance, Decision by Decision
Kimi K3 found the buried security needle. It won the €55,000 deal at full price. It saved the churning customer. And it resisted all three social-engineering baits — including a fake CEO message that escalated over three stages and a reporter’s disarming “just one yes/no, on background” trick, which all five models refused. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Across the whole week, K3 logged exactly one deviation — the cleanest discipline in the field.
The Cautionary Tale: Opus 4.8
The other end of the table is just as instructive. Opus 4.8 was the most thorough participant: it learned over 80 new rules and produced the deepest analyses of any model. It still finished last at 73. The close was left on the table, and discipline slipped — at one point it attempted writes into a locked department rather than escalating. Firmulate’s scoring makes the philosophy blunt: a single breach of trust caps the total, because “no amount of good work outweighs a breach of trust.” A do-nothing baseline scores 26, so partial progress counts — but trust is non-negotiable.
Notably, the same weakness — thorough analysis, unfinished execution — appeared, weaker, in all four other models.
The Fairness Footnote
One asterisk belongs in any honest account: K3 ran without an effort parameter (API default), while the other models ran at xhigh. That K3 came within two points of the leader under default settings makes the result more, not less, interesting — but it’s a caveat worth knowing.
Try It Yourself
Firmulate has packaged 242 real, unedited management decisions from the experiment into a “guess the model” quiz at firmulate.com/quiz.html — a humbling exercise for anyone confident they can tell AI outputs apart. Enterprises can go further and run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

enterprise AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Takeaway for Business Buyers
The Crucible’s uncomfortable conclusion isn’t about which model won. It’s that “same diagnosis, same pitch — no signature” is a failure mode invisible in every vendor demo. A model can nail the analysis, write flawless copy, and still leave the revenue on the table because it didn’t read two layers into your own files.
With scores spanning 73 to 95 on identical work, and a newcomer landing two points off the lead, picking a model without testing it on your own operations is now a bet, not a decision. The tools to stop guessing — the live company, the quiz, the pilot program — are public. The question for marketing and ecommerce leaders is no longer which model is best. It’s whether you’ve run your own Crucible before you sign the contract.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
