
The Demo That Won’t Tell You What You Need to Know
If you’re shopping for an AI agent to touch your CRM, your support queue, or your forecast, you’ve probably seen the demos: fluent answers, clever code, a leaderboard ranking. Chat arenas and coding benchmarks measure one thing — how well a model answers. They say almost nothing about how it behaves when the week goes wrong: a churn wave, a price increase, a downround, a PR crisis, and a tempting shortcut sitting right there on the table.
That gap is exactly what Firmulate, a live public experiment, was built to expose. Its pitch is blunt: it measures management quality, not chat quality. And the first set of results suggests the two are barely correlated.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Worst Week, Four Different AI Bosses
The setup is elegant. Four frontier AI models were each handed the same small software company and the same brutal week — same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so the whole thing can be replayed and checked.
The final league table from the July 2026 run: gpt-5.6-sol finished first with 95, Kimi K3 second with 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. A do-nothing baseline scored 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: “no amount of good work outweighs a breach of trust.”
The Headline Finding
Here’s what should make any executive sit up. All four models spotted every crisis. All four refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The models did the hard analytical work and then simply didn’t finish the job.
The Buried Fact
The deal turned on something subtler than the customer conversation itself. The decisive competitor weakness was sitting two document references deep in the company’s own files — not in the customer event. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. In business terms: the winners weren’t smarter talkers. They were the ones who did their homework before the meeting.
The Social Engineering Test
The week also included fake CEO messages escalating over three stages, plus a reporter offering an easy out: “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning was admirably paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Cautionary Tale
Then there’s Opus 4.8 — the most thorough participant in the field, generating the deepest analyses and 80+ learned rules, yet finishing dead last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more faintly, in all four models. Effort and diligence, it turns out, don’t automatically convert into outcomes.
One fairness note worth recording: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.
As an affiliate, we earn on qualifying purchases.
Not a Slide Deck — A Company You Can Watch Lose Money
What makes Firmulate more than a one-off report is that the company is real software that runs every business day. It has 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and a self-learned playbook that has grown past 680 rules. Every workday is versioned, and you can watch it live at firmulate.com.
Want to test your own instincts? 242 real, unedited management decisions from the runs power a “guess the model” quiz. And enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever written back to real systems. Full benchmark results are published at firmulate.com/benchmarks.html.

As an affiliate, we earn on qualifying purchases.
The Question That Actually Matters
If AI agents are going to run parts of your business, the question is not “does it write well?” It’s: does it finish what it starts, does it read your files before it acts, does it stay honest under pressure — and what does a unit of useful work actually cost?
The Firmulate results suggest that chat fluency is table stakes, and that the differentiators — closing discipline, document diligence, resistance to impersonation — only show up when you put models in charge of consequences that unfold over days. That’s a new curriculum for evaluation: churn waves, price increases, downrounds, PR crises. Leaderboards measure the answer. This measures the manager.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.