
Two AI agents gave the same pitch. Only one closed.
Imagine two of your sales reps walk into the same meeting with the same diagnosis, the same proposal, and the same €55,000 opportunity. One walks out with a signed contract at full price. The other walks out with nothing — and can’t explain why. The difference wasn’t talent, charm, or effort. It was homework. One had read a file buried two references deep in your company’s own documents; the other hadn’t.
That scenario isn’t hypothetical. It’s the headline finding from a live, public experiment by Firmulate, which runs frontier AI models as complete companies through identical worst-case weeks and scores them on management quality — not chat quality. The result reframes what buyers should actually be shopping for in an AI agent.
As an affiliate, we earn on qualifying purchases.
The experiment: same company, same crises, same temptations
Firmulate handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing rested on vibes.
The final July 2026 league table tells a surprising story:
- 1. gpt-5.6-sol — 95 points. Found the buried fact, closed the deal — the complete performance.
- 2. Kimi K3 — 93 points. The Moonshot newcomer closed the deal too, with the cleanest discipline of the field.
- 3. Sonnet 5 — 88 points. Closed the deal, with a few process slips.
- 4. Fable 5 — 77 points and 5. Opus 4.8 — 73 points.
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.” One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh, which makes its second-place finish even more striking.
AI-powered enterprise document analysis
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buried fact that decided €55,000
Here’s where it gets interesting for anyone who buys or deploys AI agents. The decisive competitor weakness in that €55,000 deal wasn’t in the customer meeting, the email thread, or the CRM notes from the call. It sat two document references deep in the company’s own files.
The models that followed the references and actually read the file won the deal — at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, lost it automatically. Same diagnosis, same pitch — no signature.
All five models spotted every crisis and refused every manipulation attempt, including a three-stage fake-CEO escalation and a reporter’s “just one yes/no, on background” trick. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” Honesty under pressure turned out to be table stakes. Finishing the job — doing the unglamorous reading — was the differentiator.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness isn’t the same as follow-through
The most counterintuitive profile belongs to Opus 4.8: the most thorough participant in the field, with the deepest analyses and the most learned rules (+80), yet last place. The close was left on the table, and discipline slipped — including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.
That’s a caution for anyone evaluating agents on demos of analytical brilliance. The gap between a great analysis and a signed deal is invisible in a chat showcase — and it’s exactly where revenue lives or dies.
As an affiliate, we earn on qualifying purchases.
Why this is watchable, not hypothetical
Firmulate isn’t a slide deck. It’s a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it running at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The buying question has changed
If AI agents will touch your CRM, your support queue, or your forecast, the question is no longer “does it write well?” It’s: does it finish what it starts, does it read your files before answering, does it stay honest under pressure — and what does a unit of useful work actually cost?
The Firmulate experiment makes “reads your files first” a measurable, purchase-deciding property — worth exactly €55,000 in one week at one small software company. Before you sign an agent vendor, ask them to prove it the same way: same scenario, buried fact, auditable decisions. Full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html