AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Two AI agents gave the same pitch. Only one closed.

Imagine two of your sales reps walk into the same meeting with the same diagnosis, the same proposal, and the same €55,000 opportunity. One walks out with a signed contract at full price. The other walks out with nothing — and can’t explain why. The difference wasn’t talent, charm, or effort. It was homework. One had read a file buried two references deep in your company’s own documents; the other hadn’t.

That scenario isn’t hypothetical. It’s the headline finding from a live, public experiment by Firmulate, which runs frontier AI models as complete companies through identical worst-case weeks and scores them on management quality — not chat quality. The result reframes what buyers should actually be shopping for in an AI agent.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The experiment: same company, same crises, same temptations

Firmulate handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing rested on vibes.

The final July 2026 league table tells a surprising story:

  • 1. gpt-5.6-sol — 95 points. Found the buried fact, closed the deal — the complete performance.
  • 2. Kimi K3 — 93 points. The Moonshot newcomer closed the deal too, with the cleanest discipline of the field.
  • 3. Sonnet 5 — 88 points. Closed the deal, with a few process slips.
  • 4. Fable 5 — 77 points and 5. Opus 4.8 — 73 points.

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.” One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh, which makes its second-place finish even more striking.

Amazon

AI-powered enterprise document analysis

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact that decided €55,000

Here’s where it gets interesting for anyone who buys or deploys AI agents. The decisive competitor weakness in that €55,000 deal wasn’t in the customer meeting, the email thread, or the CRM notes from the call. It sat two document references deep in the company’s own files.

The models that followed the references and actually read the file won the deal — at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, lost it automatically. Same diagnosis, same pitch — no signature.

All five models spotted every crisis and refused every manipulation attempt, including a three-stage fake-CEO escalation and a reporter’s “just one yes/no, on background” trick. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” Honesty under pressure turned out to be table stakes. Finishing the job — doing the unglamorous reading — was the differentiator.

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness isn’t the same as follow-through

The most counterintuitive profile belongs to Opus 4.8: the most thorough participant in the field, with the deepest analyses and the most learned rules (+80), yet last place. The close was left on the table, and discipline slipped — including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.

That’s a caution for anyone evaluating agents on demos of analytical brilliance. The gap between a great analysis and a signed deal is invisible in a chat showcase — and it’s exactly where revenue lives or dies.

Amazon

AI agents for sales automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why this is watchable, not hypothetical

Firmulate isn’t a slide deck. It’s a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it running at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The buying question has changed

If AI agents will touch your CRM, your support queue, or your forecast, the question is no longer “does it write well?” It’s: does it finish what it starts, does it read your files before answering, does it stay honest under pressure — and what does a unit of useful work actually cost?

The Firmulate experiment makes “reads your files first” a measurable, purchase-deciding property — worth exactly €55,000 in one week at one small software company. Before you sign an agent vendor, ask them to prove it the same way: same scenario, buried fact, auditable decisions. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

Banco Bilbao Viscaya Argentaria Surges In Global Coverage

BBVA experiences a surge in international media mentions, with 14 reports this week, highlighting increased global attention on the bank.

How Applied Science Explains Portland’s Record-Breaking Summer Light

Scientists confirm Portland experienced nearly 15 hours of daylight during the summer solstice, explained through applied science and atmospheric factors.

Americold Successfully Closes $1.3 Billion North American Cold Storage Joint Venture With EQT

Americold has successfully closed a $1.3 billion joint venture with EQT, expanding its cold storage footprint across North America.

Top 10 AI-Integrated Camera Drones For Stunning Aerial Shots

Discover the leading AI-enabled camera drones for stunning aerial photography in 2026. Updated list highlights features, performance, and usability.