AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The Demo That Won’t Tell You What You Need to Know

If you’re shopping for an AI agent to touch your CRM, your support queue, or your forecast, you’ve probably seen the demos: fluent answers, clever code, a leaderboard ranking. Chat arenas and coding benchmarks measure one thing — how well a model answers. They say almost nothing about how it behaves when the week goes wrong: a churn wave, a price increase, a downround, a PR crisis, and a tempting shortcut sitting right there on the table.

That gap is exactly what Firmulate, a live public experiment, was built to expose. Its pitch is blunt: it measures management quality, not chat quality. And the first set of results suggests the two are barely correlated.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Worst Week, Four Different AI Bosses

The setup is elegant. Four frontier AI models were each handed the same small software company and the same brutal week — same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so the whole thing can be replayed and checked.

The final league table from the July 2026 run: gpt-5.6-sol finished first with 95, Kimi K3 second with 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. A do-nothing baseline scored 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: “no amount of good work outweighs a breach of trust.”

The Headline Finding

Here’s what should make any executive sit up. All four models spotted every crisis. All four refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The models did the hard analytical work and then simply didn’t finish the job.

The Buried Fact

The deal turned on something subtler than the customer conversation itself. The decisive competitor weakness was sitting two document references deep in the company’s own files — not in the customer event. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. In business terms: the winners weren’t smarter talkers. They were the ones who did their homework before the meeting.

The Social Engineering Test

The week also included fake CEO messages escalating over three stages, plus a reporter offering an easy out: “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning was admirably paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Cautionary Tale

Then there’s Opus 4.8 — the most thorough participant in the field, generating the deepest analyses and 80+ learned rules, yet finishing dead last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more faintly, in all four models. Effort and diligence, it turns out, don’t automatically convert into outcomes.

One fairness note worth recording: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

Amazon

enterprise AI chatbot for CRM

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Not a Slide Deck — A Company You Can Watch Lose Money

What makes Firmulate more than a one-off report is that the company is real software that runs every business day. It has 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, and a self-learned playbook that has grown past 680 rules. Every workday is versioned, and you can watch it live at firmulate.com.

Want to test your own instincts? 242 real, unedited management decisions from the runs power a “guess the model” quiz. And enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever written back to real systems. Full benchmark results are published at firmulate.com/benchmarks.html.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Question That Actually Matters

If AI agents are going to run parts of your business, the question is not “does it write well?” It’s: does it finish what it starts, does it read your files before it acts, does it stay honest under pressure — and what does a unit of useful work actually cost?

The Firmulate results suggest that chat fluency is table stakes, and that the differentiators — closing discipline, document diligence, resistance to impersonation — only show up when you put models in charge of consequences that unfold over days. That’s a new curriculum for evaluation: churn waves, price increases, downrounds, PR crises. Leaderboards measure the answer. This measures the manager.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Volatus Aerospace Releases Q2 2026 Financial Results

Volatus Aerospace announced its Q2 2026 financial results, showing revenue growth and operational updates. Details remain preliminary as the company releases official figures.

OlmoEarth Embeddings: Precise Data Exports For AI Efficiency

OlmoEarth Studio now supports on-demand generation and export of satellite embedding vectors, enhancing Earth-observation analysis and AI efficiency.

Grok 4.6: The Frontier Is Now A Price War

Grok 4.6, released by SpaceXAI, advances in intelligence but maintains flat pricing, igniting a significant price competition at the AI frontier.

Cosmos Health Projects Over $6M In Annual Revenue And $4.3M In Gross Profit From Oliv18™ Across The U.S. Cardiovascular Health And Antioxidant Categories

Cosmos Health projects over $6 million in annual revenue and $4.3 million in gross profit from Oliv18™ sales across the U.S. in cardiovascular and antioxidant markets.