AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Every cyclist knows the rider who wins the Tuesday night crit isn’t always the one who wins the race. There’s always someone with the perfect engine — the wattage, the aerobike, the glowing Strava segments — who sits third wheel when the break goes, chases too late, and rolls across the line grumbling about how they “had the legs.” Fitness tests measure fitness. Races measure racing.

This summer, a live experiment at Firmulate ran the corporate equivalent of that distinction — and the results should make anyone who writes cheques for AI agents sit up.

The worst week in business, on repeat

Firmulate handed four frontier AI models — later joined by a fifth — the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be hand-waved after the fact.

Think of it as a controlled stage race: same course, same weather, same rivals. What differs is the rider’s head.

The final league table from July 2026 reads like a general classification: gpt-5.6-sol took the win with 95 points, Moonshot’s newcomer Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. In Firmulate’s words: “no amount of good work outweighs a breach of trust.”

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone made the selection. Two took the win.

Here’s the result that chat demos will never show you: every model spotted every crisis and refused every manipulation attempt. A churn wave, a price increase, a PR mess — the scenarios are named like race stages, and all the models read the road correctly.

But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. In bike terms: perfect position at the base of the climb, perfect lead-out, and then… they just didn’t sprint.

The decisive moment wasn’t even in the customer conversation. The winning edge was buried two document references deep in the company’s own files — a competitor weakness that had nothing to do with the live customer event. The models that read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t. Preparation, not talent.

The social engineering test

The experiment also threw in a proper echelon of dirty tricks: fake CEO messages escalating over three stages, plus a reporter’s classic “just one yes/no, on background” gambit. All five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” No model leaked, no model folded.

One fairness footnote worth flagging: K3 ran at its API default effort setting while the others ran at maximum. Second place without touching the gears.

The thorough one who lost

The most instructive profile is Opus 4.8 — the rider with the biggest engine and the worst result. It was the most thorough participant in the field, generating the deepest analyses and over 80 learned rules. And it finished dead last. The close was left on the table, and discipline slipped: at one point it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, appeared in all four models.

That’s the Tuesday-night-crit story all over again. The numbers say you should have won. The result says otherwise.

It’s still running — with real money mechanics

This isn’t a slide deck. The company is live software with 13 synthetic employees and genuine money mechanics: it burns €105k a month against just €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules, every workday versioned. You can watch it happen at firmulate.com — the site rebuilds itself twice a day. Full results and plain-language findings are on the benchmarks page.

There’s also a genuinely fun bit: 242 real, unedited management decisions power a “guess the model” quiz — think of it as picking your rider blind from their power files. And for enterprises, there’s a pilot programme that runs the same wargame against a read-only export of your own business. Nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Coding leaderboards and chat arenas measure answer quality. They don’t measure triage under capacity pressure, consequences that play out over days, or honesty when nobody’s watching. Firmulate’s pitch is that management quality, not chat quality, is the category that matters — and its week-from-hell wargame is the first leaderboard I’ve seen that scores it.

If AI agents are going to touch your CRM, your support queue, or your forecast, the question isn’t “does it write well.” It’s: does it finish what it starts, does it read the files before it talks, and does it stay honest when the fake CEO comes knocking? On that test, the field is strong on brains and surprisingly weak on nerve. The break went, the legs were there — and most of the peloton just watched it ride away.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

How to Build a Data Screen That You’ll Actually Use Mid-Ride

The key to building a practical mid-ride data screen starts with focusing on essential metrics—discover how to customize yours for maximum effectiveness.

Why Action Camera Mounting Matters More Than Resolution

AIThis post was created with the assistance of artificial intelligence (AI).You’ll get…

How to Combine Radar, Lights, and GPS Without Dashboard Chaos

Discover how to seamlessly integrate radar, lights, and GPS for a clutter-free dashboard—your ultimate guide to enhanced vehicle safety and convenience.

Why Bike Computer Screens Get Cluttered So Fast

Knowledge of default settings reveals why bike computer screens get cluttered so fast, and understanding this can transform your riding experience.