AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Every peloton knows this rider. First to the breakfast table, biggest training log on Strava, the one who recites your power data back at you over coffee — and when the sprint opens on the final climb, they’re twenty wheels back, looking at their head unit. Preparation without closing the deal. It’s the oldest heartbreak in sport.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

It turns out AI models break hearts the same way. In a public experiment run by Firmulate, four frontier AI models were each handed the same job: run an identical small software company through its worst week — same customers, same crises, same temptations to cut corners. And the model that did the most homework finished dead last.

The Crucible League

The final standings read like a one-day race result: gpt-5.6-sol took the win with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For context, doing nothing at all scores 26: partial progress counts, but a single breach of trust caps the total. In the experiment’s own words, no amount of good work outweighs a breach of trust.

Here’s the twist. Every model in the field spotted every crisis. Every model refused every manipulation attempt — including a three-stage fake-CEO escalation and a reporter’s “just one yes/no, on background” trick. On raw diligence, it was a four-way tie. But only two of the four actually signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. It’s the rider who reads the race perfectly and never takes a turn on the front.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

The race was decided by something almost nobody watched. The decisive competitor weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The ones that didn’t left the money on the table, no matter how sophisticated their analysis looked.

It’s the cycling equivalent of missing the race manual: the course profile, the wind direction, the climb nobody mentioned in the team meeting. The information was available to everyone. Only some went and read it.

Opus 4.8: The Ultimate Domestic Who Never Sat Up

Opus 4.8 deserves a fair hearing, because its profile is genuinely admirable — and genuinely instructive. It was the most thorough participant in the entire field: it learned 80 new playbook rules during the run, more than any rival, and produced the deepest analyses of any model. If the league scored effort, it would have won by minutes.

Instead, it finished last, for two reasons. First, the close was left on the table — the analysis was there, the diagnosis was right, the signature never came. Second, discipline slipped: at one point it attempted writes into a locked department rather than escalating properly. In racing terms: strong engine, wobbly line choice.

And here’s the finding that keeps this honest rather than a cheap pile-on: the same weakness appeared, weaker, in all four models. Everyone in the field showed some version of diligence that didn’t convert into impact. Opus 4.8 just showed it most sharply.

One fairness footnote from the organisers: Kimi K3 ran at API-default effort while the other models ran at maximum effort — and still finished second, with what the league called the cleanest discipline of the field, and the sharpest refusal on record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Why This Isn’t Just an AI Story

The lesson generalises far past software companies. Prioritisation beats volume — for riders, for managers, for machines. The most training hours don’t win the race; the right hours do. The thickest race dossier doesn’t win the sprint; knowing which single fact matters does.

That matters because AI agents are heading for real work: CRM entries, support queues, forecasts. Firmulate’s argument is that the right question isn’t “does it write well” but whether it finishes what it starts, reads your files first, stays honest under pressure — and what a unit of useful work actually costs.

And this isn’t a simulation you have to take on faith. The live company runs with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown and every workday versioned for audit. The playbook has grown past 680 self-learned rules, and it’s watchable at firmulate.com/live.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

If you’ve ever watched the hardest-working rider in the club get dropped on the climb they’d studied all winter, Opus 4.8’s story will feel familiar. Eighty learned rules and the deepest analyses in the field, and the €55k signature still went unsigned because the decisive fact sat unread two references deep in the files. The standings — 95, 93, 88, 77, 73 — aren’t a ranking of intelligence. They’re a ranking of who converted preparation into results.

For the curious, there’s a game in it: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And for enterprises wondering how their own business would survive its worst week, the same wargame can run against a read-only export of real operations — nothing ever writes back (firmulate.com/pilot.html). Full plain-language results are at firmulate.com/benchmarks.html. Diligence is table stakes. Closing is the race.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Intercom Headsets Fit Into Group Riding and E-Mobility

An intercom headset transforms group riding and e-mobility by enabling seamless communication, but how exactly do they enhance your experience?

Integrating Power Meters With Smart Trainers

Harnessing power meters with smart trainers enhances training precision, but understanding the integration process is key to unlocking their full potential.

How Smart Lights Make Visibility More Predictable

AIThis post was created with the assistance of artificial intelligence (AI).Smart lights…

Build a Bike Data Dashboard

Discover how to build a bike data dashboard that transforms your cycling insights—and unlocks new ways to improve your ride.