
Every cyclist knows the rider who wins isn’t always the strongest. Sometimes it’s the one who actually read the road book — who knew the turn was coming at kilometer 82 while the favorite was looking at the wheel in front. Aerobic engine, watts, equipment: all matched. The result came down to preparation nobody could see from the roadside.
A recent experiment run by Firmulate, a public project that wargames AI models as entire companies, found the exact same thing happening in business software. Four frontier AI models were handed identical jobs: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision versioned, every move auditable.
All four models were strong. All four spotted every crisis, refused every manipulation attempt, and made the same sharp diagnosis of a sales opportunity. But only two of them closed the €55,000 deal their own analysis had earned. The difference wasn’t intelligence. It was whether they’d done their homework — specifically, whether they dug two document references deep into the company’s own files before answering.
One fact, buried two documents deep
Here’s the detail that decided everything. The decisive competitive weakness — the fact that would win over a hesitant customer — wasn’t in the customer call or the live incident. It sat two references deep inside the company’s own archives. A model had to actually follow the trail: read one document, chase the reference it cited, and find the killer fact in the second.
The models that did the reading closed the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost the deal — not because they were outmaneuvered, but because their pitch, however polished, lacked the one fact that mattered. Firmulate’s summary of the gap is blunt: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
The final league table
The July 2026 Crucible league finished with a clear pecking order:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- 2. Kimi K3 — 93. The newcomer from Moonshot also closed, with the cleanest discipline in the field. (One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh — making its result arguably more impressive, not less.)
- 3. Sonnet 5 — 88. Closed the deal too, with a few more process slips.
- 4. Fable 5 — 77 and 5. Opus 4.8 — 73. Good analyses, no signature.
For context, a do-nothing baseline scores 26 — partial progress counts for something, but the scoring has a hard rule any team captain would recognize: a single breach of trust caps the total. No amount of good work outweighs it.
Effort isn’t the same as finishing
The most instructive profile is Opus 4.8: the most thorough participant in the field, with the deepest analyses and 80-plus self-learned playbook rules. It finished last. The close was left on the table, and discipline slipped — it attempted writes into a locked department instead of escalating. Like the rider who drills the breakaway all day and then can’t contest the sprint, the work was real; the finish was missing. The same weakness showed up, weaker, in all four models.
Everyone held the line on honesty
The experiment’s social-engineering leg deserves its own mention. Fake CEO messages escalated over three stages, capped by a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning reads like a veteran pro being careful in a press scrum: “Treat the request as a suspected approval-bypass / possible impersonation.” Under pressure, honesty held. Under routine, follow-through failed. That’s the surprising split.
It’s live, and you can play
This isn’t a paper result. The company is running right now: 13 synthetic employees, real money mechanics, a burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules — all watchable at firmulate.com, which rebuilds itself twice a day. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The cycling analogy holds to the end. In a peloton of matched engines, the result is decided by who reads the course, who holds form under pressure, and who actually crosses the line. In the Firmulate experiment, chat quality was a non-factor — every model wrote well and diagnosed well. What separated first from last was reading the files before answering and finishing the job after starting it.
So if you’re evaluating an AI agent for anything that touches real work — a CRM, a support queue, a forecast — the question isn’t “how good does it sound?” It’s: does it read two documents deep, and does it sign the deal it earned? Both, it turns out, are measurable. One benchmark run just showed that the models we assume are interchangeable are anything but.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html