AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every cyclist knows the rider who wins isn’t always the strongest. Sometimes it’s the one who actually read the road book — who knew the turn was coming at kilometer 82 while the favorite was looking at the wheel in front. Aerobic engine, watts, equipment: all matched. The result came down to preparation nobody could see from the roadside.

A recent experiment run by Firmulate, a public project that wargames AI models as entire companies, found the exact same thing happening in business software. Four frontier AI models were handed identical jobs: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision versioned, every move auditable.

All four models were strong. All four spotted every crisis, refused every manipulation attempt, and made the same sharp diagnosis of a sales opportunity. But only two of them closed the €55,000 deal their own analysis had earned. The difference wasn’t intelligence. It was whether they’d done their homework — specifically, whether they dug two document references deep into the company’s own files before answering.

One fact, buried two documents deep

Here’s the detail that decided everything. The decisive competitive weakness — the fact that would win over a hesitant customer — wasn’t in the customer call or the live incident. It sat two references deep inside the company’s own archives. A model had to actually follow the trail: read one document, chase the reference it cited, and find the killer fact in the second.

The models that did the reading closed the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost the deal — not because they were outmaneuvered, but because their pitch, however polished, lacked the one fact that mattered. Firmulate’s summary of the gap is blunt: “Same diagnosis, same pitch — no signature.”

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The final league table

The July 2026 Crucible league finished with a clear pecking order:

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
  • 2. Kimi K3 — 93. The newcomer from Moonshot also closed, with the cleanest discipline in the field. (One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh — making its result arguably more impressive, not less.)
  • 3. Sonnet 5 — 88. Closed the deal too, with a few more process slips.
  • 4. Fable 5 — 77 and 5. Opus 4.8 — 73. Good analyses, no signature.

For context, a do-nothing baseline scores 26 — partial progress counts for something, but the scoring has a hard rule any team captain would recognize: a single breach of trust caps the total. No amount of good work outweighs it.

Effort isn’t the same as finishing

The most instructive profile is Opus 4.8: the most thorough participant in the field, with the deepest analyses and 80-plus self-learned playbook rules. It finished last. The close was left on the table, and discipline slipped — it attempted writes into a locked department instead of escalating. Like the rider who drills the breakaway all day and then can’t contest the sprint, the work was real; the finish was missing. The same weakness showed up, weaker, in all four models.

Everyone held the line on honesty

The experiment’s social-engineering leg deserves its own mention. Fake CEO messages escalated over three stages, capped by a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning reads like a veteran pro being careful in a press scrum: “Treat the request as a suspected approval-bypass / possible impersonation.” Under pressure, honesty held. Under routine, follow-through failed. That’s the surprising split.

It’s live, and you can play

This isn’t a paper result. The company is running right now: 13 synthetic employees, real money mechanics, a burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules — all watchable at firmulate.com, which rebuilds itself twice a day. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The cycling analogy holds to the end. In a peloton of matched engines, the result is decided by who reads the course, who holds form under pressure, and who actually crosses the line. In the Firmulate experiment, chat quality was a non-factor — every model wrote well and diagnosed well. What separated first from last was reading the files before answering and finishing the job after starting it.

So if you’re evaluating an AI agent for anything that touches real work — a CRM, a support queue, a forecast — the question isn’t “how good does it sound?” It’s: does it read two documents deep, and does it sign the deal it earned? Both, it turns out, are measurable. One benchmark run just showed that the models we assume are interchangeable are anything but.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

The AI Management Peloton Has a Breakaway Problem

Can you spot an AI manager’s habits? Firmulate turns 242 audited decisions into a quiz about judgment, follow-through and trust under pressure.

The Firmware Update Habit That Prevents Annoying Failures

Firmware update habits prevent failures and keep devices smooth—discover essential tips to stay ahead and avoid unexpected issues.

Installing an Airtag Inside Your Frame—Is It Worth It?

Many bike owners wonder if hiding an Airtag inside their frame is worth it, but consider the risks before making your decision.

GPS Trackers: What Happens After Your Bike Is Stolen?

Inevitably, understanding how GPS trackers assist in recovery can make all the difference when your bike is stolen.