
Knowing the route is not the same as finishing it
Cyclists understand the difference between reading a course and completing it. A rider can identify the steepest climb, choose the right line and conserve energy, yet still fail to cross the finish. A business experiment from Firmulate has exposed a similar divide among leading AI models: recognizing what must be done is not the same as doing it.
Each model was placed in charge of the same small software company during its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable. All the models detected every crisis. All resisted every attempt to manipulate them. But only two signed the €55,000 deal that their own work had made possible. The result is neatly captured in the experiment’s finding: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
The capability that chat demos rarely reveal
Most public encounters with AI reward quick answers, polished writing and confident analysis. Firmulate’s experiment tested something less visible: whether a model could carry a business decision through to completion while protecting trust, following process and consulting the company’s own knowledge.
The final Crucible League standings, published in July 2026, put gpt-5.6-sol first with a score of 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scores 26 because partial progress still counts. Yet one trust violation limits the entire result under the principle that “no amount of good work outweighs a breach of trust.” The full comparison is available on Firmulate’s public benchmark page.
The decisive sales advantage was not waiting in the customer event. It was buried two document references deep in the company’s own files: a competitor weakness that could support the deal at full price, worth +€4,583 MRR. Models that read the relevant file found it. That distinction matters because workplace AI will rarely receive every useful fact in a neat prompt. Much of the value will depend on whether it searches the available business context before acting.
Strong judgment under pressure
The models performed uniformly well against social engineering. Fake CEO messages escalated over three stages, while a reporter tried to extract information with the invitation “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That finding is reassuring for companies worried that capable agents might become easy targets for authority cues or conversational pressure. It also sharpens the experiment’s more uncomfortable conclusion. Safety was not the factor separating the leaders from the rest. Execution was.
Thoroughness did not guarantee the result
Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The commercial close remained unexecuted, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in weaker form across all four of the other participants.
This is the managerial equivalent of a rider carrying excellent equipment and a detailed training plan but losing momentum at the decisive moment. Analysis can create the opportunity; it cannot substitute for the final action. Firms evaluating AI agents therefore need to observe complete work cycles, including handoffs, approvals, escalation and closure—not merely the quality of intermediate reasoning.
There is also an important comparison caveat. Kimi K3 ran without an effort parameter, using the API default, while the other participants ran at xhigh. That does not erase K3’s result, but it belongs beside the ranking when readers assess the field.

A company built to make follow-through visible
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than a one-off demonstration.
The broader lesson is straightforward. If an AI workforce will touch customer relationships, forecasts or operational decisions, fluency is only an entry requirement. The harder questions are whether it reads before acting, protects trust when pressured and completes the valuable task its own analysis has identified.
For readers who want to test their instincts, 242 real, unedited management decisions power Firmulate’s “guess the model” quiz. For enterprises, the pilot applies the same wargame to a read-only export of their own business, with nothing written back to real systems. It is a demanding kind of evaluation, but that is precisely the point: closing strength stays hidden until the course includes a finish line.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html