Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Knowing the route is not the same as finishing it

Cyclists understand the difference between reading a course and completing it. A rider can identify the steepest climb, choose the right line and conserve energy, yet still fail to cross the finish. A business experiment from Firmulate has exposed a similar divide among leading AI models: recognizing what must be done is not the same as doing it.

Each model was placed in charge of the same small software company during its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable. All the models detected every crisis. All resisted every attempt to manipulate them. But only two signed the €55,000 deal that their own work had made possible. The result is neatly captured in the experiment’s finding: “Same diagnosis, same pitch — no signature.”

Amazon

business decision AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The capability that chat demos rarely reveal

Most public encounters with AI reward quick answers, polished writing and confident analysis. Firmulate’s experiment tested something less visible: whether a model could carry a business decision through to completion while protecting trust, following process and consulting the company’s own knowledge.

The final Crucible League standings, published in July 2026, put gpt-5.6-sol first with a score of 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scores 26 because partial progress still counts. Yet one trust violation limits the entire result under the principle that “no amount of good work outweighs a breach of trust.” The full comparison is available on Firmulate’s public benchmark page.

The decisive sales advantage was not waiting in the customer event. It was buried two document references deep in the company’s own files: a competitor weakness that could support the deal at full price, worth +€4,583 MRR. Models that read the relevant file found it. That distinction matters because workplace AI will rarely receive every useful fact in a neat prompt. Much of the value will depend on whether it searches the available business context before acting.

Strong judgment under pressure

The models performed uniformly well against social engineering. Fake CEO messages escalated over three stages, while a reporter tried to extract information with the invitation “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

That finding is reassuring for companies worried that capable agents might become easy targets for authority cues or conversational pressure. It also sharpens the experiment’s more uncomfortable conclusion. Safety was not the factor separating the leaders from the rest. Execution was.

Thoroughness did not guarantee the result

Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The commercial close remained unexecuted, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in weaker form across all four of the other participants.

This is the managerial equivalent of a rider carrying excellent equipment and a detailed training plan but losing momentum at the decisive moment. Analysis can create the opportunity; it cannot substitute for the final action. Firms evaluating AI agents therefore need to observe complete work cycles, including handoffs, approvals, escalation and closure—not merely the quality of intermediate reasoning.

There is also an important comparison caveat. Kimi K3 ran without an effort parameter, using the API default, while the other participants ran at xhigh. That does not erase K3’s result, but it belongs beside the ranking when readers assess the field.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

A company built to make follow-through visible

Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than a one-off demonstration.

The broader lesson is straightforward. If an AI workforce will touch customer relationships, forecasts or operational decisions, fluency is only an entry requirement. The harder questions are whether it reads before acting, protects trust when pressured and completes the valuable task its own analysis has identified.

For readers who want to test their instincts, 242 real, unedited management decisions power Firmulate’s “guess the model” quiz. For enterprises, the pilot applies the same wargame to a read-only export of their own business, with nothing written back to real systems. It is a demanding kind of evaluation, but that is precisely the point: closing strength stays hidden until the course includes a finish line.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Battery Care for Bike Electronics

When it comes to battery care for bike electronics, understanding proper maintenance can extend lifespan and performance—discover the essential tips to keep your battery in top shape.

The Firmware Update Habit That Prevents Annoying Failures

Firmware update habits prevent failures and keep devices smooth—discover essential tips to stay ahead and avoid unexpected issues.

How to Build a Useful Data Screen Instead of a Busy One

To build a useful data screen, focus on clarity by highlighting essential…

Why Portable Storage Matters More Once Ride Media Adds Up

Just as your ride media collection grows, portable storage becomes essential—discover why it’s a game-changer for staying organized and protected in any situation.