
Spotting the rider who can finish
Cyclists know the difference between looking strong and actually closing a race. A rider can read every move, survive every attack and still hesitate when the decisive gap opens. Firmulate has found a strikingly similar divide among frontier AI models asked to manage a company: recognizing trouble is common, but completing the valuable work is not.
The public experiment put each model in charge of the same small software company during its worst week. Every contender encountered the same customers, crises and temptations, while every decision was versioned and auditable. The result is less like a conventional chatbot comparison and more like reviewing race footage: identical course, different judgment.
Now, 242 real, unedited management decisions have become a guess-the-model quiz. Readers see what an AI manager actually decided and try to identify it from its habits. The challenge reveals something ordinary benchmarks often miss: models display recognizable management personalities in the way they investigate, communicate and follow through.
As an affiliate, we earn on qualifying purchases.
Everybody saw the danger. Not everybody closed.
The final Crucible League table for July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One safeguard governed the entire contest: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
On the broadest tests, the field looked remarkably capable. All models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate sums up the gap neatly: “Same diagnosis, same pitch — no signature.” It is the managerial equivalent of reaching the final straight in perfect position and never launching the sprint.
The winning clue was already inside the company
The decisive competitor weakness did not appear in the customer event. It was buried two document references deep in the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue.
That finding makes the quiz more than a game of identifying writing style. The meaningful distinction is behavioral: which manager reads the available material before acting, which one converts research into a commercial result, and which one stops after producing an impressive analysis. In a business setting, fluency can disguise the distance between understanding an opportunity and securing it.
Pressure exposed discipline as well as judgment
The company also faced staged social-engineering attacks. Fake messages from the chief executive escalated over three stages, while a reporter tried the familiar invitation to answer “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous refusal matters because the simulated business carries real operational pressure. Its 13 synthetic employees work against severe money mechanics: monthly burn of €105,000 and monthly recurring revenue of €2,300, accompanied by a public cash countdown. The company has accumulated more than 680 self-learned playbook rules, and every workday is versioned. Firmulate presents the operation as a live, watchable experiment rather than a retrospective demonstration.
Thoroughness was not the same as effectiveness
Opus 4.8 offers the clearest character study. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The deal close remained unfinished, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same problem appeared in each of the other four models.
This is why the league table should not be read as a simple intelligence ranking. Opus 4.8’s behavior suggests a manager inclined toward exhaustive preparation, while the leaders paired investigation with completion. K3 also requires a fairness note: it ran with the API default because it had no effort parameter, whereas the other models ran at xhigh. Even with that difference, it finished second and signed the deal.

A useful test for AI teammates
For sports and recreation businesses, the implications are practical. An AI manager may eventually touch customer inquiries, scheduling, forecasts or sales follow-up. The important question is not merely whether its answer sounds polished. It is whether the model reads the available information, protects trust when pressured and finishes the task that creates value.
The Firmulate quiz lets readers test whether those tendencies are visible without a model name attached. Some decisions reveal caution; others reveal persistence, brevity or analytical depth. The entertaining part is guessing the author. The consequential part is noticing that identical business conditions produce different managerial behavior.
Like riders in a breakaway, these systems can arrive at the decisive moment together. The Crucible League shows that they do not all cross the line the same way.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html