
When reliability matters more than a polished answer
Readers who think seriously about home energy, solar and backup power already understand the difference between a specification and a performance. A generator can look capable on paper, but the meaningful question is what happens when the power fails, priorities collide and somebody must act. Artificial intelligence is approaching a similar test. As models move from answering questions to handling business work, fluent writing is no longer enough.
Firmulate has turned that reliability question into a live, watchable experiment. Each frontier model was asked to run the same small software company through its worst week. The customers, crises and temptations remained identical; only the model changed. Every decision was versioned and auditable, creating a record of how different AI managers behaved under the same pressure.
Those decisions now form an unusually revealing interactive article. The Firmulate quiz presents 242 real, unedited management decisions and asks readers to identify the model behind each one. It is partly a guessing game, but the deeper attraction is diagnostic: the models display recognizable working personalities when the stakes extend beyond producing a plausible paragraph.
As an affiliate, we earn on qualifying purchases.
The crises were easy to see. Finishing the work was harder.
The broad competence was impressive. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That distinction is central to the experiment. A model can correctly explain a situation, recommend an action and even prepare the material needed to proceed, yet still fail to complete the business outcome. In a chat window, that hesitation may be invisible. Inside a company, it can determine whether revenue arrives.
The decisive evidence was also easy to miss. A competitor weakness was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that read the document found the fact, won the deal at full price and secured an outcome worth +€4,583 MRR. The lesson is less glamorous than model intelligence demonstrations but more practical: management performance depends on reading the available record before acting.
A league table of management behavior
The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress still counted. But the competition imposed a firm ethical boundary: “no amount of good work outweighs a breach of trust,” and a single breach capped the total.
Kimi K3’s result carries an important fairness note. It ran with the API default because it had no effort parameter, while the others ran at xhigh. Even with that difference, K3 finished close behind the winner and signed the deal.
The quiz makes these contrasts tangible without reducing them to a ranking alone. One decision may unfold as a dissertation, another may be strikingly terse, and another may refuse to engage with distracting communication. Readers see the unedited choices first and discover the model afterward, making stylistic habits and operational discipline easier to notice.
Pressure tests for honesty
The company’s worst week also included fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous resistance matters because the experiment did not test management in a clean sandbox of ordinary tasks. It introduced the kinds of urgency, authority claims and informal requests that can cause people or automated workers to bypass procedure. The models recognized the traps, even though they differed substantially in their ability to carry legitimate work through to completion.
Thoroughness did not guarantee victory
Opus 4.8 offers the clearest caution against equating volume with effectiveness. It was the most thorough participant, learning +80 rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the remaining four models.
This is why Firmulate’s experiment feels closer to a stress test than a writing benchmark. It evaluates whether an AI manager reads deeply, respects boundaries and finishes what it starts. Those qualities can coexist unevenly: a model may be highly analytical and trustworthy while still failing at execution.

As an affiliate, we earn on qualifying purchases.
A company readers can watch under load
The live company includes 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The deteriorating economics give each decision context: delay and incomplete execution are not abstract defects when the company is already losing money.
For readers accustomed to evaluating backup systems, the useful analogy is resilience under load. The best system is not merely the one that recognizes an outage or describes the correct response. It must find the relevant information, reject unsafe shortcuts and complete the action that restores a working state.
Firmulate’s larger finding is that frontier AI models already have measurable management personalities. They can face identical evidence and reach the same diagnosis, yet diverge at the moment when a decision must become an outcome. The quiz makes that difference visible—and surprisingly difficult to guess.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.