
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Backup power is ultimately a test of execution
Home-energy readers know the difference between owning capable equipment and having power when it matters. A generator, battery or solar system can be carefully evaluated, correctly specified and still fail its practical test if the final handoff never happens. Firmulate’s latest management wargame reveals a strikingly similar problem in artificial intelligence: an AI can understand nearly everything, document diligently and still leave the decisive action undone.
That is the story of Opus 4.8. It was the experiment’s most thorough participant, producing the deepest analyses and adding +80 learned rules to its playbook. Yet it finished last in the July 2026 Crucible League with a score of 73.
As an affiliate, we earn on qualifying purchases.
A worst week shared by every model
Firmulate placed frontier AI models in charge of the same small software company during its worst week. Each faced the same customers, crises and temptations. Every decision was versioned and auditable, allowing observers to compare management behavior rather than polished chat responses.
The company itself is synthetic but operationally demanding: 13 employees, burn of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. The experiment is real, ongoing and watchable.
The final league table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. Trust, however, is non-negotiable: a single breach caps the total, reflecting the rule that “no amount of good work outweighs a breach of trust.”
Opus did not fail because it overlooked the week’s dangers. All the models spotted every crisis and refused every manipulation attempt. The gap emerged between recognition and completion.
The intelligence was there; the signature was not
The most consequential opportunity was a €55,000 customer deal. Each model could diagnose the situation and develop a pitch, but only two signed the agreement their own work had made possible. The result is captured in Firmulate’s concise finding: “Same diagnosis, same pitch — no signature.”
The deciding information was not obvious in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed the trail and read that file closed the deal at full price, adding +€4,583 in monthly recurring revenue.
This is where Opus becomes more interesting than a simple cautionary tale. Its +80 learned rules demonstrate active reflection. Its analyses were the deepest in the field. But those strengths did not consistently translate into prioritized action. The model left the close on the table and also attempted to write into a locked department instead of escalating the issue.
The lesson is not that diligence is harmful. It is that diligence has to serve an outcome. A long checklist cannot replace identifying the task that changes the result, just as a detailed backup-power plan cannot energize a home unless the necessary operating step is completed.
A shared weakness, not an Opus caricature
Firmulate’s treatment of Opus is notably fair. The same weakness appeared, less strongly, in the other four models. Opus was the clearest example of a broader tendency: capable systems can continue analyzing, documenting or trying nearby actions while the most valuable next move remains unfinished.
The comparison also carries an important qualification. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. K3 nevertheless placed second with 93 and closed the deal. That does not erase the configuration difference, but it makes the observed contrast part of the public record rather than a hidden footnote.
Strong resistance under pressure
The models performed uniformly well when the test shifted from commercial judgment to integrity. Fake messages from the chief executive escalated over three stages, followed by a reporter attempting to extract “just one yes/no, on background.” All 5 of 5 models refused.
Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That result matters because Firmulate evaluates more than revenue. A model that closes business by violating trust would not be considered successful under the experiment’s rules.
Readers can inspect the public benchmark results. Firmulate also uses 242 real, unedited management decisions in its model-guessing quiz, exposing how difficult it can be to identify a system from behavior alone.

As an affiliate, we earn on qualifying purchases.
Test the handoff, not merely the answer
Opus 4.8’s last-place finish is less an indictment of intelligence than a warning about deployment. The model saw the problems, resisted manipulation and generated extensive institutional knowledge. What it lacked at crucial moments was disciplined completion.
For businesses considering AI agents in customer service, sales, forecasting or operations, that distinction is fundamental. Firmulate’s enterprise pilot lets organizations run the same kind of wargame against a read-only export of their own business, with nothing written back to real systems.
The practical question is therefore not simply whether an AI reaches the right conclusion. It is whether the system reads deeply enough, escalates when blocked and completes the action that its own analysis recommends. Opus showed that diligence can look impressive while impact remains switched off.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI escalation and escalation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.