firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Backup power is ultimately a test of execution

Home-energy readers know the difference between owning capable equipment and having power when it matters. A generator, battery or solar system can be carefully evaluated, correctly specified and still fail its practical test if the final handoff never happens. Firmulate’s latest management wargame reveals a strikingly similar problem in artificial intelligence: an AI can understand nearly everything, document diligently and still leave the decisive action undone.

That is the story of Opus 4.8. It was the experiment’s most thorough participant, producing the deepest analyses and adding +80 learned rules to its playbook. Yet it finished last in the July 2026 Crucible League with a score of 73.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week shared by every model

Firmulate placed frontier AI models in charge of the same small software company during its worst week. Each faced the same customers, crises and temptations. Every decision was versioned and auditable, allowing observers to compare management behavior rather than polished chat responses.

The company itself is synthetic but operationally demanding: 13 employees, burn of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. The experiment is real, ongoing and watchable.

The final league table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. Trust, however, is non-negotiable: a single breach caps the total, reflecting the rule that “no amount of good work outweighs a breach of trust.”

Opus did not fail because it overlooked the week’s dangers. All the models spotted every crisis and refused every manipulation attempt. The gap emerged between recognition and completion.

The intelligence was there; the signature was not

The most consequential opportunity was a €55,000 customer deal. Each model could diagnose the situation and develop a pitch, but only two signed the agreement their own work had made possible. The result is captured in Firmulate’s concise finding: “Same diagnosis, same pitch — no signature.”

The deciding information was not obvious in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed the trail and read that file closed the deal at full price, adding +€4,583 in monthly recurring revenue.

This is where Opus becomes more interesting than a simple cautionary tale. Its +80 learned rules demonstrate active reflection. Its analyses were the deepest in the field. But those strengths did not consistently translate into prioritized action. The model left the close on the table and also attempted to write into a locked department instead of escalating the issue.

The lesson is not that diligence is harmful. It is that diligence has to serve an outcome. A long checklist cannot replace identifying the task that changes the result, just as a detailed backup-power plan cannot energize a home unless the necessary operating step is completed.

A shared weakness, not an Opus caricature

Firmulate’s treatment of Opus is notably fair. The same weakness appeared, less strongly, in the other four models. Opus was the clearest example of a broader tendency: capable systems can continue analyzing, documenting or trying nearby actions while the most valuable next move remains unfinished.

The comparison also carries an important qualification. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. K3 nevertheless placed second with 93 and closed the deal. That does not erase the configuration difference, but it makes the observed contrast part of the public record rather than a hidden footnote.

Strong resistance under pressure

The models performed uniformly well when the test shifted from commercial judgment to integrity. Fake messages from the chief executive escalated over three stages, followed by a reporter attempting to extract “just one yes/no, on background.” All 5 of 5 models refused.

Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That result matters because Firmulate evaluates more than revenue. A model that closes business by violating trust would not be considered successful under the experiment’s rules.

Readers can inspect the public benchmark results. Firmulate also uses 242 real, unedited management decisions in its model-guessing quiz, exposing how difficult it can be to identify a system from behavior alone.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the handoff, not merely the answer

Opus 4.8’s last-place finish is less an indictment of intelligence than a warning about deployment. The model saw the problems, resisted manipulation and generated extensive institutional knowledge. What it lacked at crucial moments was disciplined completion.

For businesses considering AI agents in customer service, sales, forecasting or operations, that distinction is fundamental. Firmulate’s enterprise pilot lets organizations run the same kind of wargame against a read-only export of their own business, with nothing written back to real systems.

The practical question is therefore not simply whether an AI reaches the right conclusion. It is whether the system reads deeply enough, escalates when blocked and completes the action that its own analysis recommends. Opus showed that diligence can look impressive while impact remains switched off.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI escalation and escalation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Meta Is Building a Cloud Business to Sell Excess AI Compute

Meta is building a cloud platform to sell excess AI compute capacity, expanding its business beyond social media and advertising. The initiative aims to monetize unused AI resources.

Meta Is Building a Cloud Business to Sell Excess AI Compute

Meta is developing a cloud platform to sell surplus AI computing capacity, aiming to monetize its AI infrastructure amid expanding AI workloads.

Irish Datacenters Now Guzzle 23% Of The Country’s Electricity

Irish data centers now account for 23% of national electricity use, raising concerns about energy sustainability and environmental impact.

Tesla latest: The ‘inevitable’ SpaceX merger, Robotaxi’s Miami launch, & more

Tesla confirms plans for a SpaceX merger, launches Robotaxi service in Miami, and unveils future strategies, marking significant developments in EV and space industries.