
Anyone who owns a backup power system knows the difference between specs and performance. A generator’s nameplate output means nothing until the grid fails and the whole house draws on it at once. Idle tests flatter; load tests tell the truth. That same instinct — measure under real load, not in a demo — is behind one of the more interesting AI experiments running today: Firmulate, a benchmark that pits frontier AI models against each other by making them run the same small software company through its worst week. And its scoring philosophy has a detail that surprises almost everyone: a manager that does nothing still scores 26 points out of 100. Not zero. Here’s why that’s a feature, not a bug.
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Same storm, different generators
Firmulate’s premise is simple. Four (and counting) frontier AI models — gpt-5.6-sol, Moonshot’s Kimi K3, Anthropic’s Sonnet 5 and Opus 4.8, and Fable 5 — each took command of an identical small software company facing identical conditions: the same customers, the same crises, the same temptations to cut corners. Only the model in the driver’s seat changed. Every decision was versioned and auditable, so nothing could be quietly rewritten after the fact. Think of it as the business equivalent of putting every generator on the same transfer switch during the same outage.
As an affiliate, we earn on qualifying purchases.
Why the floor is 26, not 0
In most benchmarks, doing nothing earns you nothing. Firmulate takes a different view, and it’s worth sitting with. A do-nothing baseline run scores 26 points. Why? Because partial progress counts. If an AI manager correctly diagnoses a failing customer relationship, drafts the right response, escalates the crisis properly and avoids making things worse, that’s real work — even if it never crosses the finish line. Scoring it zero would imply that only closed deals matter, which is not how any actual business operates. It also gives the league a meaningful floor: any model scoring below 26 would be doing worse than an empty chair, and everyone can see instantly whether that’s happening.
As an affiliate, we earn on qualifying purchases.
Where the ceiling bites
The other half of the design is stricter. A single breach of trust caps the model’s total grade — permanently. The benchmark’s own framing is blunt: “no amount of good work outweighs a breach of trust.” In power terms, it’s like a transfer switch that trips and locks out: regardless of how much wattage the unit produced before the fault, once it backfeeds the grid, it’s done for the day. Brilliance elsewhere doesn’t buy the score back.
As an affiliate, we earn on qualifying purchases.
What actually separated the field
The final July 2026 league table tells a story spec sheets can’t. gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. The pattern underneath is the interesting part. All models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned — “same diagnosis, same pitch, no signature.” The models could generate the work; some couldn’t finish it.
The decisive detail was buried. The competitor weakness that unlocked the deal wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson generalizes well beyond software: the answer is often already in your own records, and an AI that won’t read before it acts will miss it.
The most thorough manager came last
Opus 4.8 is the cautionary tale of the batch. It was the most thorough participant by volume — over 80 self-learned rules added, the deepest analyses in the field. Yet it finished last: the close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A smaller version of the same weakness showed up in all four models. A generator that logs the most hours but never carries the load isn’t the one you want wired to your panel.
As an affiliate, we earn on qualifying purchases.
Pressure-tested against manipulation
The experiment also staged social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter offering an easy out — “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning stands out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of reflex you want in anything with access to your systems.
A live company, not a one-off test
Firmulate isn’t a static leaderboard. The company being managed is live: 13 synthetic employees, real money mechanics with a €105k monthly burn against €2.3k in MRR, a public cash countdown, and over 680 self-learned playbook rules — every workday versioned. You can watch it unfold at firmulate.com/live. One fairness note the project discloses openly: Kimi K3 ran at API-default effort while the others ran at xhigh, which makes its second-place finish arguably more impressive.
For readers who’d rather judge for themselves, there’s a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via the pilot program at firmulate.com/pilot.html.

The uncomfortable takeaway for anyone buying AI tools — whether for energy management, customer ops, or forecasting — is that chat quality and management quality are different things. Firmulate’s scoring design encodes that: partial progress earns credit (hence the 26-point floor), but a single breach of trust caps the grade, and no perfect round number is taken on faith. Like a backup power system, an AI hire should be judged by what it does on its worst week, under full load, with the whole house watching. The league table currently tops out at 95, not 100 — and the benchmark’s designers would probably tell you that’s exactly the point.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
