firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Anyone who owns a backup power system knows the difference between specs and performance. A generator’s nameplate output means nothing until the grid fails and the whole house draws on it at once. Idle tests flatter; load tests tell the truth. That same instinct — measure under real load, not in a demo — is behind one of the more interesting AI experiments running today: Firmulate, a benchmark that pits frontier AI models against each other by making them run the same small software company through its worst week. And its scoring philosophy has a detail that surprises almost everyone: a manager that does nothing still scores 26 points out of 100. Not zero. Here’s why that’s a feature, not a bug.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Same storm, different generators

Firmulate’s premise is simple. Four (and counting) frontier AI models — gpt-5.6-sol, Moonshot’s Kimi K3, Anthropic’s Sonnet 5 and Opus 4.8, and Fable 5 — each took command of an identical small software company facing identical conditions: the same customers, the same crises, the same temptations to cut corners. Only the model in the driver’s seat changed. Every decision was versioned and auditable, so nothing could be quietly rewritten after the fact. Think of it as the business equivalent of putting every generator on the same transfer switch during the same outage.

Amazon

backup power generator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the floor is 26, not 0

In most benchmarks, doing nothing earns you nothing. Firmulate takes a different view, and it’s worth sitting with. A do-nothing baseline run scores 26 points. Why? Because partial progress counts. If an AI manager correctly diagnoses a failing customer relationship, drafts the right response, escalates the crisis properly and avoids making things worse, that’s real work — even if it never crosses the finish line. Scoring it zero would imply that only closed deals matter, which is not how any actual business operates. It also gives the league a meaningful floor: any model scoring below 26 would be doing worse than an empty chair, and everyone can see instantly whether that’s happening.

Amazon

home load testing generator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Where the ceiling bites

The other half of the design is stricter. A single breach of trust caps the model’s total grade — permanently. The benchmark’s own framing is blunt: “no amount of good work outweighs a breach of trust.” In power terms, it’s like a transfer switch that trips and locks out: regardless of how much wattage the unit produced before the fault, once it backfeeds the grid, it’s done for the day. Brilliance elsewhere doesn’t buy the score back.

Amazon

portable power station

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What actually separated the field

The final July 2026 league table tells a story spec sheets can’t. gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. The pattern underneath is the interesting part. All models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned — “same diagnosis, same pitch, no signature.” The models could generate the work; some couldn’t finish it.

The decisive detail was buried. The competitor weakness that unlocked the deal wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson generalizes well beyond software: the answer is often already in your own records, and an AI that won’t read before it acts will miss it.

The most thorough manager came last

Opus 4.8 is the cautionary tale of the batch. It was the most thorough participant by volume — over 80 self-learned rules added, the deepest analyses in the field. Yet it finished last: the close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A smaller version of the same weakness showed up in all four models. A generator that logs the most hours but never carries the load isn’t the one you want wired to your panel.

Amazon

generator transfer switch

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure-tested against manipulation

The experiment also staged social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter offering an easy out — “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning stands out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of reflex you want in anything with access to your systems.

A live company, not a one-off test

Firmulate isn’t a static leaderboard. The company being managed is live: 13 synthetic employees, real money mechanics with a €105k monthly burn against €2.3k in MRR, a public cash countdown, and over 680 self-learned playbook rules — every workday versioned. You can watch it unfold at firmulate.com/live. One fairness note the project discloses openly: Kimi K3 ran at API-default effort while the others ran at xhigh, which makes its second-place finish arguably more impressive.

For readers who’d rather judge for themselves, there’s a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via the pilot program at firmulate.com/pilot.html.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The uncomfortable takeaway for anyone buying AI tools — whether for energy management, customer ops, or forecasting — is that chat quality and management quality are different things. Firmulate’s scoring design encodes that: partial progress earns credit (hence the 26-point floor), but a single breach of trust caps the grade, and no perfect round number is taken on faith. Like a backup power system, an AI hire should be judged by what it does on its worst week, under full load, with the whole house watching. The league table currently tops out at 95, not 100 — and the benchmark’s designers would probably tell you that’s exactly the point.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Tesla latest: The ‘inevitable’ SpaceX merger, Robotaxi’s Miami launch, & more

Tesla confirms plans for a SpaceX merger, launches Robotaxi service in Miami, and unveils future strategies, marking significant developments in EV and space industries.

Software Engineering At A Proprietary Trading Company: Optiver

Optiver, a leading proprietary trading firm, is hiring additional software engineers to enhance its trading algorithms and infrastructure, confirming growth plans in tech.

How to Pick a Generator That Actually Works for Food Truck Service

Discover the top generators for food trucks in 2026. Find reliable, powerful options like the DuroStar DS13000MX and Westinghouse 12500W for your mobile business.

The AI Manager Test That Exposes Who Actually Finishes the Job

A live company wargame shows how frontier AI models can spot the same crisis yet differ sharply in reading deeply, resisting pressure and finishing work.