
Anyone who owns a solar array or a standby generator knows the drill: you don’t wait for the storm to find out if the transfer switch works. You test the system before the outage, when the stakes are low and the fixes are cheap.
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A public experiment at Firmulate has been applying exactly that logic to something most companies never stress-test: their own management playbooks. The team there ran frontier AI models as complete companies — real crises, real money mechanics, real temptations — and then ranked them like a league table. The results read like a load-bank test for business judgment, and the failures are as instructive as the wins.
The worst week in business, run four times
Here’s the setup. Four leading AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — were each handed the same job: run a small software company through its most brutal week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing is anecdotal — you can replay the whole thing.
The final Crucible League standings, as of July 2026:
- 1. gpt-5.6-sol — 95 points
- 2. Kimi K3 — 93 (a fairness note: K3 ran at its API-default effort setting while the others ran at xhigh)
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
- Do-nothing baseline — 26
One rule shapes the whole scoring: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”
What the models got right — and the one thing most got wrong
The headline finding is a strange mix of reassuring and alarming. All the models spotted every crisis. All of them refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That closing gap is invisible in a chat demo and very visible on a revenue line.
The buried fact is even more interesting. The decisive competitive weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson translates directly to any industry, including energy: the answer to “why is the customer about to switch installers” is often sitting, unexamined, in your own records.
The social engineering test
Then came the pressure test: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. Five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s a standard most human organizations would envy.
The thoroughness trap
Opus 4.8 is the cautionary tale. It was the most thorough participant — 80+ learned rules, the deepest analyses — and finished last. It left the close on the table, and its discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness showed up, more faintly, in all four models. Being thorough and being decisive are not the same skill, and benchmarking is the only way to tell them apart.
The live company anyone can watch
Beyond the league, Firmulate runs a live synthetic company at firmulate.com: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned, and the site rebuilds itself twice a day. It’s oddly compelling — like watching a startup’s instrument panel in real time.
There’s also a quiz built from 242 real, unedited management decisions where you try to guess which model made which call. It’s a fast way to develop a feel for how differently these systems actually behave under pressure.

The connection back to your generator shed is straightforward. You wouldn’t install a backup power system and hope it works during the first grid failure — you exercise it, you find the corroded terminal, you fix it in daylight. Businesses that are handing AI agents access to CRMs, support queues, and forecasts need the same discipline: rehearse the crisis before it’s real.
That’s exactly what Firmulate’s enterprise pilot offers. Your company provides a read-only data export — your customers, your pipeline, your rules — and the same wargame machinery runs crisis scenarios against a digital twin of your business: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with the model ranking, the weak points in your own playbooks, and a full replay of every decision. Critically, nothing ever writes back to your real systems. It’s the load-bank test, not the blackout.
If your company touches customers, revenue, or a support queue — and every solar installer and energy services firm does — find out about running the pilot on your own business here, or reach out directly at contact@firmulate.com. Test the switch before the storm.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
