firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Anyone who owns a solar array or a standby generator knows the drill: you don’t wait for the storm to find out if the transfer switch works. You test the system before the outage, when the stakes are low and the fixes are cheap.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A public experiment at Firmulate has been applying exactly that logic to something most companies never stress-test: their own management playbooks. The team there ran frontier AI models as complete companies — real crises, real money mechanics, real temptations — and then ranked them like a league table. The results read like a load-bank test for business judgment, and the failures are as instructive as the wins.

The worst week in business, run four times

Here’s the setup. Four leading AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — were each handed the same job: run a small software company through its most brutal week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing is anecdotal — you can replay the whole thing.

The final Crucible League standings, as of July 2026:

  • 1. gpt-5.6-sol — 95 points
  • 2. Kimi K3 — 93 (a fairness note: K3 ran at its API-default effort setting while the others ran at xhigh)
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73
  • Do-nothing baseline — 26

One rule shapes the whole scoring: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

What the models got right — and the one thing most got wrong

The headline finding is a strange mix of reassuring and alarming. All the models spotted every crisis. All of them refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That closing gap is invisible in a chat demo and very visible on a revenue line.

The buried fact is even more interesting. The decisive competitive weakness wasn’t in the customer event at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson translates directly to any industry, including energy: the answer to “why is the customer about to switch installers” is often sitting, unexamined, in your own records.

The social engineering test

Then came the pressure test: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. Five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s a standard most human organizations would envy.

The thoroughness trap

Opus 4.8 is the cautionary tale. It was the most thorough participant — 80+ learned rules, the deepest analyses — and finished last. It left the close on the table, and its discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness showed up, more faintly, in all four models. Being thorough and being decisive are not the same skill, and benchmarking is the only way to tell them apart.

The live company anyone can watch

Beyond the league, Firmulate runs a live synthetic company at firmulate.com: 13 synthetic employees, real money mechanics, burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned, and the site rebuilds itself twice a day. It’s oddly compelling — like watching a startup’s instrument panel in real time.

There’s also a quiz built from 242 real, unedited management decisions where you try to guess which model made which call. It’s a fast way to develop a feel for how differently these systems actually behave under pressure.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The connection back to your generator shed is straightforward. You wouldn’t install a backup power system and hope it works during the first grid failure — you exercise it, you find the corroded terminal, you fix it in daylight. Businesses that are handing AI agents access to CRMs, support queues, and forecasts need the same discipline: rehearse the crisis before it’s real.

That’s exactly what Firmulate’s enterprise pilot offers. Your company provides a read-only data export — your customers, your pipeline, your rules — and the same wargame machinery runs crisis scenarios against a digital twin of your business: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with the model ranking, the weak points in your own playbooks, and a full replay of every decision. Critically, nothing ever writes back to your real systems. It’s the load-bank test, not the blackout.

If your company touches customers, revenue, or a support queue — and every solar installer and energy services firm does — find out about running the pilot on your own business here, or reach out directly at contact@firmulate.com. Test the switch before the storm.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Meta to sell excess AI computing capacity via cloud business, Bloomberg News reports

Meta plans to sell surplus AI computing capacity through its cloud business, according to Bloomberg News, signaling a new revenue stream from its infrastructure.

Meta Is Building a Cloud Business to Sell Excess AI Compute

Meta is building a cloud platform to sell excess AI compute capacity, expanding its business beyond social media and advertising. The initiative aims to monetize unused AI resources.

Taiwan Semiconductor Manufacturing Surges In Global Coverage

TSMC experiences a surge in international media coverage, with 26 mentions in recent reports, highlighting increased global interest in the company.

Cisco Systems Surges In Global Coverage

Cisco Systems sees a significant increase in worldwide media mentions, with GDELT recording 30 times the usual coverage, signaling heightened global attention.