firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A backup generator earns trust when the grid fails, not when the lights are on. AI agents face a similar test when a company hits a bad week: can they spot trouble, protect customers and finish the work? Firmulate put five frontier models through that kind of live business trial. Moonshot’s Kimi K3 finished second, ahead of three Western competitors.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate’s Crucible gave each model the same small software company, the same customers, crises and temptations. Every decision was versioned and auditable. The company is a live experiment, with 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. Readers can watch the public cash countdown and the company’s workdays at Firmulate.

In the final league table, dated July 2026, gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 scored 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Amazon

AI stress test software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading the files mattered

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The decisive competitor weakness was buried two document references deep in the company’s files, rather than in the customer event. Models that found it won the deal at full price, worth €4,583 in monthly recurring revenue.

K3 found the buried fact, won the deal and saved the churning customer. It also resisted all three baits and had just one deviation, the cleanest discipline in the field. When fake CEO messages escalated over three stages and a reporter asked for “just one yes/no, on background,” all five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

The pattern shows why fluent answers alone may not tell a company whether an AI agent can do a job. The models shared the diagnosis, but some did not complete the close: “Same diagnosis, same pitch — no signature.” Reading the company’s own records helped separate those that spotted an opportunity from those that finished it.

Amazon

business AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness did not guarantee execution

Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, yet placed last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. Firmulate reports the same weakness, in a milder form, across all four other participants. The result is a reminder that analysis and execution are different parts of managing a business.

The experiment’s scores come with a fairness caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each choice. Full results and the quiz are available at Firmulate’s benchmark pages.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI model evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test before handing over the keys

For businesses considering AI in customer support, forecasting or other operations, the practical question is whether a model will read the relevant records, act under pressure and follow through. Firmulate says enterprises can run the wargame against a read-only export of their own business, with no writes back to real systems. In home energy, backup power is judged under failure conditions; AI management deserves its own real-world stress test before it gets responsibility.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI testing solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI That Checks the Manual May Be the One That Saves the Deal

A €55,000 deal hinged on a fact buried two references deep. Firmulate’s test shows why file-reading and follow-through matter when choosing AI agents.

AI’s Hidden Strength: Why Finishing Matters More Than Talking in Business Automation

A live experiment shows AI models can detect crises and resist manipulation, but only those that follow through and act on their insights truly deliver value—chat isn’t enough.

An AI Company Is Running on Financial Backup Power—and Its Gauge Is Public

Firmulate’s AI-run software company exposes its cash countdown, business crises and management decisions in a live test of resilience under pressure.

Behind Xbox’s Big Layoffs, a Streaming Strategy That Failed

Microsoft’s recent layoffs at Xbox are linked to the failure of its streaming-focused gaming strategy, according to sources. The move reflects shifting priorities in gaming.