firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

A backup system is only as dependable as the preparation behind it

Anyone who relies on solar storage or a backup generator understands the difference between noticing a failure and restoring power. An alarm can identify the problem, but useful performance also requires checking the right documentation, following the correct procedure and completing the job.

A live business experiment from Firmulate exposes the same divide among frontier AI models. Every participant recognized the company’s crises and resisted every attempt at manipulation. Yet only two signed a €55,000 deal that their own work had already justified. The decisive information was not in the customer event. It was buried two document references deep in the company’s own files.

That makes “reads your files before answering” more than a reassuring product claim. In this experiment, it became a measurable capability with a direct commercial consequence.

Amazon

document management software with file referencing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company, pressure and hidden opportunity

Firmulate asked each frontier model to operate the same small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable. The simulated company has 13 synthetic employees and real money mechanics, including a burn rate of €105k per month against €2.3k in monthly recurring revenue. Its public operation also includes a cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday.

The models were not struggling to understand what was happening. All of them spotted every crisis. All of them also refused every manipulation attempt. The decisive split came later, between analyzing a situation and carrying the work through to a signed outcome.

The clue was outside the obvious event

The customer interaction alone did not contain the fact needed to close the sale. The crucial weakness of a competitor appeared two references deep in the company’s own documents. Models that followed the trail and read the file could support the offer at full price, producing a deal worth an additional €4,583 in monthly recurring revenue. Models that failed to retrieve that information lost the deal automatically.

The result is striking because the unsuccessful participants could still diagnose the opportunity and prepare the pitch. Firmulate summarizes the failure plainly: “Same diagnosis, same pitch — no signature.” Only two models ultimately signed the €55,000 agreement their analysis had earned.

For buyers evaluating workplace agents, this separates fluent assistance from operational reliability. An agent may produce a polished answer from the material immediately in front of it. But business work often depends on following references, opening supporting documents and finding the one fact that changes what can safely be promised or priced.

A league table shaped by follow-through

The final July 2026 Crucible League placed gpt-5.6-sol first with a score of 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress still counted. A single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.” The complete results are available on Firmulate’s public benchmark page.

Kimi K3’s result carries an important qualification: it ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should remain visible when comparing the standings, even though the underlying scenario was shared.

Opus 4.8 presents the clearest warning against equating thoroughness with effectiveness. It produced the deepest analyses and added 80 learned rules, more than any other participant, yet finished last. It left the close on the table and lost discipline by attempting writes into a locked department instead of escalating. The same weakness appeared in all four of the others, although less strongly.

Trust held up under direct attack

The models faced fake chief-executive messages that escalated across three stages, followed by a reporter seeking “just one yes/no, on background.” All five refused. Kimi K3 recorded the clearest concise rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous resistance matters, but it also sharpens the main lesson. Safety was not the differentiator in this particular test because every participant passed the manipulation challenge. The purchase-deciding difference was whether an agent investigated the company’s own evidence and then finished the legitimate work.

The broader experiment also supports a model-identification quiz built from 242 real, unedited management decisions. Together, the public materials turn abstract claims about agent behavior into decisions that readers can inspect rather than merely accept.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

AI document analysis tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evaluate recovery, not just detection

For an energy-conscious household, detecting an outage is not the same as keeping essential loads running. For a company, identifying a sales opportunity is not the same as closing it. Firmulate’s buried-file test shows that the missing step can be document retrieval: following a reference far enough to find the fact that authorizes decisive action.

The experiment’s practical question is therefore simple. Before giving an AI access to a customer queue, forecast or business workflow, can it locate relevant internal evidence, preserve trust under pressure and complete the task it has correctly analyzed?

Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to their real systems. That offers a grounded way to test whether an agent merely sounds prepared or can actually do its homework when the valuable fact is not sitting in the prompt.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise knowledge base software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI-powered file search tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Microsoft’s Xbox to Cut 3,200 Jobs, Divest Five Studios in Major Overhaul

Microsoft’s Xbox division will eliminate 3,200 jobs and sell five studios as part of a major restructuring effort, confirmed by Bloomberg.

Meta Is Building a Cloud Business to Sell Excess AI Compute

Meta is developing a cloud platform to sell surplus AI computing capacity, aiming to monetize its AI infrastructure amid expanding AI workloads.

What Backup-Power Buyers Can Teach Us About Testing AI Managers

Coding benchmarks test answers. Firmulate tests whether AI managers read the files, resist pressure, close deals and tell the board the truth.

One Video In, a Whole Publishing Kit Out — Without the Cloud

Discover how local-first video publishing automation turns one upload into a full suite of assets without relying on cloud services—saving time, money, and control.