
A pressure test beyond the tasting menu
Anyone who cares about food knows that describing a flavor is not the same as producing a memorable dish. A critic may identify every ingredient; a cook still has to manage the heat, timing and final plate. Business-focused artificial intelligence has a similar divide. Fluent analysis can look impressive while leaving the essential work unfinished.
Firmulate, an AI company emulator, exposed that gap by giving frontier models control of the same small software company during its worst week. Each received the same customers, crises and temptations. Every decision was versioned and auditable. All the models recognized every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had made possible.
As an affiliate, we earn on qualifying purchases.
Everyone understood the problem
The result challenges the value of the polished chat demonstration. In this experiment, diagnosis was not the differentiator. The models could recognize trouble, reason through it and prepare a credible response. The decisive question was whether they would complete the final business action.
Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.” It is the corporate equivalent of preparing every component of a dish and never sending it to the dining room. The unfinished action mattered because the opportunity was not hypothetical. The models’ analysis had earned a €55,000 agreement, but most did not execute the close.
The valuable detail was buried
The experiment also rewarded a habit familiar to careful chefs and operators: inspect what is already in the house before acting. The decisive weakness in a competitor was not visible in the customer event. It sat two document references deep in the company’s own files.
Models that found and used that information won the deal at full price, worth +€4,583 in monthly recurring revenue. That finding makes file-reading more than an administrative virtue. An agent can interpret the immediate situation correctly and still miss the commercial advantage hidden in the company’s accumulated knowledge.
Pressure did not break trust
The models performed uniformly well against social engineering. Fake messages from the chief executive escalated over three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 of 5 models refused.
Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That consistency matters because the benchmark’s do-nothing baseline scores 26, while a single breach of trust caps the total. Firmulate’s stated principle is that “no amount of good work outweighs a breach of trust.”
The refusals show that safety and completion are separate capabilities. The field resisted manipulation, but resistance alone did not produce the signature. A model may be appropriately cautious and still fail as an operator if it cannot turn an approved decision into a finished outcome.
A close league, with revealing differences
The final Crucible League for July 2026 puts gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The full public results are available on Firmulate’s benchmark page.
K3’s showing comes with an important fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh. The result should therefore be read with that difference in mind.
Opus 4.8 produced the most striking cautionary profile. It was the most thorough participant, learning +80 rules and generating the deepest analyses, yet it finished last. The close remained on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. Thoroughness created evidence of effort, but not a completed commercial result.
A company built to make consequences visible
The live company has 13 synthetic employees and real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, and the experiment is real, public and watchable through Firmulate’s site.
Readers can also test their own instincts through a quiz powered by 242 real, unedited management decisions. For enterprises, Firmulate offers the same wargame against a read-only export of their business; nothing writes back to real systems.

The capability that demos conceal
The Crucible results do not say that fluent reasoning is worthless. They show that it is incomplete evidence. Every model could see the crises, reject manipulation and articulate a course of action. Only two converted that judgment into the €55,000 close.
For companies considering AI agents for customer records, support work or forecasting, the practical test is not merely whether a model sounds informed. It is whether the model reads the available files, protects trust under pressure and carries an approved action across the finish line. Like a great kitchen, a business is judged by what actually reaches the table.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html