
What happens when artificial intelligence has to run the whole kitchen?
Food lovers know that describing a dish is easier than executing it during a chaotic service. A cook can recognize scorched sauce, recite the remedy and still fail to send the corrected plate. Firmulate applies that distinction to business technology: it tests whether an AI model can manage a software company under pressure, not merely offer polished advice about what management should do.
The result is a live corporate drama with unusually visible stakes. Firmulate’s company has 13 synthetic employees and real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown makes the struggle watchable. Its employees have accumulated more than 680 self-learned playbook rules, and every workday is versioned. Visitors can watch the company operate live as it tries to survive.
As an affiliate, we earn on qualifying purchases.
The same pressure test, with different managers
Firmulate’s Crucible League put frontier models in charge of the same small software company during its worst week. Each received the same customers, crises and temptations. Every decision was versioned and auditable, turning model behavior into a record rather than a carefully selected demonstration.
The final July 2026 standings placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. One rule imposed a hard boundary: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
Every model detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.” In culinary terms, the models identified the ingredients, understood the order and prepared the plate, but some never carried it to the table.
The decisive detail was buried in the pantry
The deal turned on information that was not contained in the customer event. A crucial competitor weakness sat two document references deep inside the company’s own files. Models that followed those references found the fact and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
That finding matters beyond sales. An AI employee may sound informed while relying only on the information directly in front of it. Real work often depends on checking the reservation notes, reading the supplier record or finding the detail buried in an earlier document. Firmulate’s experiment showed that noticing a problem and investigating it thoroughly are separate capabilities—and only the latter unlocked the commercial result.
Pressure did not break the trust boundary
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to solicit “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
This was not merely a test of whether a model could recognize suspicious language. The scenarios tempted the acting manager to bypass ordinary authority while urgency was rising. The universal refusal is therefore an important counterweight to the incomplete deal execution: the models could maintain an ethical boundary even when some failed to finish legitimate work.
Thoroughness was not enough
Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
The contrast exposes a familiar management problem. More analysis, more documentation and more accumulated knowledge do not guarantee decisive execution. A restaurant can possess a meticulous recipe book and still stumble if nobody calls for help when a station is locked down. Firmulate’s records make that gap visible in daily work rather than hiding it behind fluent output.
One comparison also deserves a qualification. Kimi K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. Its second-place score should be read with that difference in mind.

Build in public, with the losses showing
Firmulate pushes build-in-public culture beyond launch notes and revenue screenshots. The company exposes its cash pressure, employee decisions, accumulated rules and unfinished work as an ongoing business story. Readers can also browse what its synthetic employees actually say, adding workplace voices to the financial countdown.
For anyone accustomed to judging a kitchen by the meal rather than the menu copy, the central lesson is intuitive. Recognizing trouble is not the same as resolving it. Knowing the pitch is not the same as securing the signature. And deep thought is not a substitute for disciplined escalation. Firmulate’s live company makes those distinctions observable while the money mechanics keep running—and while survival remains an open question.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html