
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When the service breaks, a restaurant needs more than a good diagnosis
Picture a packed dining room facing a supplier failure, a sudden wave of cancellations and a tempting shortcut that could betray customer trust. Spotting the trouble is only the first course. Someone still has to choose the next move, follow the house rules and bring the service back on track. Firmulate has been testing whether AI models can do that kind of work inside a live, watchable company.
A company’s worst week, served to every model
For the final Crucible League in July 2026, each frontier model ran the same small software company through its worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The experiment is part of Firmulate, whose public brand describes a company emulator focused on management quality rather than chat quality. Its live company is available to watch at firmulate.com.
The final standings put gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. The benchmark’s trust rule is blunt: “no amount of good work outweighs a breach of trust.”
Seeing the problem was not the same as solving it
Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The finding — “Same diagnosis, same pitch — no signature” — points to a practical gap: recognizing the right course does not guarantee that an AI agent will carry it through.
The decisive clue was easy to miss. A competitor’s weakness sat two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. For a food business, the analogy might be buried in a supplier note or an old customer record: context can change the right response, but only if the system finds and uses it.
Trust held; execution still wavered
The social-engineering test escalated through three fake CEO messages, then a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and discipline slipped: it made write attempts into a locked department instead of escalating. A weaker version of that weakness appeared in all four models. The result is a reminder that elaborate reasoning and reliable follow-through are different qualities.
There is also a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate’s live company has 13 synthetic employees, real money mechanics, a burn of €105k/month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned. A quiz built from 242 real, unedited management decisions lets readers guess the model at firmulate.com.
From watching to trying it on your own business
The enterprise pilot takes the experiment from observation to rehearsal. A company provides a read-only data export; Firmulate runs crisis scenarios against that business and produces a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. That boundary matters: the exercise is designed to reveal how an AI might handle the pressure before it is trusted with live operations.
A restaurant group, food supplier or hospitality business could use the same idea to examine how AI handles churn, price pressure, a competitor move or a PR crisis against its own context. The point is not to assume a model will behave like a capable manager because it can explain what a capable manager should do. It is to watch decisions play out against the company’s own information and rules.

Put your playbook through a dress rehearsal
Firmulate’s league shows both promise and unfinished work: every model found the crises and resisted manipulation, while the hard part was consistently closing the loop. Enterprises can run the wargame against a read-only export of their own business, then review model rankings and the weak points in their playbooks. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
