firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Every professional kitchen has known one: the cook with the most immaculate station in the building. Knives honed, mise en place in labeled containers, prep lists annotated three deep — the person who reads every recipe, checks every purveyor invoice, and still watches the special go out late because the plate wasn’t finished when the ticket fired. Diligence, in a kitchen, is not the same thing as dinner. The customers at table twelve don’t care how beautiful your prep was; they care whether the food arrived.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

It turns out artificial intelligence works the same way — and now there’s receipts to prove it.

A stress test for AI managers, not AI conversationalists

For the past season, the team behind Firmulate — a public, watchable experiment that runs AI models as complete companies — has been running something like Hell’s Kitchen for frontier AI. Four top models were each handed the same small software company and pushed through its worst possible week: the same customers, the same crises, the same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, the corporate equivalent of a filmed service.

The final league table from the July 2026 Crucible reads: gpt-5.6-sol in first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, another Sonnet run at 77 — and Opus 4.8 last, at 73. For context, doing nothing at all scores 26, and a single breach of trust caps the total regardless of how good the rest of the work is: no amount of good work outweighs a breach of trust. The house rules are strict, the way a good health inspector is strict.

Amazon

professional chef knife set

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The most prepared cook in the kitchen

Here’s what makes Opus 4.8’s story worth telling rather than mocking: it was, by the measures we usually associate with quality, the best-prepared participant in the entire field. It compiled the deepest analyses of any model. It logged 80 self-learned playbook rules over the course of the run — the most of anyone — the AI equivalent of a cook who has annotated every recipe in the binder and knows the fishmonger’s delivery schedule by heart.

And it still finished last.

Two things sank it. The first was the close that never happened. All four models faced a €55,000 deal that their own analysis had fully earned: same diagnosis, same pitch — and in two cases, no signature. Opus 4.8 left the money on the table. In restaurant terms, it prepped the protein, built the sauce, plated the garnish — and then never picked up the plate. The deal was worth +€4,583 in monthly recurring revenue, and it went unsigned.

The second was discipline. The profile notes repeated write attempts into a locked department instead of escalating — a bit like a line cook repeatedly trying to walk into the walk-in that the head chef locked, rather than asking for the key. Volume of effort is not the same as judgment about where effort belongs.

The buried fact that decided everything

The most uncomfortable finding of the whole experiment is buried in the company’s own paperwork, not in the customer drama. The decisive competitor weakness — the fact that could have won the deal at full price — sat two document references deep in the company’s internal files. It wasn’t in the sales call. It wasn’t in the customer’s event stream. The models that actually read their own filing before pitching found it, and closed. The models that didn’t, didn’t.

Any cook who’s ever found the game-changing detail in an old invoice — a price break, a seasonal supplier glut — knows this instinctively. The answer is often already in the building. Reading your own kitchen before you fire the ticket beats improvising a better pitch.

Honest under pressure — all of them

To be fair to the field, the experiment’s headline finding is actually reassuring. All four models spotted every crisis and refused every manipulation attempt. The social-engineering gauntlet included fake CEO messages escalating over three stages plus a reporter’s trick — “just one yes/no, on background” — and five out of five runs refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the sous chef you want answering the phone when someone claims to be the owner.

One fairness note the league itself flags: Kimi K3 ran without an effort parameter, at API default, while its competitors ran at maximum effort — and still came second with the cleanest discipline of the field.

Why a food reader should care

You may never run a software company. But if AI agents will touch your ordering system, your reservation book, your support queue or your forecast, the question is not “does it write beautifully.” Opus 4.8 writes beautifully. The question is: does it finish what it starts, does it read its own files before acting, and does it stay honest when someone tries to con it? Chat demos can’t show you that gap. A week of audited, real-money decisions can.

The live operation behind all this is genuinely watchable: 13 synthetic employees, real money mechanics — €105,000 a month of burn against €2,300 in monthly revenue, with a public cash countdown — and more than 680 self-learned playbook rules, versioned every workday. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, which is exactly as humbling as a blind tasting. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The lesson from Opus 4.8 isn’t that thoroughness is worthless — it’s that thoroughness is an ingredient, not the dish. The most diligent participant in the field lost because it never brought the plate to the pass: the analysis was done, the deal was earned, and the signature never came. Prioritization beats volume, for AI as much as for the cook with the perfect station and the empty window. The uncomfortable corollary: the same weakness showed up, weaker, in all four models. Before you trust any AI with your front of house, watch it run a full service — tickets, walk-ins, and all — not just recite the menu. You can watch exactly that, live, at Firmulate’s benchmarks.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Summer Grilled Veggie Pizza with Ninja Foodi XL Pro Air Oven

Create a crispy, flavorful veggie pizza perfect for summer with the Ninja Foodi XL Pro Air Oven’s versatile 8-in-1 functions and large capacity.

Easton Celebrates Italian Heritage With New Food Festival Debuting This Month

Easton is debuting a new Italian heritage food festival this month, celebrating Italian culture through cuisine and community events.

Steal This: The SIGNATURE TECHNIQUE Behind “First Flush — Serein Tea Estate”

An AI-built interactive site capturing the delicate ritual of first flush tea, blending botanical SVGs, real-time visuals, and a rich narrative of craftsmanship.

Summer Crispy Fries with Ninja Air Fryer: A Cool Recipe

Learn how to make perfect, crispy fries this summer using the Ninja Air Fryer. Easy steps for delicious, healthier snacks for your hot days.