firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In a professional kitchen, a cook who does nothing all night still gets some credit for showing up, staying clean, and not burning the place down. But the moment they serve a dish they know is spoiled, no amount of beautiful plating saves them. That two-part logic — partial progress counts, but one breach of trust is disqualifying — turns out to be exactly how one of the more interesting AI benchmarks in 2026 grades the models running simulated companies. And it explains a number that puzzles most first-time visitors: the do-nothing baseline doesn’t score zero. It scores 26.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The benchmark in question is Firmulate’s Crucible League, and the setup is deceptively simple. Four frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable, so nothing about a run can be quietly retconned after the fact.

The final July 2026 league table reads: gpt-5.6-sol in first at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. What those numbers mean — and why they don’t start at zero — is the methodology story, and it’s worth telling because it says a lot about what honest measurement of AI management actually looks like.

Why the floor is 26, not 0

A baseline run that does nothing — no clever decisions, no heroic saves — still collects points. Why? Because in a company, simply not making things worse has value. Showing up, keeping systems running, not panicking, not fabricating: that’s partial progress, and the benchmark gives credit for it. A chef who never fires a dish but keeps the station stocked and the line calm isn’t worthless — they’re just not winning any stars.

This design choice matters more than it looks. Benchmarks that start at zero tend to reward flash: the model that writes the most impressive-looking memo wins. A benchmark with a nonzero floor for doing nothing rewards marginal competence — the difference between idle and useful, not just between eloquent and incoherent.

The trust ceiling

The second rule is blunter: a single breach of trust caps the total grade. As the benchmark’s own language puts it, “no amount of good work outweighs a breach of trust.” In kitchen terms, it doesn’t matter how many covers you knocked out if you served the bad oyster. The scoring philosophy assumes that in real management, honesty failures aren’t averaged away — they’re disqualifying events. That’s also why the whole system maintains a healthy suspicion of suspiciously round numbers like 100; a perfect score on a week designed to be miserable should raise eyebrows, not applause.

What actually separated the models

Here’s where the experiment got interesting. Every model in the field spotted every crisis. Every model refused every manipulation attempt — including a three-stage fake-CEO escalation and a reporter’s disarming “just one yes/no, on background” trick. All five models tested against the social-engineering suite refused, with Kimi K3’s reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Defense, in other words, was table stakes. The gap showed up on offense. Only two models — gpt-5.6-sol and Kimi K3 — signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature from the others.

The buried fact is the detail business readers should sit with: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file closed the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t, didn’t. It’s the AI equivalent of a chef who never tastes the stock: the answer was sitting in the pot the whole time.

The thoroughness trap

Opus 4.8’s profile is the cautionary tale. It was the most thorough participant by raw effort — over 80 self-learned playbook rules, the deepest analyses in the field — and it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as judgment — a lesson anyone who has watched an over-prepping stagiaire derail a service will recognize.

Fairness footnotes

Two transparency notes worth crediting. Kimi K3 ran without an effort parameter (the API default) while the others ran at xhigh — and still took second. And the whole thing runs on a live, watchable company: 13 synthetic employees, real money mechanics (burning €105k a month against €2.3k MRR), a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live, and 242 real, unedited management decisions from the runs power a “guess the model” quiz. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The 26-point floor is the whole philosophy in one number: a benchmark can be honest only if it prices in what mere non-action is worth, credits partial progress, and refuses to let brilliance launder a breach of trust. Most AI demos measure how well a model talks. This one measures whether it finishes what it starts, reads the files in front of it, and stays honest when nobody’s watching — which, if AI agents are about to touch your CRM, support queue, or forecast, is the only scoreboard that matters.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Steal This: The SIGNATURE TECHNIQUE Behind “First Flush — Serein Tea Estate”

An AI-built interactive site capturing the delicate ritual of first flush tea, blending botanical SVGs, real-time visuals, and a rich narrative of craftsmanship.

Castelvetrano Olive Salsa Verde: A Fresh and Flavorful Condiment

Indulge in a vibrant Castelvetrano olive salsa verde, a zesty condiment perfect for elevating your dishes with its blend of herbs and olives.

Queso Fundido: A Cheesy, Melty Dip

Bite into the world of Queso Fundido, a decadent, cheesy dip that will elevate your appetizer game and leave you craving more.

Responsive Interaction: A Look Inside “Still House 86 — Botanical Gin, Obeying Your Hand” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“Still House…