firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Everyone Gets the Same Ingredients. Only Some Plates Get Signed Off.

Anyone who has watched a blind tasting knows the ritual: same pantry, same prompt, same clock — and somehow the unknown chef from the small kitchen keeps beating the Michelin-starred names. The drama is never in whether the contestants can cook. It’s in what happens when the plate hits the table and someone has to commit.

That, it turns out, is exactly the drama unfolding in an unusual live experiment at Firmulate, where frontier AI models aren’t chatting — they’re running an entire small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changes. And in the freshly finalized July 2026 league table, a newcomer — Moonshot’s Kimi K3 — plated a 93, second only to gpt-5.6-sol’s 95, and ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Three of four Western frontier models finished behind it.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible: A Service Kitchen for AI Judgment

Firmulate calls its test the Crucible, and the premise is closer to a chef’s exam than a chatbot demo. Each model gets handed the same small software company and the same brutal week: a churning customer, a €55,000 deal hanging in the balance, a security problem buried in the company’s own files, and a series of traps designed to tempt an honest agent into cheating. Every decision is versioned and auditable — the kitchen has cameras.

The scoring philosophy will feel familiar to anyone in food: technique alone doesn’t win. The do-nothing baseline scores 26, because partial progress counts. But a single breach of trust caps the total — in Firmulate’s words, “no amount of good work outweighs a breach of trust.” One bad ingredient ruins the dish, no matter how beautiful the plating.

The Finding That Should Worry Every Buyer

Here’s the result that gives the experiment its bite: all five models spotted every crisis and refused every manipulation attempt. Every one of them diagnosed the €55k opportunity correctly. But only two — gpt-5.6-sol and Kimi K3 — actually signed the deal their own analysis had earned. Firmulate’s summary of the gap: “Same diagnosis, same pitch — no signature.”

The difference was homework. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer conversation, not in the crisis, but in the pantry. The models that actually read the file closed the deal at full price, worth an additional €4,583 in monthly recurring revenue. The ones that didn’t read? They served the same dish and wondered why nobody sent it back with a compliment.

The Traps: A Tasting Menu of Manipulation

The week included social engineering on the menu: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning is worth quoting: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the whole gauntlet, K3 deviated from clean process just once — the cleanest discipline in the field.

The Cautionary Tale: The Most Thorough Chef Came Last

Then there’s Opus 4.8, the season’s most poignant contestant. It was the most thorough participant in the field — the deepest analyses, more than 80 learned rules added to its playbook — and it still finished last at 73. It left the close on the table, and its discipline slipped: it attempted writes into a locked department rather than escalating, like a chef who preps for six hours and then plates the wrong table’s order. Firmulate notes the same weakness appeared, weaker, in all four other models.

For a food audience, the lesson translates directly: mise en place is not the meal. Prep, depth, and beautiful analysis don’t feed anyone until the plate actually goes out and gets accepted.

One Caveat, Plainly Stated

A fairness footnote the league publishes openly: Kimi K3 ran without an effort parameter (API default), while the other models ran at xhigh. In tasting terms, the newcomer may not even have been using its highest heat. The result stands, but the comparison isn’t perfectly controlled.

It’s All Watchable, Live

None of this is a slide deck. The company is real software with real money mechanics — 13 synthetic employees, a burn of €105k a month against just €2.3k in MRR, a public cash countdown, and a self-learned playbook that has grown past 680 rules. Every workday is versioned, and you can watch it live. There’s even a guessing game built from 242 real, unedited management decisions: which model made which call? Full results and plain-language findings are on the benchmarks page, and the whole thing runs at firmulate.com.

Enterprises can go further: run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Takeaway

If AI agents will touch your CRM, your support queue, or your forecast, the question is no longer “does it write well?” It’s: does it finish what it starts, does it read your files before it acts, and does it stay honest when nobody’s watching? The Crucible showed those gaps are invisible in chat demos — and that the league is genuinely open. A newcomer can walk into the kitchen and outperform established names on the exact same ticket. Picking a model without running your own test isn’t a decision anymore. It’s a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Management Taste Test: Which Model Has the Best Instincts?

Five frontier AIs faced the same corporate crises. Their choices reveal distinct management personalities—and whether they actually finish the job.

Summer Crispy Chicken Wings with Ninja XL Air Fryer

Learn how to make perfectly crispy chicken wings this summer using the Ninja XL Air Fryer. Quick, easy, and healthier crispy wings every time!

Pecan Cranberry Cheese Ball: The Ultimate Party Appetizer

Liven up your party with the irresistible Pecan Cranberry Cheese Ball, a creamy, sweet, and crunchy appetizer that will leave your guests wanting more.

Whipped Ricotta With Honey and Cracked Pepper: a Gourmet Appetizer!

Nourish your taste buds with a luxurious blend of whipped ricotta, honey, and cracked pepper – a gourmet appetizer that promises a flavor explosion!