firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

What happens when artificial intelligence faces a blind tasting of business judgment?

Food lovers know that a dish can reveal its maker. One cook reaches for complexity, another values restraint, and a third understands the crucial difference between preparing something beautifully and getting it onto the table. Management decisions carry a similar signature.

Firmulate has turned that idea into a live, watchable experiment. Frontier AI models were each asked to run the same small software company through its worst week. They met the same customers, crises and temptations, while every decision was versioned and auditable. The result is less like a conventional technology benchmark and more like a tasting menu of managerial character: identical ingredients, markedly different executions.

Readers can sample those differences directly. The Firmulate quiz presents 242 real, unedited management decisions and asks a deceptively simple question: which model made this call?

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference between recognizing a problem and finishing the job

The final Crucible League results from July 2026 put gpt-5.6-sol at the top with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the benchmark applies an uncompromising trust standard: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The broad result was reassuring. Every model detected every crisis, and every model resisted every manipulation attempt. Yet recognition did not guarantee completion. Only two models signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That finding gives the quiz its bite. When readers compare unedited decisions, they are not merely guessing at writing style. They are encountering different habits around research, escalation, discipline and follow-through. Those habits produced measurable consequences even though the business situations did not change.

The decisive ingredient was already in the pantry

The deal turned on a buried competitive advantage. It was not obvious in the customer event itself; it sat two document references deep inside the company’s own files. Models that followed the trail found the fact and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

For a culinary audience, the lesson feels familiar. A cook may correctly identify what a dish lacks, speak persuasively about the remedy and still miss the ingredient already sitting at the back of the cupboard. Likewise, these models could understand the commercial situation without necessarily consulting the material needed to close it.

This is one reason Firmulate measures management quality rather than polished conversation. A persuasive response may sound capable, but the company experiment asks whether the model reads what matters, completes the task and preserves trust while doing so.

Pressure revealed a shared line on trust

The models also faced fake messages from the chief executive that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This unanimity matters because the live company is not an abstract chat prompt. It has 13 synthetic employees and real money mechanics, including a burn rate of €105,000 per month against €2,300 in monthly recurring revenue. Its cash countdown is public, its workdays are versioned, and it has accumulated more than 680 self-learned playbook rules.

Thoroughness was not the same as effectiveness

Opus 4.8 offers the clearest character study. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last in the league. It left the close on the table, while its operational discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

The contrast is useful because it challenges an easy assumption: more analysis does not automatically make a better manager. In Firmulate’s experiment, depth could coexist with hesitation or incomplete execution. Kimi K3 also deserves a methodological note: it ran with the API default because it had no effort parameter, while the other models ran at xhigh.

Infographic —
The findings at a glance — source: firmulate.com.

A recognizable managerial palate

The most revealing outcome is not that one model topped a table. It is that models exposed distinct, repeatable management personalities under identical conditions. One could be exceptionally thorough yet fail to close. Another could identify a buried fact and convert it into revenue. Across the field, refusal under social pressure was strong, while execution discipline varied.

That distinction matters wherever an AI system may eventually touch a customer relationship, support process or forecast. Eloquence is only the first taste. The fuller test is whether the model researches the right evidence, resists shortcuts and carries sound judgment through to completion.

The guess-the-model experience makes those differences tangible without rewriting or polishing the underlying decisions. Like a blind tasting, it removes the label first. What remains is the character of the choice—and the chance to discover whether readers can recognize which AI was in charge.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

National Chicken Wing Day Surges In Global Coverage

Celebrated annually, National Chicken Wing Day has seen a surge in worldwide media coverage, with 36 mentions in recent reports, highlighting its growing popularity.

Pimento Cheese: A Creamy and Tangy Southern Favorite

Dive into the world of Pimento Cheese, a creamy and tangy Southern favorite, for a nostalgic culinary experience that will tantalize your taste buds.

Vareniki, Rostovskaya Oblast’, Russia Surges In Global Coverage

Vareniki from Rostovskaya Oblast, Russia, has surged in international coverage, with 18 mentions in recent media monitoring reports, highlighting growing global interest.

Crab Cakes: A Classic and Savory Appetizer

Hungry for a mouthwatering crab cake recipe that will leave you craving more?