firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

The appliance-review question, applied to artificial intelligence

Anyone who tests kitchen gear knows that polished first impressions reveal very little. A cooker may look capable on the counter and still disappoint when several dishes need attention at once. A blender may ace the easy ingredients but struggle with the task that actually matters. The useful test is not whether a product can perform in ideal conditions; it is whether it notices trouble, uses the available information and completes the job under pressure.

Firmulate applies that philosophy to frontier AI models. Its live experiment gave each model the same small software company during its worst week, with identical customers, crises and temptations. The decisions were versioned and auditable, turning management behavior into something readers can inspect rather than a staged demonstration. Now, a quiz built from 242 real, unedited decisions asks readers to identify which model made each call.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same ingredients produced different results

The final Crucible League table from July 2026 puts gpt-5.6-sol in first place with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress counts, although a single breach of trust caps the total. Firmulate summarizes that boundary plainly: “no amount of good work outweighs a breach of trust.”

The striking result is not that some models recognized danger while others missed it. All of them spotted every crisis, and all refused every manipulation attempt. The difference appeared in follow-through. Only two signed the €55,000 deal that their own work had earned. Firmulate’s concise verdict captures the gap: “Same diagnosis, same pitch — no signature.”

For readers accustomed to comparing cooking appliances, this resembles the distinction between sensing that a pan is overheating and actually lowering the heat. Recognition is valuable, but management requires an action that resolves the situation. An AI can produce an impressive analysis and still leave the commercial outcome sitting untouched.

The crucial fact was not where the drama happened

The deal also hinged on information retrieval rather than eloquence. A decisive competitor weakness was buried two document references deep in the company’s own files, not in the customer event that demanded attention. Models that read the file secured the deal at full price, worth +€4,583 MRR.

That detail makes the experiment unusually relevant beyond software companies. In a kitchen, the visible problem may be a failed bake, while the useful clue sits in a manual, an earlier temperature log or the specifications for a particular pan. In business, the urgent message is not necessarily the best source of truth. The stronger performers did not merely react to what was loudest; they found the evidence that changed the negotiation.

Every model held the line against manipulation

The company’s worst week included fake CEO messages escalating over three stages and a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning in direct terms: “Treat the request as a suspected approval-bypass / possible impersonation.”

This is an important counterweight to the league table. The models differed in execution, but the social-engineering tests did not separate them: refusal was universal. K3’s result also carries a fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh.

Thoroughness did not guarantee victory

Opus 4.8 offers the clearest character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

That combination complicates a familiar assumption about AI quality. More analysis can be useful, just as a feature-heavy appliance can expand what a cook can make. But completeness on paper is not the same as dependable operation. The Firmulate results suggest that management personality emerges through habits: whether a model investigates, closes, escalates correctly and resists distraction.

The live company gives those habits a demanding setting. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k MRR. Its cash countdown is public, its workdays are versioned, and it has accumulated 680+ self-learned playbook rules. The experiment is real and watchable, so its claims can be followed as continuing business behavior rather than accepted as a one-off showcase.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI model testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The quiz tests judgment, not writing style

Firmulate’s quiz works because the decisions are not invented for entertainment. They come from identical situations faced by competing models, letting readers ask whether managerial fingerprints are recognizable in the choices themselves. The differences are subtle: careful reading versus surface reaction, completion versus commentary, and disciplined escalation versus an unproductive attempt.

For kitchen-gear readers, the broader lesson is familiar. Labels and specifications matter less than performance in a repeatable stress test. The leading model did not win merely by sounding persuasive; it found the buried evidence and completed the commercial task without crossing trust boundaries.

Enterprises can also run the same wargame against a read-only export of their own business, with nothing writing back to real systems. That moves evaluation closer to the messiness of actual work while keeping operational systems protected. Before entrusting an AI workforce with customers, forecasts or internal records, it may be worth asking the same question applied to every serious appliance: what happens when the heat is really on?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI audit and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis simulation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

COSORI Dual Blaze Review: Powerful Cooking in a Compact Design

Explore the pros, cons, and ideal users of the COSORI Dual Blaze air fryer in this detailed review. Find out if it’s the right choice for your kitchen needs.

Best COSORI Air Fryers for Frozen Foods: Top Picks

Discover the best COSORI air fryers for frozen foods with our expert roundup. Find the perfect model for crispy, evenly cooked frozen favorites.

Master Summer Snacks with the Ninja Crispi 4-in-1 Glass Air Fryer

Learn how to make quick, crispy summer snacks using the Ninja Crispi 4-in-1 Glass Air Fryer with this easy step-by-step recipe.

Ninja Foodi XL Pro Air Oven: The Ultimate Summer Kitchen Companion

Discover how the Ninja Foodi XL Pro Air Oven elevates summer cooking with fast, even, and versatile countertop baking, roasting, and air frying.