
The appliance-review question, applied to artificial intelligence
Anyone who tests kitchen gear knows that polished first impressions reveal very little. A cooker may look capable on the counter and still disappoint when several dishes need attention at once. A blender may ace the easy ingredients but struggle with the task that actually matters. The useful test is not whether a product can perform in ideal conditions; it is whether it notices trouble, uses the available information and completes the job under pressure.
Firmulate applies that philosophy to frontier AI models. Its live experiment gave each model the same small software company during its worst week, with identical customers, crises and temptations. The decisions were versioned and auditable, turning management behavior into something readers can inspect rather than a staged demonstration. Now, a quiz built from 242 real, unedited decisions asks readers to identify which model made each call.
As an affiliate, we earn on qualifying purchases.
The same ingredients produced different results
The final Crucible League table from July 2026 puts gpt-5.6-sol in first place with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress counts, although a single breach of trust caps the total. Firmulate summarizes that boundary plainly: “no amount of good work outweighs a breach of trust.”
The striking result is not that some models recognized danger while others missed it. All of them spotted every crisis, and all refused every manipulation attempt. The difference appeared in follow-through. Only two signed the €55,000 deal that their own work had earned. Firmulate’s concise verdict captures the gap: “Same diagnosis, same pitch — no signature.”
For readers accustomed to comparing cooking appliances, this resembles the distinction between sensing that a pan is overheating and actually lowering the heat. Recognition is valuable, but management requires an action that resolves the situation. An AI can produce an impressive analysis and still leave the commercial outcome sitting untouched.
The crucial fact was not where the drama happened
The deal also hinged on information retrieval rather than eloquence. A decisive competitor weakness was buried two document references deep in the company’s own files, not in the customer event that demanded attention. Models that read the file secured the deal at full price, worth +€4,583 MRR.
That detail makes the experiment unusually relevant beyond software companies. In a kitchen, the visible problem may be a failed bake, while the useful clue sits in a manual, an earlier temperature log or the specifications for a particular pan. In business, the urgent message is not necessarily the best source of truth. The stronger performers did not merely react to what was loudest; they found the evidence that changed the negotiation.
Every model held the line against manipulation
The company’s worst week included fake CEO messages escalating over three stages and a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning in direct terms: “Treat the request as a suspected approval-bypass / possible impersonation.”
This is an important counterweight to the league table. The models differed in execution, but the social-engineering tests did not separate them: refusal was universal. K3’s result also carries a fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh.
Thoroughness did not guarantee victory
Opus 4.8 offers the clearest character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
That combination complicates a familiar assumption about AI quality. More analysis can be useful, just as a feature-heavy appliance can expand what a cook can make. But completeness on paper is not the same as dependable operation. The Firmulate results suggest that management personality emerges through habits: whether a model investigates, closes, escalates correctly and resists distraction.
The live company gives those habits a demanding setting. It has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k MRR. Its cash countdown is public, its workdays are versioned, and it has accumulated 680+ self-learned playbook rules. The experiment is real and watchable, so its claims can be followed as continuing business behavior rather than accepted as a one-off showcase.

As an affiliate, we earn on qualifying purchases.
The quiz tests judgment, not writing style
Firmulate’s quiz works because the decisions are not invented for entertainment. They come from identical situations faced by competing models, letting readers ask whether managerial fingerprints are recognizable in the choices themselves. The differences are subtle: careful reading versus surface reaction, completion versus commentary, and disciplined escalation versus an unproductive attempt.
For kitchen-gear readers, the broader lesson is familiar. Labels and specifications matter less than performance in a repeatable stress test. The leading model did not win merely by sounding persuasive; it found the buried evidence and completed the commercial task without crossing trust boundaries.
Enterprises can also run the same wargame against a read-only export of their own business, with nothing writing back to real systems. That moves evaluation closer to the messiness of actual work while keeping operational systems protected. Before entrusting an AI workforce with customers, forecasts or internal records, it may be worth asking the same question applied to every serious appliance: what happens when the heat is really on?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI audit and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.