
Get kitchen gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Would you let an AI run your kitchen shop through its worst week?
If you sell cookware or rely on a busy kitchen, a bad week can mean late orders, unhappy customers and a supplier problem arriving at the same time. An AI assistant may sound confident when asked what to do. But would it follow your playbook, spot a hidden opportunity and actually close a deal under pressure? Firmulate has built a live experiment to test that kind of management in a small software company.
Same company, same crises, different models
In the final Crucible League, in July 2026, frontier AI models ran the same small software company through its worst week. They faced the same customers, crises and temptations. Their decisions were versioned and auditable. The experiment is designed to show what happens when models have to manage a company, not just produce a polished answer.
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding was blunt: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not guarantee that a model would carry it through.
The clue was buried in the company’s own files
The decisive competitor weakness was two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That is a useful test for any business considering AI: the relevant information may already exist, but finding and applying it under pressure is another matter.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
More thorough did not mean more effective
The final league ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust”.
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and discipline slipped: it attempted to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four. One fairness detail: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The live experiment is watchable at firmulate.com. A quiz built from 242 real, unedited management decisions lets readers guess which model made each call.
From watching to trying it on your business
The enterprise pilot is the next step: run crisis scenarios against a digital twin made from a read-only export of your own business. The result is a board report with model rankings and the weak points in your playbooks. Nothing writes back to real systems. The idea is to see how an AI workforce handles your company’s customers, rules and pressure before trusting it with live operations.

Put your playbooks to the test
Firmulate’s pilot lets enterprises wargame their business using a read-only export, examine model performance and surface weak points without writing back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
