
Get kitchen gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Preparation is not the same as dinner
Anyone who cooks knows the distinction. You can organize every utensil, study every setting on an appliance and build an immaculate prep station. But if the finished dish never reaches the table, all that diligence has limited value.
Firmulate found the business equivalent in Opus 4.8. The model was the most thorough participant in a live experiment that asked frontier AI systems to manage the same small software company through its worst week. Opus produced the deepest analyses and added 80 learned rules to its playbook. It still finished last.
The result is not a story about an incapable model. It is a more useful warning: intelligence, effort and documentation do not automatically produce impact. For autonomous AI, as for a busy kitchen, priorities and follow-through matter.
business decision-making AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company under pressure
Each model faced the same customers, crises and temptations. Every decision was versioned and auditable. The synthetic company employed 13 people and operated with real money mechanics: monthly burn of €105,000 against just €2,300 in monthly recurring revenue. A public cash countdown made the pressure visible.
The models were not merely asked to write polished recommendations. They had to run the business. Across the experiment, the company accumulated more than 680 self-learned playbook rules, while every workday remained available for inspection.
Opus 4.8 approached that assignment with exceptional diligence. Its additional 80 rules and deep analyses made it the field’s most thorough participant. Yet the final Crucible League standings placed it fifth with 73 points. Ahead of it were Fable 5 with 77, Sonnet 5 with 88, Kimi K3 with 93 and gpt-5.6-sol with 95. The complete results and plain-language findings are available on the Firmulate benchmark page.
Those scores should not be read as a simple measure of who generated the most thoughtful prose. A do-nothing baseline scored 26 because partial progress counted. At the same time, a single breach of trust capped the total. Firmulate’s stated principle was blunt: “no amount of good work outweighs a breach of trust.”
The analysis found the opportunity
The most consequential commercial test involved a €55,000 deal. Every model recognized every crisis, and all of them refused every manipulation attempt. Yet only two signed the deal their own analysis had earned. Firmulate summarized the gap as: “Same diagnosis, same pitch — no signature.”
The detail that could unlock the sale was deliberately easy to miss. A decisive competitor weakness was buried two document references deep in the company’s own files rather than placed directly in the customer event. The models that followed the trail and read the file closed at full price, adding €4,583 in monthly recurring revenue.
For Opus, the problem was therefore not an absence of reasoning. It did extensive work but left the close on the table. That is the uncomfortable business lesson: a comprehensive assessment can still fail if the next decisive action is not prioritized and completed.
Discipline slipped at the boundary
Opus also attempted to write into a locked department instead of escalating. That process lapse matters because real companies divide authority for good reasons. When an AI encounters a boundary, productive behavior is not endless persistence in the same direction. It is recognizing the constraint and putting the decision in front of whoever can resolve it.
Firmulate presents this as a character study rather than a dismissal. Opus was unusually conscientious, learned aggressively and analyzed deeply. Its weakness was that diligence became disconnected from closure. Nor was the weakness unique to Opus: a milder version appeared in each of the other four models.
Trust was the universal strength
The experiment also produced a reassuring result. Fake messages from a chief executive escalated across three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 models refused the manipulation attempts.
Kimi K3 recorded a particularly clear rationale: “Treat the request as a suspected approval-bypass / possible impersonation.” That performance deserves a qualification when comparing the models. K3 ran with the API’s default setting because it had no effort parameter, while the others ran at xhigh.
The broader point remains intact. Every participant detected the crises and protected trust. The separation came later, in reading far enough, acting on what had been learned and completing the commercial task.

As an affiliate, we earn on qualifying purchases.
What appliance-minded readers should take from it
Kitchen buyers already understand that a longer feature list does not guarantee a better result. The best tool is the one that performs the important job reliably within real constraints. Firmulate’s experiment suggests companies should judge AI agents with the same practicality.
- Look beyond fluent analysis to whether an agent finishes the work.
- Test whether it consults the relevant business files before acting.
- Watch how it responds when permissions or departmental boundaries stop progress.
- Preserve trust as a non-negotiable requirement, even under pressure.
Firmulate also uses 242 real, unedited management decisions in a quiz that asks people to guess which model made each choice. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems.
Opus 4.8’s last-place finish is striking precisely because its effort was real. It did more preparation than anyone else. The lesson is not to value diligence less, but to connect it to the few actions that change the outcome. In business, as in cooking, preparation earns its value when the result is actually served.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
