firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The instruction manual can decide the outcome

Anyone who buys kitchen gear knows that the decisive detail is not always printed on the box. It may sit in a manual, a compatibility note or a document referenced by another document. An appliance can look capable in a demonstration yet disappoint because someone never followed that trail.

Firmulate has measured the business equivalent. Its live experiment placed frontier AI models in charge of the same small software company during its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable. The most revealing test was not whether an AI could produce an impressive answer. It was whether it would examine the company’s own files deeply enough to act on what it found.

A €55,000 deal depended on precisely that behavior. The crucial weakness in a competitor was absent from the customer event. It was buried two document references deep in the company’s files. Models that reached it could support the full-price offer, worth +€4,583 in monthly recurring revenue. Models that stopped early lost the deal automatically.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone saw the crisis, but only two completed the sale

The result isolates a subtle weakness in AI agents. All models spotted every crisis and rejected every manipulation attempt. Their diagnosis of the sales opportunity was also strong enough to produce the right pitch. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”

That distinction matters because chat demonstrations tend to reward articulate responses. In business, a polished explanation is only an intermediate product. Useful work may require checking supporting material, connecting separated facts and completing the authorized action. Here, reading the file was not a bonus for thoroughness. It determined whether revenue arrived.

The final July 2026 Crucible League makes the performance differences visible. The published benchmark ranks gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress counts, although a breach of trust caps the result. The governing principle is explicit: “no amount of good work outweighs a breach of trust.”

Thoroughness did not guarantee completion

Opus 4.8 provides the clearest warning against confusing visible effort with operational success. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close remained on the table, while discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

That profile will feel familiar to anyone comparing elaborate cooking appliances. More settings, more displays and a thicker manual do not necessarily produce a better meal. The relevant question is whether the machine performs the complete job reliably. For AI agents, extensive reasoning can coexist with a failure to retrieve the decisive fact, respect an operating boundary or finish the task.

The agents resisted pressure more consistently than they closed

The experiment also tested whether models would abandon proper process when prompted to do so. Fake CEO messages escalated over three stages, and a reporter tried the line “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning as: “Treat the request as a suspected approval-bypass / possible impersonation.”

This clean result makes the sales failure more informative. The models were not oblivious to danger, and none accepted the manipulation attempts. The differentiator was ordinary diligence: following references through internal material and then carrying a sound decision through to completion.

K3’s strong result also carries an important fairness note. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. That difference should remain visible when readers compare placements.

A company designed to make behavior observable

The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. The setup turns vague claims about capable agents into behavior that can be watched and audited.

The surrounding evidence is unusually accessible. A quiz uses 242 real, unedited management decisions and asks visitors to guess the model behind each one. Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems, allowing evaluation against company-specific material without granting the experiment operational control.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI data retrieval tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

File-reading belongs on the buying checklist

For buyers of AI agents, the buried fact changes the evaluation question. Fluency, crisis recognition and resistance to manipulation are important, but they do not establish that an agent will complete valuable work. A serious trial should reveal whether it opens relevant files, follows references, notices commercially decisive details and finishes the action its analysis supports.

That is as practical as checking whether a countertop appliance fits the cookware and workflow it will actually meet. The best-looking demonstration can conceal a costly gap. In Firmulate’s test, that gap had a clear price: a €55,000 signature and +€4,583 MRR. Reading before answering was not a personality trait. It was a measurable business capability.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI-powered knowledge management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI for business decision support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Italian Restaurant And Delicatessen Bocce Ristorante Opens In East Sacramento

Italian restaurant and delicatessen Bocce Ristorante has opened in East Sacramento, offering traditional Italian cuisine and specialty products to the community.

Meal Prepping Protein in the Air Fryer

Discover how to effortlessly meal prep protein in the air fryer for quick, tender, and flavorful meals that will keep you coming back for more.

Creative Vegetable Dishes: Roasting Chickpeas and More

Creative vegetable dishes like roasted chickpeas and vibrant veggies inspire flavorful, eye-catching meals—discover how to elevate your cooking today.

Air Fryer Plant-Based Proteins: Tofu and Tempeh

Balanced air frying of tofu and tempeh enhances flavor and texture, but discover how to unlock even more delicious plant-based protein secrets.