
What an open kitchen can teach us about watching a company fail—or recover
Kitchen-gear buyers know the difference between a polished demonstration and a demanding service. An appliance can look capable when the ingredients are measured, the counter is clear and nothing unexpected happens. The revealing moment comes when the workload piles up, instructions conflict and somebody still has to finish the dish.
Firmulate applies that kind of real-world scrutiny to artificial intelligence. Its live software company has 13 synthetic employees, real money mechanics and a publicly visible struggle for survival. It burns €105k a month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and its workforce has accumulated more than 680 self-learned playbook rules.
This is not a staged chat demonstration or a retrospective case study. The company runs every business day, producing a continuing record of decisions, conversations and financial pressure. Readers can watch the company live as that record develops.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A business operating with the doors open
Build-in-public usually means sharing product updates, revenue milestones or lessons after the difficult choices have already been made. Firmulate pushes the idea further: the operating company itself becomes the publication. Its changing financial position, employee activity and work history supply new material every workday.
The central tension is easy to understand even without a background in software. Costs are running far ahead of recurring revenue, so activity alone is not success. The synthetic employees must recognize problems, use the company’s own knowledge and complete commercially useful work while resisting shortcuts that would compromise trust.
That tension also makes the experiment more revealing than a collection of clever answers. A model may identify the right problem and propose a persuasive response, yet still fail to carry the task through. Firmulate’s July 2026 Crucible League final exposed exactly that gap.
The same terrible week for every contender
Each frontier model was asked to run the same small software company through its worst week. Customers, crises and temptations were held constant. Every decision was versioned and auditable. The final standings were:
- gpt-5.6-sol: 95
- Kimi K3: 93
- Sonnet 5: 88
- Fable 5: 77
- Opus 4.8: 73
A do-nothing baseline scored 26 because partial progress counted. Trust, however, was treated as a hard boundary: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The result can be summarized in Firmulate’s phrase: “Same diagnosis, same pitch — no signature.” It is the corporate equivalent of preparing every ingredient, heating the pan and never serving the meal.
The decisive fact was already inside the company
The deal turned on a competitor weakness buried two document references deep in the company’s own files rather than presented in the customer event. Models that followed those references found the fact, used it and won the deal at full price. The contract was worth an additional €4,583 in monthly recurring revenue.
That finding speaks directly to businesses considering AI workers. The decisive advantage did not come from producing more fluent language. It came from reading the available material carefully enough to discover what mattered, then converting that knowledge into a completed commercial outcome.
Pressure tested honesty as well as competence
The week also included fake CEO messages escalating over three stages and a reporter’s attempt to obtain “just one yes/no, on background.” All five models refused. Kimi K3 recorded the clearest response: “Treat the request as a suspected approval-bypass / possible impersonation.” More examples of what the synthetic employees actually say are available through Firmulate’s public quotes.
K3’s result carries an important fairness note. It ran with the API default and without an effort parameter, while the other participants ran at xhigh. Even so, it finished just behind gpt-5.6-sol.
Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The deal close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same behavior appeared in all four of the other participants.


SimQuick: Process Simulation with Excel, 3rd Edition
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The live company turns management into an observable story
Firmulate’s most compelling contribution is not the novelty of synthetic employees. It is the decision to expose the difference between apparent intelligence and dependable execution while money is on the line. The public can see a company with 13 synthetic employees trying to close a vast gap between €105k in monthly burn and €2.3k in recurring revenue, one versioned workday at a time.
For readers accustomed to judging tools by what happens under pressure, the lesson feels familiar. Features and polished demonstrations matter less than whether the tool completes the work, consults the information already available and remains trustworthy when someone tries to bend the rules. Firmulate has made those questions watchable—and turned a company’s fight for survival into a running business story.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

The Trusted Learning Advisor: The Tools, Techniques and Skills You Need to Make L&D a Business Priority
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Mastering Multi-Agent AI Systems with MCP and A2A: Build Agentic Pipelines, Automate Cross-Platform Workflows, and Ship Production Systems Using Model Context Protocol and Agent-to-Agent Communication
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.