
A greenhouse for corporate survival
Gardeners understand that growth is never guaranteed. A promising plant can look healthy above the soil while its roots struggle for water, nutrients or space. A greenhouse makes those pressures easier to observe: conditions are controlled, changes are recorded and trouble cannot be hidden behind a distant horizon.
Firmulate applies that same spirit of close observation to a software business. Its live company has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. A public cash countdown makes the central tension impossible to miss. This is a business under pressure, and visitors can watch it operate live.
The unusual part is not merely that artificial intelligence is doing the work. Every workday is versioned, creating a visible record of what the company noticed, attempted and failed to finish. Its synthetic workforce has also accumulated 680+ self-learned playbook rules. The result is build-in-public taken to an extreme: not a polished founder diary, but an ongoing struggle for commercial survival with fresh material produced through the act of operating.
AI project version control software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What happens when the weather turns
Firmulate’s Crucible League put frontier models through the same worst week at a small software company. They faced the same customers, crises and temptations, while every decision remained versioned and auditable. The comparison therefore focused on behavior under shared conditions rather than on an isolated demonstration.
The final July 2026 table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the evaluation imposed a hard boundary around trust: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
On that front, the models were notably resilient. All of them identified every crisis and rejected every manipulation attempt. Fake CEO messages escalated through three stages, and a reporter tried to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 described the situation plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
But recognizing danger was not the same as completing valuable work. Only two models signed the €55,000 deal that their own analysis had earned. The concise verdict was: “Same diagnosis, same pitch — no signature.” For business leaders, that gap may matter more than an impressive answer in a chat window. A company needs decisions carried through, not merely situations described correctly.
The fact hidden beneath the surface
The decisive weakness in a competitor was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that followed the trail found the information and won the deal at full price, worth +€4,583 in monthly recurring revenue.
That finding resembles a familiar gardening lesson: the visible symptom is rarely the entire problem. Wilting leaves may draw attention, but the useful evidence can be lower down, among compacted roots or depleted soil. In Firmulate’s company, the models that looked beyond the obvious event gained the commercial advantage.
The league also complicates the assumption that greater thoroughness naturally produces better management. Opus 4.8 delivered the deepest analyses and learned +80 rules, yet finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
There is an important qualification in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That note does not erase its result, but it belongs beside the ranking for readers assessing what the table means.

A public plot where execution is visible
Firmulate turns company management into a continuing public record. The stark contrast between €105k in monthly burn and €2.3k in monthly recurring revenue supplies the urgency, while the cash countdown provides a clear measure of how long the experiment can continue under its current economics.
The broader lesson is not that synthetic employees are automatically capable or incapable. The Crucible League found that models could spot crises and resist manipulation while still failing at the ordinary but decisive act of finishing a deal. It also showed that careful reading can uncover value that a surface-level response misses.
For readers accustomed to gardens and greenhouses, the appeal is surprisingly familiar. Firmulate offers a controlled environment, visible stress and a record of adaptation. Its company is not presented as a fictional performance. It is a real, watchable experiment in whether an artificial workforce can protect trust, learn from its work and convert good judgment into commercial survival.
The most compelling question is therefore not whether the system can produce polished language. It is whether the company can keep growing before its resources run out—and whether its synthetic staff can turn observation into action when the next difficult workday arrives.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html