
A stress test for artificial managers
Gardeners know the difference between a promising seedling and a plant that survives a punishing week. Healthy growth in controlled conditions says little about how it will handle drought, pests or a sudden temperature swing. Business software faces a similar test: an AI can sound capable in conversation, yet still fail when judgment must become action.
Firmulate put that distinction on public display by asking frontier AI models to run the same small software company through its worst week. Each encountered the same customers, crises and temptations. Every decision was versioned and auditable. The striking result was not that some models failed to understand the problems. All of them spotted every crisis and resisted every manipulation attempt. The separation came at the finish line: only two signed the €55,000 deal their own work had earned.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same diagnosis, different outcome
The experiment challenges a comfortable assumption about AI capability. Chat demonstrations reward fluent answers, quick analysis and persuasive writing. Managing a company also requires follow-through: reading the relevant material, finding decisive evidence, maintaining trust and completing an approved action.
In Firmulate’s words, the result was: “Same diagnosis, same pitch — no signature.” The models could identify the opportunity and prepare the case, but most did not convert that preparation into the final commercial act.
The decisive information was not sitting prominently in the customer event. A competitor weakness was buried two document references deep in the company’s own files. The models that read the file won the deal at full price, worth +€4,583 MRR. That makes the episode less a test of clever improvisation than of ordinary managerial diligence: consult the company record before acting.
The final league
The July 2026 Crucible League results placed the participants in this order:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
A do-nothing baseline scored 26 because partial progress still counted. But Firmulate imposed a firm limit on misconduct: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” That matters because the week included deliberate attempts to push the models past normal approval boundaries.
Pressure without capitulation
The social-engineering sequence used fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest concise interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”
This was a meaningful success across the field. The models did not trade trust for convenience, even while other demands competed for attention. Yet universal resistance also sharpened the larger finding. Safety alone did not determine the winner because every participant cleared that hurdle. The decisive difference was whether sound analysis became completed work.
Why thoroughness was not enough
Opus 4.8 offers the clearest warning against equating visible effort with business effectiveness. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close remained on the table, while discipline slipped through write attempts into a locked department instead of escalation. A weaker version of the same problem appeared in all four.
That profile will feel familiar to managers who have watched a project accumulate research, process notes and recommendations without reaching its commercial objective. Thoroughness can support execution, but it does not substitute for it. The experiment exposed that distinction because the company kept moving after the model had produced an impressive answer.
There is also an important fairness qualification. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The league should therefore be read as a documented comparison under those stated conditions, rather than as an abstract declaration about every possible configuration.
A company built to reveal behavior
The live business has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, maintains a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing visitors to watch the experiment rather than relying on a polished retrospective.
Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz. The premise asks readers to judge behavior without leaning on a model’s name or reputation. For enterprises, the pilot applies the wargame to a read-only export of their own business; nothing writes back to real systems.

The capability chat windows conceal
The lesson for companies considering AI agents is not that eloquence has no value. It is that eloquence reveals only part of the job. An agent touching a CRM, support queue or forecast must finish what it starts, read company files before acting and remain honest when pressure rises.
Like a greenhouse trial that continues when conditions turn hostile, Firmulate’s live experiment makes resilience and follow-through observable. Every model could see the trouble. Every model could reject the tricks. Only two completed the €55,000 close. That final gap—the distance between knowing and doing—is precisely what an ordinary chat demo leaves hidden.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html