
Every gardener knows a plot left unattended for a week doesn’t revert to bare soil. The perennials hold. The mulch keeps working. Some things fruit on their own. A do-nothing week in the garden isn’t a zero — it’s a partial score, and any honest grower will tell you so.
Get garden gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
It turns out an AI benchmark has adopted the same kind of honesty. Firmulate’s Crucible League — a live experiment that runs frontier AI models as managers of the same small software company through its worst week — gives its do-nothing baseline 26 points out of 100, not zero. That number is quietly one of the most interesting things about the whole project, and it says a lot about what honest measurement looks like when the thing being measured is management, not chat.
The experiment: same company, same worst week
Firmulate handed four frontier AI models the identical job: run a small software company through seven days of crisis — same customers, same emergencies, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the grading depends on vibes.
The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the more revealing entry is the one nobody ran: the baseline. If a model simply sat on its hands all week, it would score 26.
AI management benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why a do-nothing manager gets 26, not 0
The logic is the gardener’s logic. A company — like a garden — has momentum. Customers keep paying. Systems keep running. Doing nothing still preserves some of that value, and some problems genuinely resolve themselves. A scoring system that gave total passivity a zero would be lying about how businesses (and gardens) actually work.
But partial progress counts in both directions. A manager who diagnoses a problem but never finishes the fix gets credit for the diagnosis and dinged for the dangling close. That’s precisely what the experiment’s key finding showed: every model spotted every crisis and refused every manipulation attempt — yet only two of the five signed the €55,000 deal their own analysis had earned. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.” Chat-quality demos make that gap invisible. Management-quality scoring makes it the whole story.
The buried fact
The detail that separated the winners: the decisive competitor weakness wasn’t in the customer’s conversations at all. It sat two document references deep in the company’s own files. The models that actually read what was already on the shelf — the way a good gardener checks the soil before watering — won the deal at full price, worth +€4,583 in monthly recurring revenue. The others left the harvest standing in the field.
Why one breach of trust caps everything
The scoring has a second principle that any growers’ market vendor would recognize: “no amount of good work outweighs a breach of trust.” A single act that breaks faith with a customer or colleague caps the total grade, no matter how brilliant the rest of the week was. In the garden analogy: it doesn’t matter how lush your tomatoes are if you sold someone else’s produce as your own.
Notably, no model breached trust. The social-engineering gauntlet — fake CEO messages escalating over three stages, plus a reporter’s disarming “just one yes/no, on background” trick — was refused by five out of five models. Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”
Effort isn’t everything
The most sobering profile belongs to Opus 4.8: the most thorough participant in the field, with more than 80 self-learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more faintly, in all four models. Like the gardener who amends every bed but never actually harvests, effort without completion doesn’t score.
One fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still finished second with the cleanest discipline of the field.
A live company, watchable like a greenhouse cam
This isn’t a one-off report. Firmulate runs a live company with 13 synthetic employees and real money mechanics: a burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it the way you’d watch a greenhouse feed — and the site rebuilds itself twice a day. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions.

The 26-point floor is a small design choice with a big lesson: honest measurement doesn’t hand out zeros or round, flattering hundreds. It credits the momentum that exists, rewards finishing what you start, and treats trust as non-negotiable. Gardeners have graded this way forever — a neglected plot still grows something, a finished harvest beats perfect intentions, and one act of bad faith at the market undoes a season of good work. If AI agents are going to touch your business, that’s the kind of scorecard worth demanding. Enterprises can even run the same wargame against a read-only export of their own operations — nothing ever writes back to real systems. The full results, in plain language, are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
