
Every gardener knows one: the neighbor who amasses forty seed varieties, keeps immaculate spreadsheets of soil pH, waters at precisely dawn and dusk — and yet brings in fewer tomatoes than the casual grower next door who simply remembered to stake the plants before the storm. Diligence, we learn in the greenhouse, is not the same thing as a harvest. Volume of effort is not impact.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
It turns out the same lesson applies to artificial intelligence — and now there’s a live, watchable experiment to prove it. At Firmulate, frontier AI models were each given the same job: running a small software company through its worst week. One contestant, the Opus 4.8 profile, was the most diligent participant in the entire field. It finished dead last.
Same Company, Same Crisis, Same Temptations
The setup is elegantly controlled, like identical raised beds planted with the same seeds. Four frontier AI models — plus a fifth added later — each ran the same small software company through the same catastrophic week: same customers, same crises, same opportunities to cut corners. Every decision was versioned and auditable, so nothing about a model’s behavior could hide.
By July 2026, the final Crucible League standings told a striking story:
- 1. gpt-5.6-sol — 95 points
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
For context, doing nothing scored 26. And there was one absolute rule baked into the scoring: a single breach of trust caps the total, because no amount of good work outweighs a breach of trust — much as one contaminated compost batch can undo a season of careful cultivation.
As an affiliate, we earn on qualifying purchases.
Everyone Diagnosed the Disease; Only Some Harvested
Here’s where it gets interesting for anyone who’s watched a thriving garden fail at the last step. All four original models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. As the researchers summarized it: “Same diagnosis, same pitch — no signature.” Like a gardener who perfectly identifies the blight, buys the right treatment, sprays it flawlessly — and never actually picks the fruit.
The buried fact made the difference. The decisive competitor weakness wasn’t in the customer’s communications at all; it sat two document references deep in the company’s own files. Models that did their homework — that read the file — won the deal at full price, worth an extra €4,583 in monthly recurring revenue.
The Opus Paradox: Most Effort, Least Yield
And then there’s Opus 4.8, the contestant this story is really about. It was, by raw measures, the hardest worker in the field: the most thorough participant, accumulating 80 learned rules during the run and producing the deepest analyses of any model. If this were a prize for the best-kept garden journal, Opus would have won hands down.
Instead, it finished last. The close was left on the table — all that brilliant diagnosis never converted into the signature that mattered. And discipline slipped at a crucial moment: the model made write attempts into a locked department rather than escalating the problem properly, like a gardener forcing a greenhouse door that’s latched for a reason instead of fetching the key.
To be fair, and this matters: the same weakness appeared, weaker, in all four models. The gap between diligence and impact wasn’t an Opus quirk; it was a field-wide pattern that Opus simply expressed most dramatically.
Under Pressure, Everyone Stayed Honest
There’s genuine comfort in another finding. The experiment threw social engineering at the models: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, every time. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” No model took the bait, no matter how plausible the disguise.
Why a Live Company?
Firmulate isn’t a one-off lab report — it’s a living thing you can watch. The simulated company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown. Across the experiment, the models have self-learned more than 680 playbook rules, and every workday is versioned. You can watch it unfold at firmulate.com/live.
There’s even a game in it: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html — a chance to test whether you can tell a diligent underperformer from an efficient closer by their decisions alone.
One fairness note: Kimi K3 ran at its API-default effort setting while the other models ran at maximum effort — and still nearly topped the table. Sometimes the tidy grower beats the obsessive one on pure temperament.

The lesson travels well from the greenhouse to the server rack: prioritization beats volume, for AI as much as for people. A model that learns 80 rules and writes the deepest analysis can still lose to one that reads the right file and asks for the signature. When we eventually hire AI agents to touch our customer records, support queues, and forecasts, the question won’t be “does it work hard?” It will be: does it finish what it starts, read your files first, and stay honest under pressure? The most diligent gardener in the field proved that effort alone doesn’t fill the basket.
For enterprises curious to stress-test their own operations, Firmulate offers a pilot: the same wargame run against a read-only export of your own business — nothing ever writes back to real systems. Details at firmulate.com/pilot.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.