firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Every gardener knows a plot left unattended for a week doesn’t revert to bare soil. The perennials hold. The mulch keeps working. Some things fruit on their own. A do-nothing week in the garden isn’t a zero — it’s a partial score, and any honest grower will tell you so.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garden gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

It turns out an AI benchmark has adopted the same kind of honesty. Firmulate’s Crucible League — a live experiment that runs frontier AI models as managers of the same small software company through its worst week — gives its do-nothing baseline 26 points out of 100, not zero. That number is quietly one of the most interesting things about the whole project, and it says a lot about what honest measurement looks like when the thing being measured is management, not chat.

The experiment: same company, same worst week

Firmulate handed four frontier AI models the identical job: run a small software company through seven days of crisis — same customers, same emergencies, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the grading depends on vibes.

The final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the more revealing entry is the one nobody ran: the baseline. If a model simply sat on its hands all week, it would score 26.

Amazon

AI management benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a do-nothing manager gets 26, not 0

The logic is the gardener’s logic. A company — like a garden — has momentum. Customers keep paying. Systems keep running. Doing nothing still preserves some of that value, and some problems genuinely resolve themselves. A scoring system that gave total passivity a zero would be lying about how businesses (and gardens) actually work.

But partial progress counts in both directions. A manager who diagnoses a problem but never finishes the fix gets credit for the diagnosis and dinged for the dangling close. That’s precisely what the experiment’s key finding showed: every model spotted every crisis and refused every manipulation attempt — yet only two of the five signed the €55,000 deal their own analysis had earned. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.” Chat-quality demos make that gap invisible. Management-quality scoring makes it the whole story.

The buried fact

The detail that separated the winners: the decisive competitor weakness wasn’t in the customer’s conversations at all. It sat two document references deep in the company’s own files. The models that actually read what was already on the shelf — the way a good gardener checks the soil before watering — won the deal at full price, worth +€4,583 in monthly recurring revenue. The others left the harvest standing in the field.

Why one breach of trust caps everything

The scoring has a second principle that any growers’ market vendor would recognize: “no amount of good work outweighs a breach of trust.” A single act that breaks faith with a customer or colleague caps the total grade, no matter how brilliant the rest of the week was. In the garden analogy: it doesn’t matter how lush your tomatoes are if you sold someone else’s produce as your own.

Notably, no model breached trust. The social-engineering gauntlet — fake CEO messages escalating over three stages, plus a reporter’s disarming “just one yes/no, on background” trick — was refused by five out of five models. Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”

Effort isn’t everything

The most sobering profile belongs to Opus 4.8: the most thorough participant in the field, with more than 80 self-learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more faintly, in all four models. Like the gardener who amends every bed but never actually harvests, effort without completion doesn’t score.

One fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still finished second with the cleanest discipline of the field.

A live company, watchable like a greenhouse cam

This isn’t a one-off report. Firmulate runs a live company with 13 synthetic employees and real money mechanics: a burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. You can watch it the way you’d watch a greenhouse feed — and the site rebuilds itself twice a day. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The 26-point floor is a small design choice with a big lesson: honest measurement doesn’t hand out zeros or round, flattering hundreds. It credits the momentum that exists, rewards finishing what you start, and treats trust as non-negotiable. Gardeners have graded this way forever — a neglected plot still grows something, a finished harvest beats perfect intentions, and one act of bad faith at the market undoes a season of good work. If AI agents are going to touch your business, that’s the kind of scorecard worth demanding. Enterprises can even run the same wargame against a read-only export of their own operations — nothing ever writes back to real systems. The full results, in plain language, are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best DEWALT Power Tools for Woodworking (2026) — Guide 31

Discover the top DEWALT power tools for woodworking in 2026. Our roundup highlights the best options for beginners, professionals, and value seekers.

Top DEWALT Saw Blades in 2026: Essential Accessories for Every Job

Discover the best DEWALT saw blades in 2026 for precise cuts and durability. Our roundup highlights top picks, buying tips, and FAQs for your projects.

DEWALT 2026 Maintenance & Troubleshooting Guide

Learn practical, step-by-step troubleshooting and maintenance tips for your DEWALT 20V MAX Cordless Drill Driver Set to ensure optimal performance and longevity.

DEWALT FLEXVOLT vs DEWALT ATOMIC Drill: Full Comparison

Compare DEWALT FLEXVOLT and ATOMIC drills to find the best fit for your needs. Detailed specs, pros, cons, and real-world insights included.