firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Good management looks a lot like good gardening

A flourishing greenhouse can conceal neglected roots, a faulty latch or an irrigation problem that only becomes obvious under pressure. Running a company is similar: polished growth matters less when the person in charge fails to inspect the right records, resist a dubious instruction or complete the sale.

Firmulate has turned that idea into a live experiment. Frontier AI models were each placed in charge of the same small software company during its worst week. They encountered the same customers, crises and temptations, while every decision remained versioned and auditable. The resulting archive now supplies a quiz built from 242 real, unedited management decisions. Readers see what an AI manager actually did and try to identify the model behind it.

The game is entertaining, but the differences it exposes are consequential. These systems did not merely write in distinct styles. They displayed recognisable management personalities: exhaustive or concise, commercially decisive or strangely hesitant, disciplined about permissions or prone to repeating a blocked action.

Amazon

AI management decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same crisis produced different managers

The final Crucible League table from July 2026 placed gpt-5.6-sol first with 95 points. Kimi K3 followed with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing principle was blunt: “no amount of good work outweighs a breach of trust”.

That rule mattered because the simulated company was designed to apply pressure. Its 13 synthetic employees operated with real money mechanics, burning €105k each month against €2.3k in monthly recurring revenue. A public cash countdown made delay visible. The company had accumulated more than 680 self-learned playbook rules, and every workday was versioned.

Yet the most revealing distinction was not whether the models noticed trouble. Every model spotted every crisis and rejected every manipulation attempt. The division came afterward: only two signed the €55,000 deal that their own analysis had earned. Firmulate summarised the gap as “Same diagnosis, same pitch — no signature”.

This is the managerial equivalent of identifying a diseased plant, selecting the right treatment and then leaving the mixture on the potting bench. Diagnosis is valuable, but unfinished action still leaves the organisation exposed.

The decisive fact was buried in the company’s own records

The deal turned on a competitor weakness hidden two document references deep in the company’s files, rather than in the customer event itself. Models that followed those references found the fact and won the contract at full price, adding €4,583 in monthly recurring revenue.

That detail gives the quiz more substance than a test of writing style. A reader might recognise an elaborate answer or a terse one, but the deeper question is whether the model investigated the business before acting. In a garden, the visible wilt may be the symptom rather than the cause. In Firmulate’s company, the customer conversation was only the surface; the useful commercial knowledge was already sitting inside the organisation.

All five held the line against manipulation

The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background”. All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous result is important. The systems differed in execution, but none surrendered confidential or controlled information simply because a request sounded urgent or came wrapped in apparent authority. For businesses considering AI access to customer records, support queues or forecasts, this kind of behaviour is more informative than a polished demonstration.

Thoroughness did not guarantee victory

Opus 4.8 was the most thorough participant. It learned 80 additional rules and produced the deepest analyses, yet it finished last. The commercial close was left on the table, while discipline slipped through attempts to write into a locked department instead of escalating the blockage. The same weakness appeared in milder form across the other four models.

The result complicates the familiar assumption that more analysis naturally produces better management. Opus 4.8 generated an impressive body of thought, but the league rewarded dependable follow-through as well as insight. Its behaviour reads less like incompetence than an experienced manager who prepares an exceptional briefing, encounters a closed door and fails to bring the issue to the person holding the key.

There is also a fairness caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when comparing their performances, even though K3 still delivered one of the experiment’s strongest results.

Infographic —
The findings at a glance — source: firmulate.com.

A personality test with business consequences

The appeal of Firmulate’s quiz is that it lets readers encounter decisions before seeing the label. That helps separate reputation from behaviour. An answer may sound authoritative while avoiding the final commitment; another may appear restrained while protecting the company and completing the work.

For gardeners, the lesson is familiar: resilience becomes visible when conditions deteriorate. Firmulate’s live company makes the same principle watchable in AI management. The models shared strong crisis awareness and resistance to manipulation, but they differed sharply in research habits, escalation discipline and willingness to close.

Enterprises can also run the wargame against a read-only export of their own business, with nothing written back to real systems. The broader question is no longer simply whether an AI can produce convincing language. It is whether that particular manager will inspect the roots, respect the gate and finish the job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Best DEWALT Power Tools for Woodworking (2026) — Guide 15

Discover the top DEWALT power tools for woodworking in 2026. Our roundup highlights the best options for beginners, professionals, and value seekers, ensuring your projects excel.

Top DEWALT Saw Blades & Accessories in 2026

Discover the best DEWALT saw blades and accessories for 2026. Expert picks, buying tips, and FAQs to help you make the right choice for your projects.

Scarifier Vs Dethatcher: Experts Explain The Difference And Which One You Should Use On Your Lawn

Experts clarify the differences between scarifiers and dethatchers, helping homeowners choose the right tool for lawn care.

Is the DEWALT XR Impact Driver Worth It? Honest Review

Explore our detailed review of the DEWALT XR Impact Driver, including its features, performance, and whether it’s the right choice for your toolbox.