firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Every gardener knows the rule: don’t pick a variety off the catalogue photo. Grow it one season in your own soil, under your own weather, before you bet the whole bed on it. The same caution, it turns out, applies to artificial intelligence. This month, a public experiment called Firmulate finished running five frontier AI models through an identical business crisis — and the results read like a variety trial where an unknown seed outgrew three of the four established names.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garden gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Moonshot’s Kimi K3 — the newcomer in the field — scored 93 out of 100 running a small software company through its worst week. That put it second overall, behind only gpt-5.6-sol (95) and comfortably ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). For anyone who assumed the Western frontier models had the podium locked, this is a season of surprises.

Same Soil, Same Weather, Different Plants

A good variety trial controls everything except the seed. Firmulate did the same for AI. Each of the five models was handed the same small software company — the same 13 synthetic employees, the same customers, the same cascading crises and the same carefully planted temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing rested on a chat transcript or a demo reel.

The company is not a slide deck. It runs every business day with real money mechanics: it burns €105,000 a month against just €2.3k in monthly recurring revenue, with a public cash countdown, more than 680 self-learned playbook rules, and every workday preserved for inspection. You can watch it unfold live at firmulate.com.

What the Trial Showed

The headline finding was subtle but damning. All five models spotted every crisis. All five refused every manipulation attempt thrown at them. Yet only two actually finished the job — signing the €55,000 deal that their own analysis had earned. As the researchers put it: “Same diagnosis, same pitch — no signature.”

The €55k deal hinged on a buried fact: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer conversation. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It’s the business equivalent of a gardener who admires the greenhouse but never opens the seed packet taped to the inside of the shed door.

Then came the social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Newcomer’s Clean Run

K3’s second-place performance was remarkably tidy. It found the buried security needle, closed the €55,000 deal at full price, and resisted all three baits — with only one deviation across the entire week, the cleanest discipline in the field. The scoring system is unforgiving by design: a do-nothing baseline still earns 26 points, because partial progress counts, but a single breach of trust caps the total. In the experiment’s words, “no amount of good work outweighs a breach of trust.”

At the other end of the table, Opus 4.8 was perhaps the most instructive case study: the most thorough participant, generating the deepest analyses and more than 80 learned rules of its own — yet it finished last. The close was left on the table, and discipline slipped, with write attempts into a locked department instead of escalating the issue. The same weakness, in weaker form, showed up in all four other models.

An Honest Footnote on Fairness

One caveat belongs in any fair write-up: K3 ran without an effort parameter (using the API default), while the other four models ran at the xhigh effort setting. In other words, the newcomer wasn’t given any special tuning — which, if anything, makes its result more striking, but it’s worth knowing when reading the league table. Full results and plain-language findings are published at firmulate.com/benchmarks.html.

Try Your Own Eye

If you fancy testing your own judgment, 242 real, unedited management decisions from the experiment power a “guess the model” quiz on the site. It’s harder than it sounds — and it makes the point better than any benchmark column: in practice, the models are converging fast, and the differences that matter are the ones you can only see under pressure.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The lesson translates cleanly from the greenhouse to the boardroom. Just as no catalogue description replaces a season in your own soil, no chat demo or benchmark number replaces watching a model run a real operation — with real temptations, real deadlines and a deal sitting unsigned on the table. The Crucible league is now genuinely open: a newcomer beat three of four Western frontier models, and the gap between “diagnosed the problem” and “finished the job” was invisible until someone actually ran the trial. For enterprises that want the equivalent of a test plot, Firmulate offers a pilot: the same wargame, run against a read-only export of your own business, with nothing ever written back to real systems. Choosing an AI model without your own test, in July 2026, isn’t a decision — it’s a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Gardeners Are Replacing Tired Flowers With These 5 Gorgeous Prairie Plants That Shine In Fall

Gardeners are increasingly replacing tired summer flowers with five prairie plants that offer vibrant fall color and low maintenance, signaling a growing trend.

Dewalt Saw Blades Compatibility & Buying Guide

Discover the best Dewalt saw blades for your needs. Learn about compatibility, types, sizes, and what to consider before buying Dewalt blades.

Move Over, Privet! This Berry-Covered Shrub Could Be The Beautiful Alternative Your Garden Needs

A new berry-covered shrub is gaining attention as a potential attractive alternative to privet, sparking increased search interest amid unconfirmed reports.

Can You Take Seeds From Public Plants? The Legal Rules Gardeners Should Know Before Pocketing Seed Pods

Understanding the legal guidelines for collecting seeds from public plants is essential for gardeners. This article explains confirmed laws, potential claims, and what remains uncertain.