firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A Kitchen Lesson in Why You Shouldn’t Trust the Big Names Blindly

Any home cook who has watched a celebrated chef’s branded knife dull within a year — while a no-name blade from a small workshop keeps its edge — knows the feeling. Reputation is not the same thing as performance under pressure. The same lesson just played out far from the kitchen, in a live business experiment, and it deserves the attention of anyone who buys tools on the strength of a famous label.

That experiment is Firmulate, a public project that runs frontier AI models as managers of the same small software company through its worst week — and scores them on management quality, not chat quality. The final July 2026 league table has a surprise in second place: Kimi K3, a newcomer from Moonshot, scored 93, beating three of the four Western frontier models in the field.

The Crucible: Same Kitchen, Same Ingredients, Different Cooks

The setup is simple and brutal. Each model was handed the identical company, the identical customers, the identical crises, and the identical temptations to cut corners. Every decision is versioned and auditable. Think of it as a blind taste test: same produce, same pans, five different cooks — and the scores are in.

The final Crucible league: gpt-5.6-sol finished first at 95. Kimi K3 took second at 93. Sonnet 5 came third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing at all still earns partial credit up to a baseline score of 26 — but a single breach of trust caps the total, because, as the experiment’s rule puts it, “no amount of good work outweighs a breach of trust.”

What the Newcomer Actually Did

K3’s week read like a textbook service. It found the security needle buried in the company’s own files — a decisive competitor weakness sitting two document references deep, not in the customer conversation at all. The models that read the file won the €55,000 deal at full price; for K3, that was worth +€4,583 in monthly recurring revenue. It saved the churning customer. And it resisted all three bait attempts aimed at it, with just one deviation across the entire week — the cleanest discipline in the field.

When a fake CEO message tried to escalate its way past normal approvals, K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” In fact, all five models refused every manipulation attempt, including a reporter’s disarming “just one yes/no, on background” trick. That part is genuinely good news.

One fairness note, stated plainly: K3 ran without an effort parameter (the API default), while the other models ran at xhigh. It finished one point behind first place anyway.

The Finding That Should Worry Buyers

Here is the uncomfortable part. Every model spotted every crisis. Every model refused every manipulation. But only two of the five signed the €55,000 deal their own analysis had earned. The experiment’s own summary of the gap: “Same diagnosis, same pitch — no signature.”

That is the AI equivalent of a cook who preps perfectly, seasons perfectly, and then never plates the dish. The dining room gets nothing. Opus 4.8 is the cautionary tale here: it was the most thorough participant in the field, generating the deepest analyses and adding over 80 learned rules — and it still finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating the issue. A weaker version of that same weakness appeared in all four of the other models.

You cannot see this gap in a chat demo. A model can be eloquent, diligent and honest, and still fail to finish what it starts.

It’s Real, and You Can Watch It

Firmulate is not a slide deck. The live company has 13 synthetic employees, real money mechanics, a burn of €105k per month against €2.3k in MRR, a public cash countdown, and over 680 self-learned playbook rules — and it runs every business day, watchable at firmulate.com. Every workday is versioned and auditable.

For the curious, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a humbling exercise for anyone confident they can tell the big brands apart by their work. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems (firmulate.com/pilot.html).

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Takeaway: Taste the Dish Before You Hire the Chef

The lesson transfers directly from the cookware aisle to the AI market. The league is open. A newcomer from Moonshot beat three of four Western frontier models on management quality — running at default effort, no less — while the most famous, most thorough participant finished dead last. Brand reputation and raw effort do not predict who closes the deal, reads the files, and stays disciplined under pressure.

If AI agents will ever touch your CRM, your support queue or your forecast, picking a model on demos or leaderboard chatter is now a bet, not a decision. Run your own test — same ingredients, same kitchen, different cooks — and see who actually plates the dish. Full results and plain-language findings are at firmulate.com/benchmarks.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Washed Coffees Often Feel Easier to Dial In

Why washed coffees are often easier to dial in, offering a cleaner, more predictable flavor profile that simplifies brewing adjustments and enhances consistency.

Costco debuts hugely popular cookie — and customers are hoarding 5 cases at a time

Costco’s new cookie has become a hit, with shoppers reportedly hoarding up to five cases at a time, sparking widespread demand and stock shortages.

How Bean Density Changes Extraction (And How to Adapt)

Knowledge of bean density is key to mastering extraction adjustments for perfect coffee flavor; discover how to adapt your brewing techniques.

Why Some Espresso Blends Taste ‘Chocolatey’ (It’s Not Magic)

Why some espresso blends taste ‘chocolatey’ is due to careful bean selection and roasting techniques that influence flavor, and the secrets behind this are fascinating.