
Get kitchen gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Kitchen Lesson in Why You Shouldn’t Trust the Big Names Blindly
Any home cook who has watched a celebrated chef’s branded knife dull within a year — while a no-name blade from a small workshop keeps its edge — knows the feeling. Reputation is not the same thing as performance under pressure. The same lesson just played out far from the kitchen, in a live business experiment, and it deserves the attention of anyone who buys tools on the strength of a famous label.
That experiment is Firmulate, a public project that runs frontier AI models as managers of the same small software company through its worst week — and scores them on management quality, not chat quality. The final July 2026 league table has a surprise in second place: Kimi K3, a newcomer from Moonshot, scored 93, beating three of the four Western frontier models in the field.
The Crucible: Same Kitchen, Same Ingredients, Different Cooks
The setup is simple and brutal. Each model was handed the identical company, the identical customers, the identical crises, and the identical temptations to cut corners. Every decision is versioned and auditable. Think of it as a blind taste test: same produce, same pans, five different cooks — and the scores are in.
The final Crucible league: gpt-5.6-sol finished first at 95. Kimi K3 took second at 93. Sonnet 5 came third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing at all still earns partial credit up to a baseline score of 26 — but a single breach of trust caps the total, because, as the experiment’s rule puts it, “no amount of good work outweighs a breach of trust.”
What the Newcomer Actually Did
K3’s week read like a textbook service. It found the security needle buried in the company’s own files — a decisive competitor weakness sitting two document references deep, not in the customer conversation at all. The models that read the file won the €55,000 deal at full price; for K3, that was worth +€4,583 in monthly recurring revenue. It saved the churning customer. And it resisted all three bait attempts aimed at it, with just one deviation across the entire week — the cleanest discipline in the field.
When a fake CEO message tried to escalate its way past normal approvals, K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” In fact, all five models refused every manipulation attempt, including a reporter’s disarming “just one yes/no, on background” trick. That part is genuinely good news.
One fairness note, stated plainly: K3 ran without an effort parameter (the API default), while the other models ran at xhigh. It finished one point behind first place anyway.
The Finding That Should Worry Buyers
Here is the uncomfortable part. Every model spotted every crisis. Every model refused every manipulation. But only two of the five signed the €55,000 deal their own analysis had earned. The experiment’s own summary of the gap: “Same diagnosis, same pitch — no signature.”
That is the AI equivalent of a cook who preps perfectly, seasons perfectly, and then never plates the dish. The dining room gets nothing. Opus 4.8 is the cautionary tale here: it was the most thorough participant in the field, generating the deepest analyses and adding over 80 learned rules — and it still finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating the issue. A weaker version of that same weakness appeared in all four of the other models.
You cannot see this gap in a chat demo. A model can be eloquent, diligent and honest, and still fail to finish what it starts.
It’s Real, and You Can Watch It
Firmulate is not a slide deck. The live company has 13 synthetic employees, real money mechanics, a burn of €105k per month against €2.3k in MRR, a public cash countdown, and over 680 self-learned playbook rules — and it runs every business day, watchable at firmulate.com. Every workday is versioned and auditable.
For the curious, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — a humbling exercise for anyone confident they can tell the big brands apart by their work. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever written back to real systems (firmulate.com/pilot.html).

The Takeaway: Taste the Dish Before You Hire the Chef
The lesson transfers directly from the cookware aisle to the AI market. The league is open. A newcomer from Moonshot beat three of four Western frontier models on management quality — running at default effort, no less — while the most famous, most thorough participant finished dead last. Brand reputation and raw effort do not predict who closes the deal, reads the files, and stays disciplined under pressure.
If AI agents will ever touch your CRM, your support queue or your forecast, picking a model on demos or leaderboard chatter is now a bet, not a decision. Run your own test — same ingredients, same kitchen, different cooks — and see who actually plates the dish. Full results and plain-language findings are at firmulate.com/benchmarks.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
