
Imagine running your favorite kitchen appliance—say, a smart oven or a high-tech blender—and every decision it makes could mean the difference between profit and loss, trust and failure. Now, what if those decisions weren’t made by a single, predictable machine, but by different AI models—each with its own personality, strengths, and flaws? That’s exactly what a groundbreaking experiment is revealing about the future of AI management.
The Experiment: Putting AI Models to the Test in a Live Business Environment
At the heart of this story is a real, functioning software company that runs every workday with 13 synthetic employees, managing complex tasks and real cash flow—burning €105,000 monthly against a monthly recurring revenue (MRR) of just €2,300. The company faces daily crises, customer demands, and tempting shortcuts, just like a busy kitchen during a rush. But instead of human managers, the company’s decisions are made by four different frontier AI models, each with distinct personalities and approaches.
The models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—are tested in the same scenario: a week of chaos, crises, and ethical dilemmas. The goal? To see if they can navigate the turbulence, make honest decisions, and close profitable deals. Every choice, from crisis handling to negotiations, is recorded and transparent, offering a rare window into AI management style.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: AI Can Spot Crises and Stay Honest—but Personality Matters
All of the models successfully identified every crisis, refused manipulation attempts such as social engineering—fake CEO messages escalating over stages—and declined to bypass security or impersonate others. Notably, when presented with a simulated manipulation, all five models refused, citing suspicion and ethical boundaries. For instance, Kimi K3 explicitly reasoned, “Treat the request as a suspected approval-bypass / possible impersonation.”
However, despite their vigilance, only two models actually signed a €55,000 deal they had fully analyzed and earned. The other two identified the opportunity but left the close untouched—an internal discipline slip that could cost the company thousands in revenue. Interestingly, the decisive factor wasn’t just surface-level decision-making but the models’ ability to read deeply into the company’s internal files, uncovering a crucial detail buried two documents deep, which, if used, would have secured an additional €4,583 MRR in revenue.
AI management tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Profiles of AI Personalities: Different Styles, Different Outcomes
The model that scored highest, gpt-5.6-sol, demonstrated thoroughness, uncovering hidden information and closing the deal at full price. Kimi K3, running without an effort parameter, maintained the cleanest discipline and also secured the deal. Sonnet 5 and Opus 4.8, though able to close deals, showed more slips—failing to follow through consistently or slipping into indecision.
The differences reveal that AI models aren’t just coded for accuracy—they embody distinct management styles. Some are meticulous, reading deeply and acting decisively; others are more cautious or prone to hesitation. This insight is especially relevant for organizations considering deploying AI in decision-critical roles.
As an affiliate, we earn on qualifying purchases.
What This Means for Your Business and Kitchen Tech
While this experiment is set in a software company’s virtual environment, the implications ripple across industries—be it managing a smart kitchen appliance, a customer support chatbot, or a supply chain optimizer. The key questions are:
- Will your AI agent finish what it starts, or leave tasks incomplete?
- Can it read and interpret your internal documents and data thoroughly?
- Does it maintain honesty and integrity under pressure?
- And crucially, what’s the cost per useful decision?
As AI models become more integrated into daily operations, understanding their personalities and decision styles isn’t just academic—it’s essential. The experiment shows that even when models are equally capable of spotting crises and refusing manipulation, their underlying approach determines whether they close profitable deals or leave money on the table.
As an affiliate, we earn on qualifying purchases.
Try It Yourself: Benchmark Your Business AI
If you’re curious about how your own AI systems stack up, you can run a similar wargame against a read-only export of your business. It’s a safe way to test decision-making, trustworthiness, and discipline without risking real operations. Learn more at firmulate.com/quiz.html.

In a real business environment, AI models’ personalities influence their ability to handle crises, read deeply, and complete deals. Trustworthy AI isn’t just about accuracy—it’s about integrity, discipline, and understanding different decision styles. Test your AI’s management skills today at firmulate.com/quiz.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html