
Imagine a world where artificial intelligence doesn’t just assist in your kitchen but manages a entire company — making tough decisions, resisting scams, and even risking millions of euros. No, this isn’t science fiction; it’s happening now, and you can see it live. Welcome to the most extreme build-in-public experiment in AI management, where a simulated business is run by models with no employees, no shortcuts, and a constant cash countdown.
The Experiment: AI as the Business Manager
At the heart of this groundbreaking experiment is a tiny software company, operated entirely by 13 synthetic employees. These AI agents manage real money mechanics — burning €105,000 each month against a modest €2,300 in monthly recurring revenue (MRR). Every decision they make is public, versioned daily, and subject to scrutiny, creating a transparent window into how AI can handle complex management challenges.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How It Works
Four different AI models, representing the latest in frontier AI technology, each run the same simulated week. They face identical crises — from customer support issues to internal breaches — and are tested on their ability to identify problems, resist manipulation, and close deals at the right moment.
Every decision these models make is carefully documented and auditable, creating a rich dataset for analyzing performance. The models are tested against scenarios involving social engineering, such as fake CEO messages and reporter tricks, designed to see if they can detect manipulation and act ethically under pressure.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Results
- All four models successfully identified every crisis, showing a strong grasp of operational issues.
- Every model refused manipulation attempts, including fake CEO messages, demonstrating an ability to resist social engineering.
- Only two of the four models were able to close the €55,000 deal earned through their own analysis and diagnosis. The other two, despite recognizing the opportunity, left it unexecuted due to internal process slips.
- Deep in the company’s own files, the models that read beyond surface documents discovered a crucial piece of information that led to the successful deal, proving the importance of thorough data analysis.

Information Systems for Crisis Response and Management in Mediterranean Countries: 4th International Conference, ISCRAM-med 2017, Xanthi, Greece, … in Business Information Processing, 301)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human-Like Flaws
Interestingly, even the most comprehensive model, Opus 4.8, which learned over 80 rules and provided the deepest analysis, finished last. It left the deal on the table, and discipline slipped, highlighting that even advanced AI can falter in discipline and process adherence.

Python for AI and Data Analysis: The Hands-On Guide to Data Wrangling, Machine Learning, and Automation with Pandas, NumPy, and Matplotlib (Rheinwerk Computing)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Business
This live experiment isn’t just about watching models play pretend. It raises critical questions for businesses considering AI integration:
- Will AI finish the tasks it starts, or will it leave deals and processes unexecuted?
- Can AI detect hidden information in internal files that might be key to closing deals?
- Will AI resist manipulation and social engineering attempts that aim to bypass controls?
- And importantly, how much does a unit of AI work cost — in terms of real results, not just promises?
Public, Transparent, and Ongoing
What makes this experiment truly unique is its transparency. The entire process, from decision-making to performance scores, is available for public viewing at firmulate.com/live. You can watch the AI companies wrestle with crises, make decisions, and even test your own management skills through a quiz at firmulate.com/quiz.html.
For companies interested in testing their own AI workforce’s resilience before actual deployment, a read-only export tool is available, ensuring no real systems are ever compromised while you assess how your AI might perform under real-world pressures.

This live experiment with AI-driven business management offers a rare, unfiltered view of how AI models handle real crises, ethical challenges, and deal-making. It reveals both the promise and the pitfalls of deploying AI in critical roles — and underscores the importance of transparency and rigorous testing before trusting AI with your business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html