
Imagine running your restaurant kitchen with an AI that scores 26 out of 100 just by doing nothing. Sounds useless, right? But in AI benchmarking, that 26 is a crucial starting point — a baseline that reveals how trustworthy, honest, and effective these models truly are. For business leaders, understanding this number is key to knowing what to expect from AI in real-world scenarios.
Get kitchen gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the Baseline Score: Why 26?
In a recent open experiment conducted by Firmulate, four leading AI models were tested by running a simulated small software company through its worst week. They faced the same customer crises, temptations to cheat, and manipulative tactics. The surprising result? Even a model that does nothing at all, just by being present in the system, scores a baseline of 26 points out of 100. This score isn’t arbitrary; it reflects the minimum level of partial progress that an AI can achieve simply by being engaged, even if it doesn’t actively solve problems.
Why isn’t it zero? Because even a passive AI reads documents, detects crises, and processes information, which counts as partial progress. This baseline ensures that models aren’t just scored on their ability to produce impressive chat responses but on their capacity to handle real management tasks with honesty and discipline.
As an affiliate, we earn on qualifying purchases.
The Power of Progress and the Limits of Trust
The experiment underscores a crucial rule: any breach of trust caps the total score. Even if an AI model performs well elsewhere, a single slip — such as attempting to manipulate data or signing a deal it shouldn’t — can trigger a trust penalty that limits the overall grade. This rule mirrors real business priorities: honesty and integrity are non-negotiable, and no amount of good work can outweigh a breach of trust.
AI trust and integrity assessment software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Models Performed: Not Just About Crises
All models successfully identified every crisis and refused manipulative tactics, including fake CEO messages and reporter tricks. Interestingly, the key advantage was found not in the obvious customer crises but in reading internal documents. The competitor that read two document references deep into the company’s files won the deal at full price — worth over €4,583 monthly recurring revenue. This highlights that effective AI must go beyond surface-level tasks to truly understand and leverage internal data.
As an affiliate, we earn on qualifying purchases.
Assessing Discipline and Decision-Making
Among the participants, Opus 4.8, with over 80 learned rules and the deepest analyses, finished last — missing the chance to close a deal and slipping into departmental silos instead of escalating critical issues. This demonstrates that thoroughness alone isn’t enough; disciplined decision-making and trustworthiness matter just as much.
AI performance benchmarking kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business Operations
For companies considering AI integration, the message is clear: performance isn’t just about generating content or handling customer queries. It’s about finishing tasks reliably, reading relevant internal data, and maintaining honesty under pressure. An AI that scores 26 just by being present already indicates a baseline level of engagement, but true usefulness requires higher scores — which only some models can achieve.
Monitoring and Testing Your AI Workforce
Firmulate offers interactive tools and live experiments to help businesses test their AI models before deploying them in real operations. These tests simulate crises, temptations, and manipulative tactics, giving a clear picture of AI’s trustworthiness and discipline. Leaders can run these scenarios against their own AI models, ensuring that what seems promising in demos withstands the rigors of actual management tasks.
Final Takeaway
The key lesson from this benchmarking is that a do-nothing baseline score of 26 isn’t just a score — it’s a warning sign. It shows the minimum engagement level, but also highlights the importance of discipline, internal understanding, and trustworthiness in AI models. For businesses, the goal isn’t simply to have an AI that can chat well; it’s to have one that can finish what it starts, stay honest, and deliver real value under pressure.
To see these tests in action and understand where your AI stands, check out Firmulate’s live platform and interactive quizzes at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
