firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine running your restaurant kitchen with an AI that scores 26 out of 100 just by doing nothing. Sounds useless, right? But in AI benchmarking, that 26 is a crucial starting point — a baseline that reveals how trustworthy, honest, and effective these models truly are. For business leaders, understanding this number is key to knowing what to expect from AI in real-world scenarios.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Baseline Score: Why 26?

In a recent open experiment conducted by Firmulate, four leading AI models were tested by running a simulated small software company through its worst week. They faced the same customer crises, temptations to cheat, and manipulative tactics. The surprising result? Even a model that does nothing at all, just by being present in the system, scores a baseline of 26 points out of 100. This score isn’t arbitrary; it reflects the minimum level of partial progress that an AI can achieve simply by being engaged, even if it doesn’t actively solve problems.

Why isn’t it zero? Because even a passive AI reads documents, detects crises, and processes information, which counts as partial progress. This baseline ensures that models aren’t just scored on their ability to produce impressive chat responses but on their capacity to handle real management tasks with honesty and discipline.

Amazon

AI model testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Power of Progress and the Limits of Trust

The experiment underscores a crucial rule: any breach of trust caps the total score. Even if an AI model performs well elsewhere, a single slip — such as attempting to manipulate data or signing a deal it shouldn’t — can trigger a trust penalty that limits the overall grade. This rule mirrors real business priorities: honesty and integrity are non-negotiable, and no amount of good work can outweigh a breach of trust.

Amazon

AI trust and integrity assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Models Performed: Not Just About Crises

All models successfully identified every crisis and refused manipulative tactics, including fake CEO messages and reporter tricks. Interestingly, the key advantage was found not in the obvious customer crises but in reading internal documents. The competitor that read two document references deep into the company’s files won the deal at full price — worth over €4,583 monthly recurring revenue. This highlights that effective AI must go beyond surface-level tasks to truly understand and leverage internal data.

Amazon

internal data analysis AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Assessing Discipline and Decision-Making

Among the participants, Opus 4.8, with over 80 learned rules and the deepest analyses, finished last — missing the chance to close a deal and slipping into departmental silos instead of escalating critical issues. This demonstrates that thoroughness alone isn’t enough; disciplined decision-making and trustworthiness matter just as much.

Amazon

AI performance benchmarking kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business Operations

For companies considering AI integration, the message is clear: performance isn’t just about generating content or handling customer queries. It’s about finishing tasks reliably, reading relevant internal data, and maintaining honesty under pressure. An AI that scores 26 just by being present already indicates a baseline level of engagement, but true usefulness requires higher scores — which only some models can achieve.

Monitoring and Testing Your AI Workforce

Firmulate offers interactive tools and live experiments to help businesses test their AI models before deploying them in real operations. These tests simulate crises, temptations, and manipulative tactics, giving a clear picture of AI’s trustworthiness and discipline. Leaders can run these scenarios against their own AI models, ensuring that what seems promising in demos withstands the rigors of actual management tasks.

Final Takeaway

The key lesson from this benchmarking is that a do-nothing baseline score of 26 isn’t just a score — it’s a warning sign. It shows the minimum engagement level, but also highlights the importance of discipline, internal understanding, and trustworthiness in AI models. For businesses, the goal isn’t simply to have an AI that can chat well; it’s to have one that can finish what it starts, stay honest, and deliver real value under pressure.

To see these tests in action and understand where your AI stands, check out Firmulate’s live platform and interactive quizzes at firmulate.com/benchmarks.html.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Backerei Wahl

Backerei Wahl has announced a change in leadership as part of its strategic growth, confirmed by company officials. Details on the future direction remain unclear.

Costco debuts hugely popular cookie — and customers are hoarding 5 cases at a time

Costco’s new cookie has become a hit, with shoppers reportedly hoarding up to five cases at a time, sparking widespread demand and stock shortages.

Monti Carlo’s Caribbean Cowboy Caviar

Monti Carlo introduces Caribbean Cowboy Caviar, a new appetizer blending tropical flavors with classic ingredients, set to hit menus nationwide.

The Stale Bean Sign Most People Miss

A subtle radiographic clue, the Stale Bean Sign hints at underlying lung issues, but its elusive nature requires careful attention to uncover its full significance.