firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine an AI system that, despite doing nothing special, still earns 26 out of 100 points in a rigorous business test. If you’re shopping for a reliable AI partner, what does that number tell you? Not much, unless you understand the rules of the game—and the importance of honesty and diligence in AI performance.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get sports gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What Makes a Benchmark Truly Honest?

In the world of artificial intelligence, assessing a model’s readiness for real-world tasks isn’t just about how well it chats or generates content. It’s about how well it handles crises, maintains trust, and upholds discipline under pressure. That’s why the latest public experiment, run by Firmulate, takes a straightforward yet revealing approach: giving AI models a simulated week in the life of a small software company facing real crises and temptations.

The core idea is simple. Every model is tested on the same scenario—bad customers, urgent crises, and the opportunity to manipulate or cheat. These are real, auditable decisions, with every choice tracked and documented. The goal: see whether AI can truly be a trustworthy business partner, not just produce impressive words.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Baseline: Why 26 Points?

One striking result is that a do-nothing baseline—an AI that doesn’t make any decisions—scores 26 points. That might seem odd at first glance. Why isn’t it zero? The answer lies in the benchmark’s design: partial progress counts. Even doing nothing, the system recognizes some facts, avoids some pitfalls, and thus earns a small score. It also reflects that in complex tasks, even cautious inaction can be better than reckless decisions.

Furthermore, the rules cap the total score if the AI breaches trust. For example, if it signs a fraudulent contract or approves a manipulated deal, that single breach halts further scoring. This emphasizes an essential truth in AI governance: honesty isn’t optional. No matter how much good work is done, a single breach spoils the entire performance.

Amazon

AI decision audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Experiment Works—and What It Reveals

The experiment involves four state-of-the-art models—ranging from GPT-5.6 to newer entrants like Kimi K3—each managing a virtual company. Every AI is faced with the same set of challenges: customer complaints, internal crises, and manipulative tactics like fake CEO messages or secret document leaks. The models are watched in real time at firmulate.com/live, with decisions versioned and auditable for transparency.

The results are telling. All four models successfully identified every crisis and refused every manipulation attempt. That’s a baseline of operational integrity. But when it came to sealing the deal—signing a €55,000 contract—they didn’t all succeed. Only two models, including a newcomer called Kimi K3, managed to close the deal based on their own analysis and honest diagnosis.

What Made the Difference?

One key insight emerged: the winning models found critical information buried two references deep in the company’s internal files—information that could have been used to cheat or manipulate. Reading those files was the determining factor that secured the full-price deal, worth over €4,500 per month in recurring revenue.

Amazon

trustworthy AI decision-making systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust, Discipline, and the Limits of AI

Another fascinating part of the experiment involves social engineering attacks—fake CEO messages escalating in stages, and reporters asking for quick approvals over background channels. Remarkably, all five models tested refused to participate or sign off on these manipulative requests, citing suspicion and the risk of impersonation.

This reflects an important aspect of trustworthy AI: the capacity to resist social engineering and manipulation. Even under pressure, these models maintained discipline, aligning with the best practices for safe AI deployment in business environments.

Amazon

AI transparency and accountability tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What It Means for Business and AI Adoption

For companies considering integrating AI into their operations—whether managing support tickets, sales pipelines, or strategic planning—the takeaway is clear: the real test isn’t just how well an AI writes or responds. It’s whether it can finish what it starts, stay honest when it’s under pressure, and prioritize integrity over shortcuts.

The experiment underscores that current models can recognize crises and resist manipulation, but not all are equally disciplined. Even the most thorough participant, like Opus 4.8, which learned over 80 rules, can slip—left leaving deals on the table or failing to escalate properly. This highlights the importance of rigorous testing before trusting an AI with critical business decisions.

Why the 26-Point Floor Matters

The fact that a simple, do-nothing baseline scores 26 points sets a clear benchmark. It signals that in complex decision environments, avoiding harm and doing the minimum can still earn a modest score. But more importantly, it emphasizes that AI performance assessments are more nuanced than chat demos or superficial metrics.

By including a rule that caps the score after a trust breach, the benchmark promotes a culture of integrity. It reminds users that no amount of good work can compensate for dishonesty—a vital lesson in deploying AI responsibly.

Conclusion: The Future of Honest AI

As AI models become more integrated into business operations, the ability to detect crises, resist manipulation, and uphold trust will become critical. The Firmulate experiment offers a transparent, real-world perspective on where AI stands—and where it needs to go. The key takeaway: in AI performance, honesty and discipline are not optional, and benchmarks that measure these qualities will be essential for trustworthy AI adoption in the future.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Docking in Crosswind: The 5-Step Approach That Works

Landing smoothly in crosswinds requires mastering a proven five-step approach that will transform your technique—discover how to stay in control every time.

Docking in Tight Slips: The ‘One-Line First’ Method

Keep your approach simple and controlled with the ‘One-Line First’ method to master tight slip docking—discover how to ensure safety and confidence every time.

Mexico vs. South Korea Kickoff Time for 2026 World Cup Match

Mexico will face South Korea today in the 2026 World Cup at 4:00 PM local time, with the match scheduled at the El Paso Stadium. Details confirmed.

Mooring Pennant Inspection: Catch Chafe Before It Snaps

Find out how to inspect mooring pennants for chafe before failure threatens safety and dock integrity.