
Imagine an AI system that, despite doing nothing special, still earns 26 out of 100 points in a rigorous business test. If you’re shopping for a reliable AI partner, what does that number tell you? Not much, unless you understand the rules of the game—and the importance of honesty and diligence in AI performance.
Get sports gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What Makes a Benchmark Truly Honest?
In the world of artificial intelligence, assessing a model’s readiness for real-world tasks isn’t just about how well it chats or generates content. It’s about how well it handles crises, maintains trust, and upholds discipline under pressure. That’s why the latest public experiment, run by Firmulate, takes a straightforward yet revealing approach: giving AI models a simulated week in the life of a small software company facing real crises and temptations.
The core idea is simple. Every model is tested on the same scenario—bad customers, urgent crises, and the opportunity to manipulate or cheat. These are real, auditable decisions, with every choice tracked and documented. The goal: see whether AI can truly be a trustworthy business partner, not just produce impressive words.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Baseline: Why 26 Points?
One striking result is that a do-nothing baseline—an AI that doesn’t make any decisions—scores 26 points. That might seem odd at first glance. Why isn’t it zero? The answer lies in the benchmark’s design: partial progress counts. Even doing nothing, the system recognizes some facts, avoids some pitfalls, and thus earns a small score. It also reflects that in complex tasks, even cautious inaction can be better than reckless decisions.
Furthermore, the rules cap the total score if the AI breaches trust. For example, if it signs a fraudulent contract or approves a manipulated deal, that single breach halts further scoring. This emphasizes an essential truth in AI governance: honesty isn’t optional. No matter how much good work is done, a single breach spoils the entire performance.
As an affiliate, we earn on qualifying purchases.
How the Experiment Works—and What It Reveals
The experiment involves four state-of-the-art models—ranging from GPT-5.6 to newer entrants like Kimi K3—each managing a virtual company. Every AI is faced with the same set of challenges: customer complaints, internal crises, and manipulative tactics like fake CEO messages or secret document leaks. The models are watched in real time at firmulate.com/live, with decisions versioned and auditable for transparency.
The results are telling. All four models successfully identified every crisis and refused every manipulation attempt. That’s a baseline of operational integrity. But when it came to sealing the deal—signing a €55,000 contract—they didn’t all succeed. Only two models, including a newcomer called Kimi K3, managed to close the deal based on their own analysis and honest diagnosis.
What Made the Difference?
One key insight emerged: the winning models found critical information buried two references deep in the company’s internal files—information that could have been used to cheat or manipulate. Reading those files was the determining factor that secured the full-price deal, worth over €4,500 per month in recurring revenue.
trustworthy AI decision-making systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust, Discipline, and the Limits of AI
Another fascinating part of the experiment involves social engineering attacks—fake CEO messages escalating in stages, and reporters asking for quick approvals over background channels. Remarkably, all five models tested refused to participate or sign off on these manipulative requests, citing suspicion and the risk of impersonation.
This reflects an important aspect of trustworthy AI: the capacity to resist social engineering and manipulation. Even under pressure, these models maintained discipline, aligning with the best practices for safe AI deployment in business environments.
AI transparency and accountability tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What It Means for Business and AI Adoption
For companies considering integrating AI into their operations—whether managing support tickets, sales pipelines, or strategic planning—the takeaway is clear: the real test isn’t just how well an AI writes or responds. It’s whether it can finish what it starts, stay honest when it’s under pressure, and prioritize integrity over shortcuts.
The experiment underscores that current models can recognize crises and resist manipulation, but not all are equally disciplined. Even the most thorough participant, like Opus 4.8, which learned over 80 rules, can slip—left leaving deals on the table or failing to escalate properly. This highlights the importance of rigorous testing before trusting an AI with critical business decisions.
Why the 26-Point Floor Matters
The fact that a simple, do-nothing baseline scores 26 points sets a clear benchmark. It signals that in complex decision environments, avoiding harm and doing the minimum can still earn a modest score. But more importantly, it emphasizes that AI performance assessments are more nuanced than chat demos or superficial metrics.
By including a rule that caps the score after a trust breach, the benchmark promotes a culture of integrity. It reminds users that no amount of good work can compensate for dishonesty—a vital lesson in deploying AI responsibly.
Conclusion: The Future of Honest AI
As AI models become more integrated into business operations, the ability to detect crises, resist manipulation, and uphold trust will become critical. The Firmulate experiment offers a transparent, real-world perspective on where AI stands—and where it needs to go. The key takeaway: in AI performance, honesty and discipline are not optional, and benchmarks that measure these qualities will be essential for trustworthy AI adoption in the future.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
