
Imagine a boat race where the fastest boat isn’t just about speed but about navigating every obstacle flawlessly. In the world of AI decision-making, diligence and volume alone don’t guarantee victory. The real key? Prioritization and discipline. A groundbreaking live experiment reveals how even the most thorough AI can falter under pressure, leaving the most critical deals on the table.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Test: AI Under Pressure in a Simulated Business Environment
To understand how AI models handle complex, real-world decisions, researchers at Firmulate ran an innovative experiment. They tasked four cutting-edge AI models with managing a small software company during its worst week—full of crises, manipulations, and temptations. Every decision was carefully recorded and auditable, creating a transparent comparison of performance.
The Participants and Performance Scores
- gpt-5.6-sol scored 95, found the critical information buried in the company’s files, and successfully closed the €55,000 deal.
- Kimi K3 scored 93, maintained the most discipline, and also closed the deal.
- Sonnet 5 scored 88, closed the deal with minor slips.
- Fable 5 scored 77, closed the deal but demonstrated more process issues.
- Baseline (do-nothing) scored 26, highlighting how little progress can be made without active decision-making.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading the Files Matters Most
While all models recognized the crises and refused manipulative tactics—such as fake CEO messages and staged reporter inquiries—the decisive factor was how deep they read into the company’s documentation. Only the top performers identified the buried data point that was critical to closing the deal. Those who read more thoroughly gained significant advantages, including a potential additional €4,583 in Monthly Recurring Revenue (MRR).
Social Engineering Tests and Integrity
All models refused to be manipulated through staged social engineering attempts, treating such requests as impersonation risks. Kimi K3 explicitly reasoned that such requests could bypass approval processes or indicate impersonation, demonstrating sound judgment under pressure.
As an affiliate, we earn on qualifying purchases.
The Real-World Implication: Diligence Is Not Enough
The experiment’s real-world relevance is profound. Many AI systems are trained to be diligent—spotting crises, refusing manipulations, and following rules. Yet, even with over 80 learned rules and deep analyses, the most thorough model—Opus 4.8—still finished last because discipline slipped during the critical closing phase. Instead of escalating issues or verifying information, the model’s decision-making process faltered, leaving valuable opportunities on the table.
What This Means for Business and AI Integration
For companies considering AI in customer support, sales, or decision-making roles, the takeaway is clear: volume and rules alone won’t ensure success. Focusing on prioritization, reading comprehension, and discipline is crucial. An AI that reads deeply, stays honest under pressure, and follows clear judgment calls will outperform a diligence-focused counterpart that simply detects crises.
As an affiliate, we earn on qualifying purchases.
Live Experiment: Watch It in Action
The entire process is observable live at firmulate.com/live. You can see how each AI model navigates real crises, handles manipulation attempts, and strives to close deals amid simulated business chaos. These insights are invaluable for anyone deploying AI in mission-critical environments.
Final Reflection: The Cost of Discipline
In this experimental league, the winning models didn’t just find the right information—they prioritized correctly and maintained discipline, avoiding shortcuts that undermine trust and success. The deep analysis by Opus 4.8, despite being the most thorough, was insufficient without disciplined escalation. This underscores a vital lesson: in AI decision-making, diligence should be coupled with strategic prioritization.

AI models can recognize crises and reject manipulation, but true success hinges on prioritization and discipline—reading deeply, verifying, and staying honest under pressure. Volume alone won’t win deals or trust. For businesses, the key is training AI to read smarter, act decisively, and prioritize what matters most.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.