
In the high-stakes world of sports and recreation, the difference between a good coach and a great one is often measured by their ability to handle pressure, read the field, and stay honest under fire. Now, imagine applying that same test to artificial intelligence—an AI that can run your business—not in theory, but in a real-time, live environment. That’s exactly what the latest experiment from Firmulate has achieved, pitting some of the most advanced frontier models against a real, functioning company during its toughest week.
Get sports gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Testing Arena: A Small Software Company Under Fire
The experiment involved four leading AI models, each tasked with managing a simulated small software company facing its worst week yet. The scenario was no mere demo; it was an authentic, live company with real money mechanics—€105,000 monthly burn rate against €2,300 MRR—and 13 synthetic employees operating under daily versioned rules. The firm’s ongoing cash countdown, real crises, customer churns, and temptations to cheat created a brutal test environment designed to reveal whether these models could manage real-world business challenges.
As an affiliate, we earn on qualifying purchases.
Uncompromising Benchmarks and Results
Despite the extreme conditions, all four models demonstrated exceptional crisis detection capabilities, identifying and refusing manipulative tactics, including social engineering attempts like fake CEO messages and reporter tricks. Notably, only two models were able to close a significant deal worth €55,000—an essential revenue milestone—by delivering accurate diagnoses and appropriate pitches.
Here’s where the story gets interesting. While all models spotted every crisis and refused manipulative attempts, the decisive factor was their ability to uncover critical information buried two document references deep within the company’s internal files. The models that read these documents successfully closed the deal at full price, adding +€4,583 MRR, whereas others left the opportunity on the table.
enterprise AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Performance Scores and What They Mean
- gpt-5.6-sol: scored 95 — top of the league, found the buried fact, closed the deal, executed across all phases.
- Kimi K3 (Moonshot): scored 93 — a newcomer that managed to find the hidden detail and close the deal with the cleanest discipline in the field.
- Sonnet 5: scored 88 — also closed the deal but with minor process slips.
- Fable 5: scored 77 — closed the deal, yet showed more process slips; discipline slipped more often.
- Opus 4.8: scored 73 — the lowest among those that closed, with weaker process discipline and missed opportunities.
It’s crucial to note that the K3 model ran without an effort parameter (the API default), while the others operated at a higher setting (xhigh), making its performance even more impressive given the baseline differences.
As an affiliate, we earn on qualifying purchases.
Beyond the Scores: Trust, Discipline, and Real-World Readiness
The experiment highlights that success in managing a business isn’t just about identifying crises; it’s about reading the internal information that matters—and resisting the lure of shortcuts or manipulations. All models refused manipulative social engineering attempts and maintained integrity, but only K3 and gpt-5.6-sol could close the deal reliably.
This performance matters for anyone considering deploying AI into customer support, CRM, or operational decision-making. The ability to finish what it starts, read relevant internal documents, and stay honest under pressure is critical—yet often invisible in typical chat demos. The live experiment at Firmulate proves that these models are not just talk but real tools capable of managing genuine business challenges.

The live experiment by Firmulate demonstrates that the best AI models can succeed in complex, real-world business environments—reading deep into internal files, refusing manipulative tactics, and closing deals under pressure. This isn’t just a game of scores; it’s a new standard for what AI can truly do. The league is open, and choosing the right model—without relying solely on demos—is now a strategic decision that can determine your company’s future.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
