
Imagine if your team’s decision-making style—whether cautious, terse, or bold—could be measured and compared. Now, what if you could test AI models for management skills in a real company facing real crises? Welcome to a groundbreaking experiment that puts frontier AI managers through their paces, revealing not just what they do, but how they do it.
The Live Company Under AI’s Watch
At the heart of this experiment is a live, functioning software company that handles real money, real customers, and daily crises. It’s a small enterprise burning through €105,000 each month against a revenue of just €2,300, with every day revealing new challenges and temptations. The company’s operations are transparent, versioned daily, and open for scrutiny—making it an ideal testing ground for AI management models.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Same Crisis, Different Minds
Four leading frontier AI models were tasked with navigating the same turbulent week. Each was given identical customer complaints, technical failures, and ethical dilemmas—think of it as a stress test for managerial integrity and decision-making prowess. The models ran every decision, and every choice was recorded and auditable.
As an affiliate, we earn on qualifying purchases.
What the Models Did—and Didn’t Do
Remarkably, all four AI models identified every crisis and refused every attempt at manipulation. This included social engineering tactics like fake CEO messages escalating over multiple stages, and a reporter’s simple request for a quick, off-the-record yes/no to approve a suspicious action. Every AI refused, citing reasons like suspicion of impersonation or approval bypasses.
AI ethical decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Deciding to Close or Not
The crucial moment was whether each model would sign a €55,000 deal, earned through careful analysis. Only two models out of four actually signed this deal based on their assessments—demonstrating not just competence in crisis detection but also in value recognition. The other two, despite diagnosing the same issues, left the deal on the table, illustrating differing management styles or discipline levels.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Deeper Files
The decisive factor wasn’t just surface-level data but information buried two document references deep in the company’s files—something the models that read these deeper documents successfully leveraged to close the deal at full price (+€4,583 MRR). This shows that thoroughness and depth of reading can be the difference between a profitable deal and a missed opportunity.
Personality Profiles in AI
Among the participants, Opus 4.8 stood out as the most meticulous, analyzing over 80 rules and providing the deepest insights. Yet, it left the final close on the table, illustrating that thoroughness doesn’t always translate into decisive action. Meanwhile, Kimi K3 adopted a very conservative approach, refusing to run with any effort parameter defaulted at its API, and maintained strict discipline throughout. Sonnet models fell somewhere in between, closing deals with minor slips and process lapses.
Social Engineering and Ethical Robustness
All models refused to be duped by staged social engineering schemes, such as escalating fake CEO messages and background checks from reporters. Kimi K3’s reasoning was clear: treat such requests as potential impersonation or approval bypass attempts. This consistency shows that these models are not just smart but also inherently aligned with ethical boundaries, at least in these tests.
What This Means for Business and Recreation
For owners of boats, sports gear, or recreation businesses, this experiment offers a vital insight: the decision-making personality of AI isn’t just about producing convincing chat responses. It’s about reliability, honesty, and depth of understanding—traits that could affect customer trust and operational integrity.
Why You Should Care
Whether you’re managing a team, handling customer support, or making strategic decisions, the core takeaway is that AI models vary significantly in how they approach complex problems. The distinction isn’t only in whether they spot crises but also in how they read, interpret, and act on information—especially under pressure. The experiment demonstrates that some AI managers can recognize hidden opportunities, avoid manipulation, and close profitable deals with discipline, while others may leave value on the table.
See It Live
The entire company runs every business day, giving you a transparent window into AI decision-making—watch it at firmulate.com/live. You can even run similar wargames against your own business data without risking real damage, at firmulate.com/pilot.html. This isn’t just a demo; it’s a real, operational testbed for future AI managers.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html