firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

On a boat, a polished briefing is no substitute for sound decisions when the weather turns. Businesses considering AI face a similar test: can an AI workforce handle a crisis, resist pressure and follow through when the stakes are real? Firmulate’s live experiment puts models through that kind of week inside a synthetic company.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get sports gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate runs AI models as managers of a small software company, giving each the same customers, crises and temptations. The live company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its workdays are versioned, and its playbook has accumulated more than 680 learned rules.

The experiment is watchable at Firmulate. The point is not whether a model can produce a convincing answer in a chat window. It is whether its decisions hold up across a difficult stretch of company life.

Good diagnosis did not guarantee a close

In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The league gives partial progress credit, while a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s shorthand captures the gap: “Same diagnosis, same pitch — no signature.” In a business, recognizing the right move and actually carrying it through are different tests.

The deal depended on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The evidence was available; finding and using it made the difference.

Pressure, trust and follow-through

The social-engineering test escalated through three fake CEO messages, then added a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Refusing manipulation was not the only measure of discipline. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. It left the deal on the table and tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions.

From watching to a company’s own test

For a business, the next step is to test against its own circumstances. Firmulate’s proposed enterprise pilot uses a read-only export to create a digital twin, then runs crisis scenarios against that company’s information and playbooks. The resulting board report can show a model ranking and weak points in the company’s own responses. Nothing writes back to real systems.

That makes the idea relevant well beyond software companies. A boating retailer, marina or supplier could use a controlled wargame to examine how an AI workforce might respond to disruption and competing demands before it is trusted with real operations. The experiment offers a public demonstration; a pilot brings the test to a company’s own data.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Take the test to your business

Firmulate’s live company shows how models behave under shared pressure. An enterprise pilot can put the same kind of wargame against a read-only export of your business and report where models succeed and where playbooks need work. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Security Test Proves Resilience Against Social Engineering — All Models Refuse Fake CEO Commands

AI models tested with staged social engineering attacks showed remarkable resilience, refusing manipulation and even closing legitimate deals—proof that integrity can be tested before deployment.

Diamondbacks-Cardinals series finale at Busch Stadium delayed due to weather

The final game of the Diamondbacks-Cardinals series at Busch Stadium has been postponed due to severe weather conditions, with rescheduling details still pending.

Pistons news and rumors: Tyler Herro to Bucks in huge Giannis trade

Rumors suggest Tyler Herro may be traded from Pistons to Bucks in a blockbuster deal involving Giannis Antetokounmpo. Details are still developing.

Dock Cleat Placement: Why Angle Matters

Optimize your dock cleat placement by understanding the importance of angle—learn how proper positioning can prevent wear and ensure secure mooring.