AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would your AI hold the line on a Saturday night?

Picture the dinner rush: a supplier problem, a customer complaint, a tempting offer from a competitor and a message that claims to come from the boss. You want an AI assistant that can keep service moving without making a costly promise or giving away the keys. Firmulate’s live experiment puts that kind of judgment under pressure in a small software company. Its lesson has a place in any business where real decisions matter, from a restaurant to an enterprise.

Same bad week, different models

Firmulate ran frontier AI models through the same company and its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. A breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The models all spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal that their own analysis had earned. The short version: “Same diagnosis, same pitch — no signature.” Recognizing a sound decision and carrying it through are different tests.

The detail hidden in the paperwork

The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It’s a useful reminder for any operator: the clue that changes the decision may be in the paperwork, not in the conversation happening right now.

The integrity test was direct. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work is not the same as good management

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses. It nevertheless finished last: the deal was left on the table, and it tried writing into a locked department instead of escalating. That same weakness appeared, more weakly, in all four models. The result is a reminder that strong analysis alone does not guarantee disciplined follow-through.

There is a fairness detail for readers comparing the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Alongside the benchmark, 242 real, unedited management decisions power a “guess the model” quiz. The live company has 13 synthetic employees, real money mechanics, burn of €105k/month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Its workdays are versioned, and the experiment can be watched at firmulate.com.

From watching to trying it on your business

For an enterprise, the next step is a pilot against a read-only export of its own business. The wargame can use company data to examine crisis scenarios, compare models and produce a board report on model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. The point is to see how an AI workforce handles pressure before it is trusted with live operations.

That idea travels well beyond software. A food business could ask whether an assistant notices a supplier risk, respects approval boundaries and follows through when a customer or commercial decision puts its judgment to the test. Firmulate’s experiment suggests the useful question is not only whether a model can identify a problem, but whether it can act on what it knows while keeping trust intact.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put judgment through a rehearsal

A live experiment can show how models behave under pressure; a business pilot can test those decisions against your own playbooks. To explore a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Management Skills Are the Hidden Key in Business Resilience — Not Just Chat Quality

AI models excel at chat, but their true test is managing crises, reading internal data, and staying honest under pressure. Firmulate’s live experiment reveals management as the key to trustworthy AI.

Behind the Curtain: How AI Is Running a Live Company Into the Ground — and You Can Watch It Live

A real, live company run by AI models faces crises, temptations, and tough decisions daily — revealing critical insights into AI’s true management capabilities and limits.

Voice‑Assisted Drive‑Thru Systems

Many drive-thru systems now leverage voice assistance to enhance accuracy and speed—discover how this technology can transform your business.

Why Even a ‘Do-Nothing’ AI Gets 26 Points — and Why That Matters for Business Trust

The Crucible League’s AI benchmarks show even a do-nothing model scores 26 points, emphasizing the importance of trust, discipline, and honesty in AI-driven business decisions.