AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would your AI hold the line on a Saturday night?

Picture the dinner rush: a supplier problem, a customer complaint, a tempting offer from a competitor and a message that claims to come from the boss. You want an AI assistant that can keep service moving without making a costly promise or giving away the keys. Firmulate’s live experiment puts that kind of judgment under pressure in a small software company. Its lesson has a place in any business where real decisions matter, from a restaurant to an enterprise.

Same bad week, different models

Firmulate ran frontier AI models through the same company and its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. A breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The models all spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal that their own analysis had earned. The short version: “Same diagnosis, same pitch — no signature.” Recognizing a sound decision and carrying it through are different tests.

The detail hidden in the paperwork

The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It’s a useful reminder for any operator: the clue that changes the decision may be in the paperwork, not in the conversation happening right now.

The integrity test was direct. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work is not the same as good management

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses. It nevertheless finished last: the deal was left on the table, and it tried writing into a locked department instead of escalating. That same weakness appeared, more weakly, in all four models. The result is a reminder that strong analysis alone does not guarantee disciplined follow-through.

There is a fairness detail for readers comparing the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Alongside the benchmark, 242 real, unedited management decisions power a “guess the model” quiz. The live company has 13 synthetic employees, real money mechanics, burn of €105k/month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Its workdays are versioned, and the experiment can be watched at firmulate.com.

From watching to trying it on your business

For an enterprise, the next step is a pilot against a read-only export of its own business. The wargame can use company data to examine crisis scenarios, compare models and produce a board report on model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. The point is to see how an AI workforce handles pressure before it is trusted with live operations.

That idea travels well beyond software. A food business could ask whether an assistant notices a supplier risk, respects approval boundaries and follows through when a customer or commercial decision puts its judgment to the test. Firmulate’s experiment suggests the useful question is not only whether a model can identify a problem, but whether it can act on what it knows while keeping trust intact.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put judgment through a rehearsal

A live experiment can show how models behave under pressure; a business pilot can test those decisions against your own playbooks. To explore a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Blockchain for Supply‑Chain Transparency

Harness blockchain to boost supply-chain transparency and discover how this innovative technology can transform your operations.

Predictive Maintenance for Kitchen Equipment

Optimize your kitchen’s performance with predictive maintenance; discover how real-time data can prevent failures and keep your appliances running smoothly.

Digital Menu Boards: ROI Analysis

Maximize your investment with digital menu boards by understanding their impressive ROI and discovering how they can transform your business.