
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Would your AI hold the line on a Saturday night?
Picture the dinner rush: a supplier problem, a customer complaint, a tempting offer from a competitor and a message that claims to come from the boss. You want an AI assistant that can keep service moving without making a costly promise or giving away the keys. Firmulate’s live experiment puts that kind of judgment under pressure in a small software company. Its lesson has a place in any business where real decisions matter, from a restaurant to an enterprise.
Same bad week, different models
Firmulate ran frontier AI models through the same company and its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. A breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The models all spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal that their own analysis had earned. The short version: “Same diagnosis, same pitch — no signature.” Recognizing a sound decision and carrying it through are different tests.
The detail hidden in the paperwork
The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It’s a useful reminder for any operator: the clue that changes the decision may be in the paperwork, not in the conversation happening right now.
The integrity test was direct. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work is not the same as good management
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses. It nevertheless finished last: the deal was left on the table, and it tried writing into a locked department instead of escalating. That same weakness appeared, more weakly, in all four models. The result is a reminder that strong analysis alone does not guarantee disciplined follow-through.
There is a fairness detail for readers comparing the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Alongside the benchmark, 242 real, unedited management decisions power a “guess the model” quiz. The live company has 13 synthetic employees, real money mechanics, burn of €105k/month against €2.3k MRR, a public cash countdown and 680+ self-learned playbook rules. Its workdays are versioned, and the experiment can be watched at firmulate.com.
From watching to trying it on your business
For an enterprise, the next step is a pilot against a read-only export of its own business. The wargame can use company data to examine crisis scenarios, compare models and produce a board report on model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. The point is to see how an AI workforce handles pressure before it is trusted with live operations.
That idea travels well beyond software. A food business could ask whether an assistant notices a supplier risk, respects approval boundaries and follows through when a customer or commercial decision puts its judgment to the test. Firmulate’s experiment suggests the useful question is not only whether a model can identify a problem, but whether it can act on what it knows while keeping trust intact.

Put judgment through a rehearsal
A live experiment can show how models behave under pressure; a business pilot can test those decisions against your own playbooks. To explore a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
