
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Following a recipe is not the same as getting dinner on the table
A cook can spot what is wrong with a dish, explain how to fix it, and still never serve it. That gap between knowing and doing matters just as much when an AI is asked to run a business. In Firmulate’s live company experiment, Moonshot’s Kimi K3 scored 93, finishing second behind gpt-5.6-sol at 95 and ahead of three Western models. The result is a reminder that polished answers alone do not tell you which AI will finish the job.
One company, one difficult week
Firmulate gave each frontier model the same small software company to run through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The live experiment is watchable at Firmulate; its published benchmark results are available on the benchmark page.
The final league table put gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The company’s refrain captures the gap: “Same diagnosis, same pitch — no signature.”
The detail hidden in the files
The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read those files won the deal at full price, worth +€4,583 MRR. K3 found the buried fact and closed. It also saved the churning customer and resisted all three baits, with one deviation—the cleanest discipline in the field.
Those baits included fake CEO messages escalating across three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work is not the same as a complete result
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. That same weakness appeared, less strongly, in all four models. A company can be busy and still miss the decision that matters.
The simulated firm has 13 employees, burns €105k a month against €2.3k MRR, and publishes a cash countdown. It has accumulated 680+ self-learned playbook rules, and every workday is versioned. These are real money mechanics inside a synthetic company, not a claim that the models ran a real-world business.
The result also comes with a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate offers a quiz built from 242 real, unedited management decisions to guess which model made each choice.

Test the model against your own work
For companies considering AI agents in customer support, sales or forecasting, the question is whether a model reads the relevant records, stays honest under pressure and completes the task. Firmulate says enterprises can run the same wargame against a read-only export of their own business; it does not write back to real systems. A strong benchmark result is useful, but choosing a model without testing it on your own work is still a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
