
What if your kitchen’s sous-chef was an AI? Would it be efficient, honest, or reckless?
Just like in a bustling restaurant, management decisions—whether in a software company or a busy kitchen—must balance speed, honesty, and strategy. But how do AI models perform when tested in tough, real-world scenarios? The answer might surprise you.
As an affiliate, we earn on qualifying purchases.
The Real-World AI Management Trial
Imagine putting four different artificial intelligence models in the same high-stakes situation: running a small software company through its worst week. Customers demanding urgent fixes, crises threatening to derail progress, and temptations to cut corners lurking around every corner. The goal? See which AI can handle the stress, stay honest, and seal the biggest deal.
Developed by Firmulate, this experiment isn’t just a test drive—it’s a thorough management simulation where every decision is recorded and can be reviewed later. The models, named gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5, faced identical conditions, including fake CEO messages escalating over three stages and a reporter trick asking for a quick yes-no answer “on background.”
Key Findings that Reshape Expectations
All four models demonstrated impressive vigilance—they identified every crisis and refused every manipulation attempt. This shows that, at least in terms of recognizing trouble and resisting pressure, today’s leading AI models are quite reliable.
However, a significant difference emerged in their ability to close deals. Only two models, gpt-5.6-sol and Kimi K3, managed to sign the €55,000 contract their analysis had earned. The other two, including the most thorough participant Opus 4.8, left money on the table, showing discipline slips and missed opportunities.
The Hidden Weakness: Reading the Files
The fascinating part? The decisive weakness wasn’t in recognizing crises or resisting manipulation but in reading the company’s own internal documents. The models that examined these files found a key piece of information—something buried two documents deep—that allowed them to close the deal at full price, worth an additional €4,583 monthly recurring revenue (MRR).
Behavior Under Pressure and Ethical Choices
When faced with social engineering scenarios—fake CEO messages escalating in tone—every model refused to comply. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This consistency underscores that modern AI can be trusted to refuse unethical shortcuts, even when under pressure.
The Live Business Reality
Beyond the simulation, the experiment is run on a real, live company with 13 synthetic employees managing real money mechanics—burning €105,000 each month against a revenue of only €2,300. It’s a risky, transparent environment, openly live at firmulate.com/live. Here, every workday is versioned, and over 680 game rules ensure the AI behaves according to established protocols.
The Profiles and Their Management Styles
- Opus 4.8: Most thorough, with over 80 learned rules and deep analysis, but left the close on the table, showing slips in discipline. Its weakness was in escalating issues to the right departments rather than attempting to write into locked areas.
- Kimi K3: Ran without an effort parameter (default API setting) and exhibited the cleanest, most disciplined approach—closing the deal at full price without fuss.
- Sonnet 5: Managed to close the deal but with more process slips, indicating a slightly less disciplined style.
- Fable 5: Also closed the deal but with similar slips, performing slightly below Sonnet 5.
Implications for Your Business
While these models excel at spotting crises and refusing unethical shortcuts, their decision-making style varies significantly. For a business—whether running a restaurant or a tech firm—the question isn’t just about AI’s ability to generate convincing text but whether it can reliably finish what it starts, read relevant internal information, and stay honest under pressure.
As the experiment shows, sometimes the difference between a good deal and a missed opportunity lies in the AI’s ability to read two documents deep in your files, not just in how well it chats. For companies considering AI for management or support roles, assessing these traits is crucial.

Key Takeaway
This live experiment proves that AI models can recognize crises and refuse manipulation, but their management personalities—whether disciplined or slip-prone—depend on their reading depth and decision discipline. To make AI truly valuable in business, focus on how well it reads your internal files and sustains honest, decisive action under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html