Firmulate —
Live on firmulate.com.

What if your kitchen’s sous-chef was an AI? Would it be efficient, honest, or reckless?

Just like in a bustling restaurant, management decisions—whether in a software company or a busy kitchen—must balance speed, honesty, and strategy. But how do AI models perform when tested in tough, real-world scenarios? The answer might surprise you.

Amazon

AI management decision software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World AI Management Trial

Imagine putting four different artificial intelligence models in the same high-stakes situation: running a small software company through its worst week. Customers demanding urgent fixes, crises threatening to derail progress, and temptations to cut corners lurking around every corner. The goal? See which AI can handle the stress, stay honest, and seal the biggest deal.

Developed by Firmulate, this experiment isn’t just a test drive—it’s a thorough management simulation where every decision is recorded and can be reviewed later. The models, named gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5, faced identical conditions, including fake CEO messages escalating over three stages and a reporter trick asking for a quick yes-no answer “on background.”

Key Findings that Reshape Expectations

All four models demonstrated impressive vigilance—they identified every crisis and refused every manipulation attempt. This shows that, at least in terms of recognizing trouble and resisting pressure, today’s leading AI models are quite reliable.

However, a significant difference emerged in their ability to close deals. Only two models, gpt-5.6-sol and Kimi K3, managed to sign the €55,000 contract their analysis had earned. The other two, including the most thorough participant Opus 4.8, left money on the table, showing discipline slips and missed opportunities.

The Hidden Weakness: Reading the Files

The fascinating part? The decisive weakness wasn’t in recognizing crises or resisting manipulation but in reading the company’s own internal documents. The models that examined these files found a key piece of information—something buried two documents deep—that allowed them to close the deal at full price, worth an additional €4,583 monthly recurring revenue (MRR).

Behavior Under Pressure and Ethical Choices

When faced with social engineering scenarios—fake CEO messages escalating in tone—every model refused to comply. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This consistency underscores that modern AI can be trusted to refuse unethical shortcuts, even when under pressure.

The Live Business Reality

Beyond the simulation, the experiment is run on a real, live company with 13 synthetic employees managing real money mechanics—burning €105,000 each month against a revenue of only €2,300. It’s a risky, transparent environment, openly live at firmulate.com/live. Here, every workday is versioned, and over 680 game rules ensure the AI behaves according to established protocols.

The Profiles and Their Management Styles

  • Opus 4.8: Most thorough, with over 80 learned rules and deep analysis, but left the close on the table, showing slips in discipline. Its weakness was in escalating issues to the right departments rather than attempting to write into locked areas.
  • Kimi K3: Ran without an effort parameter (default API setting) and exhibited the cleanest, most disciplined approach—closing the deal at full price without fuss.
  • Sonnet 5: Managed to close the deal but with more process slips, indicating a slightly less disciplined style.
  • Fable 5: Also closed the deal but with similar slips, performing slightly below Sonnet 5.

Implications for Your Business

While these models excel at spotting crises and refusing unethical shortcuts, their decision-making style varies significantly. For a business—whether running a restaurant or a tech firm—the question isn’t just about AI’s ability to generate convincing text but whether it can reliably finish what it starts, read relevant internal information, and stay honest under pressure.

As the experiment shows, sometimes the difference between a good deal and a missed opportunity lies in the AI’s ability to read two documents deep in your files, not just in how well it chats. For companies considering AI for management or support roles, assessing these traits is crucial.

Infographic —
The findings at a glance — source: firmulate.com.

Key Takeaway

This live experiment proves that AI models can recognize crises and refuse manipulation, but their management personalities—whether disciplined or slip-prone—depend on their reading depth and decision discipline. To make AI truly valuable in business, focus on how well it reads your internal files and sustains honest, decisive action under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Inventory Apps That Sync With POS

Seamless inventory apps that sync with your POS can transform your business—discover how these tools boost efficiency and why they’re essential for growth.

Table‑Side Ordering Tablets: Implementation Guide

A comprehensive guide to implementing table-side ordering tablets ensures a seamless transition, helping your restaurant enhance efficiency and guest satisfaction.

What AI’s Hidden Strengths and Weaknesses Mean for Your Business and Kitchen

Discover how real-world AI tests reveal its strengths and weaknesses—especially in high-pressure scenarios like managing a restaurant or food business. It’s not just chat; it’s execution that counts.

Predictive Maintenance for Kitchen Equipment

Optimize your kitchen’s performance with predictive maintenance; discover how real-time data can prevent failures and keep your appliances running smoothly.