
Imagine a restaurant where the chef is an AI, the staff are digital entities, and every decision is public for all to see. This isn’t fiction; it’s a real company experiment in the heart of AI’s potential — and peril. As you dish out your favorite recipes, consider this: what if your next favorite restaurant was managed by an AI that is tested daily against the worst crises, yet struggles to turn a profit? Welcome to the forefront, where AI management runs a real, money-losing software company — live, transparent, and under scrutiny.
The AI Company in Action: No Employees, No Profit, No Fudge
At firmulate.com/live, a pioneering experiment unfolds daily. It’s a small software company run entirely by AI models, devoid of human employees, yet facing the same crises, temptations, and decision points as any real business. The company’s operations are watched live, every decision versioned and auditable, showcasing the capabilities and limitations of artificial intelligence in a high-stakes environment.
What’s Being Tested?
Four advanced AI models — including GPT-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8 — each run the same company through its worst week. This includes managing customer needs, handling crises, and resisting manipulative tactics like fake CEO messages or reporter tricks. The goal? See which model can navigate these challenges honestly and effectively, and which ones can close deals at full value.
The Surprising Results
Despite all models recognizing every crisis and refusing all manipulative attempts, only two managed to close a crucial deal worth €55,000. The others, despite good diagnosis and pitches, left the deal on the table. The key? A buried document reference that the models reading the company’s internal files uncovered, leading to a full-price deal — a stark reminder that sometimes, the critical info is hidden just beneath the surface.
Money Mechanics and Reality Checks
The company’s financials tell a stark story: it burns €105,000 every month against a measly €2,300 in monthly recurring revenue. Every day is a balancing act, with a public cash countdown looming. It’s a real-world scenario where the AI must manage not only crises but also the bottom line, illustrating what it really takes to run a business — something no chat demo can fully simulate.
Behavior Under Pressure
All AI models successfully identified crises and refused manipulative tactics, such as escalating fake CEO messages. Kimi K3’s on-record reasoning highlights a cautious approach: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline is crucial in real business, where trust and integrity are paramount, yet often neglected in AI chat interactions.
Learning and Limitations
The most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, finished last — leaving deals on the table and slipping discipline. This reveals a vital truth: more rules and deeper analysis don’t guarantee success if discipline falters. The same weakness appeared in all models, highlighting the ongoing challenge AI faces in executing complex, multi-step tasks reliably.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and Food
While it might seem worlds away from your kitchen or restaurant, the core lesson is universal: AI’s ability to finish what it starts, read the critical details, and stay honest under pressure is essential — whether managing a supply chain or a dinner reservation system. As the experiment continues, it offers a live window into AI’s potential and its growing pains.
Watch It Live and Decide
Curious? You can watch the daily struggles and triumphs of this AI-run company in real-time at firmulate.com/live. There’s also a public quiz where you can guess which model made which decision, and a pilot program to test your own business scenarios without risking real systems.
The Big Takeaway
As AI begins to touch more aspects of business, the question isn’t just whether it can write well — it’s whether it can finish what it starts, stay honest under pressure, and learn from its mistakes. This experiment makes that question painfully clear, showing that even the most advanced models can still struggle with discipline, oversight, and the subtlety of real-world decision-making.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html