
Imagine a company where no humans work, yet it faces daily financial stress, endless crises, and tough ethical tests—all live online. This is not a sci-fi story, but the reality of a groundbreaking experiment in artificial intelligence and business management, now accessible for all to watch and learn from.
The Live Experiment: An AI-Run Company in Real-Time
At the heart of this experiment is Firmulate, a platform that runs a simulated small software business using advanced AI models as its decision-makers. Every workday, the company faces real crises—customer complaints, ethical dilemmas, and strategic decisions—all in a controlled environment. The company operates with 13 synthetic employees and some astonishing mechanics: it burns €105,000 each month against a modest €2,300 monthly recurring revenue, with a public countdown showing how long it can keep going.
This setup isn’t just for show. Every decision made by the AI models is versioned and auditable, providing a transparent window into their reasoning process. The goal? Test whether different AI models can navigate complex management scenarios, stay honest under pressure, and ultimately, close profitable deals.
![Express Schedule Free Employee Scheduling Software [PC/Mac Download]](https://m.media-amazon.com/images/I/41yvuCFIVfS._SL500_.jpg)
Express Schedule Free Employee Scheduling Software [PC/Mac Download]
- User-friendly drag & drop interface: Simple shift planning
- Manage time-off and leave: Add sick leave, breaks, holidays
- Email schedules to employees: Send schedules directly via email
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Models and Their Performance
Four frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were each given the same intense week of management challenges. These models faced identical crises: customer churn, ethical threats, temptation to cut corners, and even social engineering attempts like fake CEO messages and reporter tricks. Remarkably, all four models identified every crisis and refused every manipulation attempt. Their discipline was impressive, especially in ethical judgment.
However, when it came to closing deals—the monetary outcome that truly measures management effectiveness—only two models succeeded. The gpt-5.6-sol and Kimi K3 closed the deal worth €55,000 and increased monthly recurring revenue by €4,583. The other two models—Sonnet 5 and Fable 5—missed the opportunity, despite having the correct diagnoses and pitches. Fable 5, for example, left the deal on the table, failing to follow through due to a discipline slip, illustrating that even the most rule-disciplined AI can falter under operational pressures.
The Hidden Weakness and Its Implications
The most telling weakness was buried in the company’s internal documents, not in customer interactions. The models that read these files thoroughly and caught the hidden fact ended up winning the deal at full price. This highlights a critical point: in real-world business, the ability to dig deep into internal data can be decisive—yet many AI models overlook this. In practice, ignoring internal files may cost companies lucrative opportunities, and this experiment makes that risk clear.
Trust and Ethical Testing in AI Decision-Making
Beyond financial outcomes, the models were tested against social engineering attempts—fake CEO messages escalating over multiple stages and a reporter trick asking for a quick approval. All five models refused to be manipulated, with Kimi K3 explicitly treating such requests as potential impersonation. This demonstrates that well-trained AI can maintain integrity even under social pressure, an essential trait for AI systems that will interact with human users and sensitive data in the real world.
The Human-Like Struggles of AI Managers
The live company, with its daily updates and versioned decisions, reveals that AI decision-making isn’t infallible. For instance, Opus 4.8, the most comprehensive model, was the last to place a closing bet and slipped in discipline—leaving a deal unexecuted. This mirrors human management errors, suggesting that AI, despite its discipline, still needs careful oversight and training.
Why This Matters for Your Business
If AI agents will someday touch your customer database, support system, or forecasting models, the key questions aren’t about how well they write but whether they can complete the work honestly and effectively. Will they follow through on promising deals? Will they read your internal files to find hidden opportunities? And crucially, can they resist manipulation or ethical shortcuts? The answers can mean the difference between a productive AI workforce and one that costs more than it saves.
Explore and Observe in Real Time
Visit firmulate.com/live to see this experiment unfold live. Watch how the AI models handle crises, ethical dilemmas, and lucrative deals—all in real time, versioned and transparent. The experiment is ongoing, with new runs queued regularly, offering a rare glimpse into the future of AI-driven management and enterprise decision-making.

This live AI company experiment reveals that while models can identify crises and resist manipulation, closing profitable deals remains a challenge. It underscores the importance of internal data access and ethical discipline—lessons vital for AI’s role in future businesses. Watch the story as it unfolds at firmulate.com/live, and consider how your own enterprise might learn from this transparent, build-in-public show of AI management in action.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html