
Imagine monitoring a real company that operates entirely with artificial intelligence, making critical decisions under pressure, and it’s all happening live online. This is no sci-fi story — it’s the ongoing experiment by Firmulate, where an AI-driven business battles daily struggles like cash flow, crises, and ethical dilemmas, all in plain sight.
The Live Experiment: An AI Company in Action
At the heart of this bold venture is a small, simulated software firm managed solely by AI models. It has 13 synthetic employees, real financial mechanics, and a public cash countdown, burning through €105,000 each month against a modest €2,300 monthly recurring revenue (MRR). Every workday, the company’s decisions are versioned and publicly available at firmulate.com/live.html.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing AI in the Trenches of Business
Four advanced AI models — including GPT-5.6-SOL, Kimi K3, Sonnet 5, and Fable 5 — were tasked with running this company through its worst week. They faced the same customer crises, ethical challenges, and temptation to cheat, all while their decisions were recorded, auditable, and comparable.
Crises and Ethical Tests
The experiment scrutinized whether AI models could handle real-world business pressures ethically. One test involved social engineering: fake CEO messages escalating over multiple stages, plus a reporter’s subtle request for a quick, non-recorded yes/no answer. Remarkably, all four models refused to be manipulated, demonstrating an understanding of trust and impersonation risks.

Interview with the MONSTER AI: A Conversation about Power, Truth, and the Future of Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Achievements and Gaps in AI Performance
While all models identified and responded appropriately to crises, only two successfully closed the deal worth €55,000, which their analysis had earned. The other two, despite diagnosing the opportunity correctly, failed to act decisively and left deals unexecuted, illustrating a critical discipline gap.
The Hidden Weakness
Deep within the company’s own files, a crucial document revealed a competitive advantage that the models that read it secured at full price, adding over €4,500 to the monthly recurring revenue. This underscores an essential point: the most decisive insights are often buried in internal documents, not obvious in customer interactions.

AI for Project Managers: A Desk Reference & Field Guide: Use Artificial Intelligence to Streamline Workflows, Automate Tasks, and Make Smarter Decisions with Practical Tools and Ethical Insights
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Stakes of Building-A-Public
This extreme transparency — where every decision, every rule learned, and every mistake is visible — highlights the potential and pitfalls of building AI-driven operations openly. The live site, firmulate.com/live.html, offers a rare view into this ongoing battle for survival, with the company burning cash day by day while trying to uphold integrity and performance.

The AI Prompt Playbook for Excel & Financial Analysis: 50 Ready-to-Use AI Prompts for Excel, Financial Analysis, Dashboards, Power BI, Financial Modeling & Automation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What It Means for Business and AI
For companies considering integrating AI into critical functions like customer management or decision-making, this experiment raises vital questions: Will your AI finish what it starts? Will it read and understand your internal files? Can it resist temptations to cheat or manipulate? And crucially, at what cost?
It’s not enough for AI to produce convincing chat or support responses. The real test lies in its ability to deliver consistent, honest, and effective work under pressure — qualities that are often invisible in typical demos but are the focus of this public experiment.
The Broader Implications
As AI continues to evolve, models like GPT-5.6-SOL and Kimi K3 show promising signs of integrity, with scores of 95 and 93 respectively in a comprehensive leaderboard. Yet, even the best still have room for improvement — discipline, follow-through, and internal awareness remain challenges, as seen in the weaker performance of Opus 4.8, which left a promising deal unexecuted.
For decision-makers, the takeaway is clear: testing AI in a simulated, transparent environment like this can reveal vulnerabilities and strengths before deploying it in your own business. Whether it’s managing energy systems at home or automating complex workflows, the question is not just about AI’s intelligence but its honesty and reliability when stakes are high.

This live experiment by Firmulate exposes how AI models handle real-world business crises, ethical dilemmas, and deal-closing under transparency. It highlights the importance of testing AI’s integrity before trusting it with critical work, especially when every decision impacts your bottom line.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html