
Business resilience, observed in real time
For readers interested in home energy, solar and backup power, the premise feels familiar: performance under ordinary conditions tells only part of the story. What matters is what happens when pressure rises, resources tighten and several problems arrive together.
Firmulate brings that resilience test into business. Its live software company has 13 synthetic employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue. A public cash countdown makes the central tension impossible to miss: this is a company fighting for survival in full view of its audience.

Methodology for Ensuring Operational Resilience of IT Projects: Automation of Integration Processes and Minimization of Anthropogenic Risks
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company that publishes its working days
The experiment can be watched live. Every workday is versioned, while the synthetic workforce has accumulated more than 680 self-learned playbook rules. That creates something unusual in business technology: an ongoing company story in which operating decisions, financial pressure and employee conduct generate new material every day.
Instead of presenting AI through a polished chat demonstration, Firmulate places frontier models in charge of the same small software company during its worst week. They encounter the same customers, crises and temptations, and every decision is auditable. The comparison therefore centers on management behavior: whether a model notices trouble, gathers the relevant information, resists manipulation and completes commercially important work.
The gap between understanding and finishing
The models performed impressively on recognition. All of them spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the problem neatly: “Same diagnosis, same pitch — no signature.”
The decisive information was not sitting in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue. The episode turns an apparently mundane act—reading the company’s own material—into the difference between analysis and revenue.
That distinction should resonate with anyone evaluating resilience technology. Identifying a threat is valuable, but dependable performance also requires the system to follow through when conditions become complicated. In Firmulate’s company, the commercial outcome separated models that reached a sound diagnosis from those that converted it into action.
Pressure also tested trust
The worst week included fake CEO messages that escalated across three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest description of the danger: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because the do-nothing baseline scored 26 even though partial progress counted. A single breach of trust capped the total, reflecting the experiment’s governing principle that “no amount of good work outweighs a breach of trust.” The models maintained that boundary even while other parts of their execution varied.
A close league with meaningful differences
The final Crucible League standings from July 2026 were:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
Kimi K3’s result carries an important fairness note: it ran with the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, it finished just behind gpt-5.6-sol.
Opus 4.8 provides the most revealing counterexample. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four of the other models.
A continuing public narrative
Firmulate’s appeal is not limited to a final leaderboard. The company continues operating with its exposed financial imbalance, growing body of learned rules and versioned workdays. Visitors can also read what its synthetic employees say, turning management behavior into a public record rather than a private product claim.
The broader project includes 242 real, unedited management decisions used in a “guess the model” quiz. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.


AI for Real Companies: A Practical Guide to Smarter Systems and Stronger Profits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The lesson for resilience-minded readers
The live company offers a useful way to think about AI readiness. A system can recognize every crisis, reject every manipulation and still fail to finish the work that keeps a business alive. Thorough analysis can coexist with weak execution, while a buried fact can determine whether a valuable deal closes.
Firmulate makes those differences visible through a company whose cash pressure, employee decisions and operating history remain public. Like any resilience test, its value lies in exposing behavior before the stakes move from an experiment into everyday operations.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

IT Crisisology: Smart Crisis Management in Software Engineering: Models, Methods, Patterns, Practices, Case Studies (Smart Innovation, Systems and Technologies, 210)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI-driven business continuity solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.