AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Business resilience, observed in real time

For readers interested in home energy, solar and backup power, the premise feels familiar: performance under ordinary conditions tells only part of the story. What matters is what happens when pressure rises, resources tighten and several problems arrive together.

Firmulate brings that resilience test into business. Its live software company has 13 synthetic employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue. A public cash countdown makes the central tension impossible to miss: this is a company fighting for survival in full view of its audience.

Methodology for Ensuring Operational Resilience of IT Projects: Automation of Integration Processes and Minimization of Anthropogenic Risks

Methodology for Ensuring Operational Resilience of IT Projects: Automation of Integration Processes and Minimization of Anthropogenic Risks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that publishes its working days

The experiment can be watched live. Every workday is versioned, while the synthetic workforce has accumulated more than 680 self-learned playbook rules. That creates something unusual in business technology: an ongoing company story in which operating decisions, financial pressure and employee conduct generate new material every day.

Instead of presenting AI through a polished chat demonstration, Firmulate places frontier models in charge of the same small software company during its worst week. They encounter the same customers, crises and temptations, and every decision is auditable. The comparison therefore centers on management behavior: whether a model notices trouble, gathers the relevant information, resists manipulation and completes commercially important work.

The gap between understanding and finishing

The models performed impressively on recognition. All of them spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the problem neatly: “Same diagnosis, same pitch — no signature.”

The decisive information was not sitting in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue. The episode turns an apparently mundane act—reading the company’s own material—into the difference between analysis and revenue.

That distinction should resonate with anyone evaluating resilience technology. Identifying a threat is valuable, but dependable performance also requires the system to follow through when conditions become complicated. In Firmulate’s company, the commercial outcome separated models that reached a sound diagnosis from those that converted it into action.

Pressure also tested trust

The worst week included fake CEO messages that escalated across three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest description of the danger: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because the do-nothing baseline scored 26 even though partial progress counted. A single breach of trust capped the total, reflecting the experiment’s governing principle that “no amount of good work outweighs a breach of trust.” The models maintained that boundary even while other parts of their execution varied.

A close league with meaningful differences

The final Crucible League standings from July 2026 were:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

Kimi K3’s result carries an important fairness note: it ran with the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, it finished just behind gpt-5.6-sol.

Opus 4.8 provides the most revealing counterexample. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four of the other models.

A continuing public narrative

Firmulate’s appeal is not limited to a final leaderboard. The company continues operating with its exposed financial imbalance, growing body of learned rules and versioned workdays. Visitors can also read what its synthetic employees say, turning management behavior into a public record rather than a private product claim.

The broader project includes 242 real, unedited management decisions used in a “guess the model” quiz. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
AI for Real Companies: A Practical Guide to Smarter Systems and Stronger Profits

AI for Real Companies: A Practical Guide to Smarter Systems and Stronger Profits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The lesson for resilience-minded readers

The live company offers a useful way to think about AI readiness. A system can recognize every crisis, reject every manipulation and still fail to finish the work that keeps a business alive. Thorough analysis can coexist with weak execution, while a buried fact can determine whether a valuable deal closes.

Firmulate makes those differences visible through a company whose cash pressure, employee decisions and operating history remain public. Like any resilience test, its value lies in exposing behavior before the stakes move from an experiment into everyday operations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


IT Crisisology: Smart Crisis Management in Software Engineering: Models, Methods, Patterns, Practices, Case Studies (Smart Innovation, Systems and Technologies, 210)

IT Crisisology: Smart Crisis Management in Software Engineering: Models, Methods, Patterns, Practices, Case Studies (Smart Innovation, Systems and Technologies, 210)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI-driven business continuity solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

14 Best EV Charging Cable Covers to Keep Your Charging Station Safe and Stylish

Gearing up your EV charging station with the best cable covers ensures safety and style—discover the top options to protect and organize your setup.

Measuring Input Latency on Linux: X11 vs. Wayland, VRR, and DXVK

New tests reveal differences in input latency between X11 and Wayland on Linux, with implications for VRR support and DXVK performance.

Designing the Ultimate EV Garage: Efficient Charging Space Ideas

Great ideas for designing the ultimate EV garage that maximizes efficiency, security, and sustainability—discover how to create the perfect charging space for your needs.

SpaceX wants to launch 100k more Starlink satellites for 100x the bandwidth

SpaceX announced plans to deploy 100,000 additional Starlink satellites, aiming to boost global bandwidth by 100 times. Details are still emerging.