
A solar installer can have a busy pipeline, a strained support team and a cash crunch all at once. If AI agents are going to help run those operations, a polished demo is not much of a stress test. The sharper question is what they do when a customer crisis, a tempting shortcut and a hard-earned deal land in the same week.
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate’s live experiment puts AI models in charge of a small software company and makes those choices visible. Its next step for businesses is to run a similar wargame against their own company data.
A shared bad week, with real stakes inside the experiment
In the final Crucible League, published in July 2026, five entries ranged from gpt-5.6-sol at 95 to Opus 4.8 at 73. Kimi K3 scored 93, Sonnet 5 scored 88 and Fable 5 scored 77. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The experiment gave frontier models the same small software company, customers, crises and temptations. Every decision was versioned and auditable. All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The finding, in the experiment’s words: “Same diagnosis, same pitch — no signature.”
The clue was already in the company’s files
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That detail matters to businesses with their own customer histories, service records and operating rules: useful context may be present, but an agent still has to find and use it when a decision is on the line.
The pressure tests also included fake CEO messages that escalated over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough analysis did not guarantee a strong finish
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but it placed last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same issue appeared in all four. For a home-energy business, where customer commitments and operational boundaries matter, that gap between understanding a situation and acting well is the point of a wargame.
The live company makes the test watchable. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and versioned workdays. The site also offers a “guess the model” quiz built from 242 real, unedited management decisions. A fairness note accompanies the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
From watching to testing your own playbooks
For companies weighing AI in sales, support or operations, Firmulate’s proposed pilot moves the experiment onto a read-only export of the company’s own business. Teams can test crisis scenarios against their own context and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems.
That makes the exercise relevant beyond software firms. A solar company could use its own exported business information to examine how an AI workforce handles the kinds of decisions its teams face, before relying on agents in live systems. The live experiment shows what the models did in one company; a pilot is the route to seeing how they perform against yours.

Put your own company in the scenario
Firmulate’s experiment shows that spotting a crisis and refusing manipulation do not guarantee that an AI will close a deal or follow the right process. For businesses considering AI agents, a wargame against their own playbooks can make those gaps visible before deployment. To discuss a pilot using a read-only export, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
