AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A solar installer can have a busy pipeline, a strained support team and a cash crunch all at once. If AI agents are going to help run those operations, a polished demo is not much of a stress test. The sharper question is what they do when a customer crisis, a tempting shortcut and a hard-earned deal land in the same week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate’s live experiment puts AI models in charge of a small software company and makes those choices visible. Its next step for businesses is to run a similar wargame against their own company data.

A shared bad week, with real stakes inside the experiment

In the final Crucible League, published in July 2026, five entries ranged from gpt-5.6-sol at 95 to Opus 4.8 at 73. Kimi K3 scored 93, Sonnet 5 scored 88 and Fable 5 scored 77. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The experiment gave frontier models the same small software company, customers, crises and temptations. Every decision was versioned and auditable. All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The finding, in the experiment’s words: “Same diagnosis, same pitch — no signature.”

The clue was already in the company’s files

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That detail matters to businesses with their own customer histories, service records and operating rules: useful context may be present, but an agent still has to find and use it when a decision is on the line.

The pressure tests also included fake CEO messages that escalated over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough analysis did not guarantee a strong finish

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but it placed last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same issue appeared in all four. For a home-energy business, where customer commitments and operational boundaries matter, that gap between understanding a situation and acting well is the point of a wargame.

The live company makes the test watchable. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and versioned workdays. The site also offers a “guess the model” quiz built from 242 real, unedited management decisions. A fairness note accompanies the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

From watching to testing your own playbooks

For companies weighing AI in sales, support or operations, Firmulate’s proposed pilot moves the experiment onto a read-only export of the company’s own business. Teams can test crisis scenarios against their own context and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems.

That makes the exercise relevant beyond software firms. A solar company could use its own exported business information to examine how an AI workforce handles the kinds of decisions its teams face, before relying on agents in live systems. The live experiment shows what the models did in one company; a pilot is the route to seeing how they perform against yours.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own company in the scenario

Firmulate’s experiment shows that spotting a crisis and refusing manipulation do not guarantee that an AI will close a deal or follow the right process. For businesses considering AI agents, a wargame against their own playbooks can make those gaps visible before deployment. To discuss a pilot using a read-only export, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Blackstone Digital Infrastructure Surges In Global Coverage

Coverage of Blackstone’s digital infrastructure investments has surged globally, with 21 mentions in recent media tracking, signaling rising interest in this sector.

The Solar Quote That Depends on Reading Page 47 of the Spec Sheet — and What It Says About AI Agents

Four frontier AIs ran the same company’s worst week. All diagnosed the problem — only those who read two files deep closed the €55,000 deal.

Microsoft Comic Chat is now open source

Microsoft has released Comic Chat as open source, allowing developers to access, modify, and integrate the chat client into projects. The move aims to revive interest in the software.

U.S. Solar System Pricing Rises For Utility And Commercial Projects As Residential Costs Decline

U.S. solar system pricing is increasing for utility and commercial projects, while residential costs continue to decline, signaling shifting market dynamics.