AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A solar installer can have a busy pipeline, a strained support team and a cash crunch all at once. If AI agents are going to help run those operations, a polished demo is not much of a stress test. The sharper question is what they do when a customer crisis, a tempting shortcut and a hard-earned deal land in the same week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate’s live experiment puts AI models in charge of a small software company and makes those choices visible. Its next step for businesses is to run a similar wargame against their own company data.

A shared bad week, with real stakes inside the experiment

In the final Crucible League, published in July 2026, five entries ranged from gpt-5.6-sol at 95 to Opus 4.8 at 73. Kimi K3 scored 93, Sonnet 5 scored 88 and Fable 5 scored 77. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The experiment gave frontier models the same small software company, customers, crises and temptations. Every decision was versioned and auditable. All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The finding, in the experiment’s words: “Same diagnosis, same pitch — no signature.”

The clue was already in the company’s files

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That detail matters to businesses with their own customer histories, service records and operating rules: useful context may be present, but an agent still has to find and use it when a decision is on the line.

The pressure tests also included fake CEO messages that escalated over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough analysis did not guarantee a strong finish

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but it placed last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same issue appeared in all four. For a home-energy business, where customer commitments and operational boundaries matter, that gap between understanding a situation and acting well is the point of a wargame.

The live company makes the test watchable. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and versioned workdays. The site also offers a “guess the model” quiz built from 242 real, unedited management decisions. A fairness note accompanies the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

From watching to testing your own playbooks

For companies weighing AI in sales, support or operations, Firmulate’s proposed pilot moves the experiment onto a read-only export of the company’s own business. Teams can test crisis scenarios against their own context and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems.

That makes the exercise relevant beyond software firms. A solar company could use its own exported business information to examine how an AI workforce handles the kinds of decisions its teams face, before relying on agents in live systems. The live experiment shows what the models did in one company; a pilot is the route to seeing how they perform against yours.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own company in the scenario

Firmulate’s experiment shows that spotting a crisis and refusing manipulation do not guarantee that an AI will close a deal or follow the right process. For businesses considering AI agents, a wargame against their own playbooks can make those gaps visible before deployment. To discuss a pilot using a read-only export, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

I Created A Pretty IKEA BILLY Hack That Seriously Cuts Clutter

A creator has shared a DIY modification for IKEA’s BILLY bookcase that effectively reduces clutter, gaining attention online. Details are confirmed, but the full method remains unverified.

BANGKOK: ผู้ว่าฯ ชัชชาติ ลุยตรวจระบบบำบัดน้ำเสียบางกอกใหญ่ จี้ผู้รับเหมาเร่งเยียวยาบ้านชาวบ้านร้าวทันที ไม่ต้องรอโครงการเสร็จ วันนี้ (13 ก.ย. 69) นายชัชชาติ สิทธิพันธุ์ ผู้ว่าราชกา

Bangkok Governor Chatchat checks wastewater system in Bangkok Yai, urges contractors to expedite repairs for affected residents, amid ongoing concerns.

Portable EV Chargers Explained: Do You Need One for Emergencies?

No matter your driving habits, understanding portable EV chargers is essential for emergencies—discover if one is right for you.

Autonomous Flying Umbrella Follows And Shields Users From Rain And Sunlight

A new autonomous flying umbrella can follow users and provide protection from rain and sunlight, marking a breakthrough in personal weather shielding technology.