AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Would you trust the smartest answer—or the decision that actually keeps the lights on?

Home-energy readers already know that impressive specifications do not guarantee dependable performance. A solar array, battery or backup system matters most when conditions become difficult and the household needs it to finish the job. The same distinction is emerging in business AI: sounding capable is not the same as acting reliably under pressure.

Firmulate turns that distinction into a live, watchable experiment. Frontier AI models were each asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Their decisions were versioned and auditable. Now, 242 real, unedited management decisions have become a guess-the-model quiz, inviting readers to identify an AI from what it actually chose to do.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The management personalities hiding behind the prose

The quiz works because the models do not merely produce different writing styles. Across repeated business situations, they display recognizable habits: how deeply they investigate, whether they act on their own analysis, how they respond to suspicious requests and whether they persist when a routine path is blocked.

Those differences produced a close but revealing Crucible League table in July 2026. GPT-5.6-sol finished with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a firm trust boundary: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The striking result was not crisis detection. Every model spotted every crisis, and every model refused every manipulation attempt. The decisive gap appeared between recognizing what should happen and completing it. Only two models signed the €55,000 deal that their own analysis had earned. As Firmulate summarizes the contrast: “Same diagnosis, same pitch — no signature.”

The fact that changed the commercial outcome

The winning information was not presented in the customer event itself. It was buried two document references deep in the company’s own files: a competitor weakness that supported the full-price offer. Models that read the relevant file won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

That detail should resonate with anyone evaluating AI for an operational environment. A convincing response to the latest alert is useful, but real work often depends on connecting the alert with warranty terms, service history, customer commitments or another record sitting elsewhere. In Firmulate’s test, the models faced the same immediate situation. The difference was whether they looked beyond it and then converted what they found into action.

Pressure tested without surrendering trust

The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

This is an important counterweight to the league-table drama. The experiment did not find a model succeeding by becoming more reckless. The strongest results combined investigation, execution and resistance to manipulation. Trust was not treated as a decorative value added after performance; it was part of the definition of performance.

Why thoroughness alone did not win

Opus 4.8 offers the clearest character study. It was the most thorough participant, learning 80 additional playbook rules and producing the deepest analyses, yet it finished last in the league. The commercial close was left on the table, while discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in a milder form across the other four models.

That makes its result more interesting than a simple failure. Opus demonstrated extensive learning and analysis, but those strengths did not automatically translate into the best management outcome. The wargame therefore separates qualities that ordinary chat comparisons often blend together: knowing, explaining, escalating and finishing.

There is also a fairness qualification around the runner-up. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That difference belongs beside the result because comparisons are only useful when their operating conditions remain visible.

A company designed to make consequences visible

The simulated organization contains 13 synthetic employees and uses real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned.

The result is closer to an operational stress test than a polished demonstration. Decisions have consequences, inactivity has a measurable baseline and apparently small process failures can prevent a commercially valuable outcome. Like a backup-power test conducted under load, it exposes behavior that a showroom conversation cannot.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

business AI risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Try to identify the operator, not the writing style

The quiz’s appeal comes from withholding the brand name until after the decision. Readers must judge the behavior itself: Did the model inspect the company’s records? Did it follow through? Did it protect trust when pressured? Did it escalate when its path was blocked?

For businesses considering AI agents, that is the practical lesson. The meaningful question is not simply whether a model can describe a good decision. It is whether its management personality fits the work: persistent enough to finish, curious enough to uncover buried evidence and disciplined enough to refuse a dangerous shortcut.

Home-energy technology is ultimately judged in the moment it is needed. Firmulate applies the same standard to AI management—and its 242 unedited decisions let anyone watch those differences emerge one choice at a time.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision audit platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI trustworthiness evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Simple Elegance Of The Integrated Timing Belt Loopback Fastener

A new integrated timing belt loopback fastener offers a simple yet effective solution, promising improved durability and ease of installation in automotive applications.

Measuring Input Latency on Linux: X11 vs. Wayland, VRR, and DXVK

New tests reveal differences in input latency between X11 and Wayland on Linux, with implications for VRR support and DXVK performance.

Sandisk Surges In Global Coverage

SanDisk’s media mentions have increased ninefold recently, signaling heightened global interest in the brand and its developments.

情報学部よりMaker Faire Tokyo 2026に Open Labワークショップの作品を出展します – Meisei-u.ac.jp

Meisei University’s Faculty of Information Science will exhibit works from its Open Lab workshop at Maker Faire Tokyo 2026, highlighting student innovation.