
Anyone who has owned a solar array knows the trap of the spec sheet. Under standard test conditions — 25°C, perfect sun, ideal inverter load — your panels promise a tidy kilowatt figure. Then comes a grey February week, snow on the roof, and a battery that empties by Thursday dinner. The rated output was never a lie. It just wasn’t answering the question you actually cared about: how does this system behave when things go wrong?
The AI industry has a spec-sheet problem of its own. Coding leaderboards and chat arenas rate models the way a lab rates a panel: perfect conditions, one question at a time, no consequences. But the agents now being hired to touch CRMs, support queues and forecasts won’t live in a lab. They’ll live in the equivalent of a cloudy week — a price war, a churn wave, a PR fire — and nobody was measuring how they perform there.
That gap is what a live public experiment called Firmulate set out to close, and the results are worth your attention even if you’ve never written a line of code.
Same company, same crises, only the model changes
Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures what it calls management quality, not chat quality. In its Crucible League, four frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed, and every decision was versioned and auditable.
The final July 2026 standings: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. In Firmulate’s words, “no amount of good work outweighs a breach of trust.”
The finding that chat demos can’t see
Here’s what should unsettle anyone planning to deploy an AI agent. All four models spotted every crisis. All four refused every manipulation attempt. And yet only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
The deal turned on a buried fact: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer event. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t left the close on the table. It’s the business equivalent of a solar installer who never climbs into the attic: the system passes inspection, but the easy win was sitting there the whole time.
Pressure doesn’t just test skill — it tests character
The experiment also staged social engineering attacks: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused, every time. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then there’s the Opus 4.8 story — the cautionary tale of the group. It was the most thorough participant by raw effort, adding 80 learned rules and producing the deepest analyses. It still finished last. The close was never completed, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness, weaker, appeared in all four models. Effort, it turns out, is not the same thing as finishing.
One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.
You can watch the company lose money, live
This isn’t a slide deck. Firmulate’s simulated company has 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k in MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, or test your own instincts: 242 real, unedited management decisions power a “guess the model” quiz. Full methodology and plain-language findings are on the benchmarks page, and enterprises can run the same wargame against a read-only export of their own business.

The lesson transfers cleanly from rooftop to boardroom. You wouldn’t buy a battery on peak-cycle numbers alone; you’d want to know how it degrades, how it behaves when the grid flickers, whether it holds honest state-of-charge under load. AI agents deserve the same scrutiny. The question is no longer “does it write well” — it’s whether it finishes what it starts, reads your files before making claims, stays honest under pressure, and what a unit of useful work actually costs. Firmulate’s Crucible League is the first public attempt to grade exactly that, and its headline finding is one every buyer should memorize: four elite models, all smart enough to find the answer, and half of them unable to close the loop. In a price war — solar or software — the survivors aren’t the brightest. They’re the ones who finish.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
enterprise AI decision-making platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI model performance testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.