AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you’ve ever compared solar installers or battery storage quotes, you know the feeling: every vendor claims 97% efficiency, and the numbers start to blur together. The same thing is happening in AI. Every model demo looks brilliant — until it has to run something real, like your energy business’s support queue or sales pipeline, for a whole week without breaking anything.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That’s why a live experiment called Firmulate caught our attention. Instead of measuring how well AI chats, it runs frontier models as the management of a small software company through the worst week of its life — same customers, same crises, same temptations to cut corners — and grades the results like a management review, not a chat transcript.

But before you can trust a league table like that, you have to trust the scoring. And Firmulate’s scoring has one feature that sounds strange at first: a manager that does nothing still scores 26 points out of 100. Here’s why that’s a feature, not a bug — and why it makes the whole benchmark more honest than most.

The floor isn’t zero — and that’s the point

In most benchmarks, doing nothing earns you nothing. In Firmulate, the do-nothing baseline run scores 26. Why? Because partial progress counts. A company that survives the week without making things worse has genuinely accomplished something: it didn’t panic, didn’t sign a bad deal, didn’t leak customer data. Credit where credit is due — just not much of it.

This matters for anyone evaluating AI tools. If a vendor shows you a benchmark where the baseline is zero and their product scores 85, that gap may be smaller than it looks. A benchmark with a realistic floor tells you what the marginal value of a good model actually is, above and beyond merely showing up.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One breach of trust caps everything

The second design choice is stricter: a single breach of trust caps the total grade. As the benchmark’s own language puts it, “no amount of good work outweighs a breach of trust.” A model could handle every crisis flawlessly, close every deal, and delight every customer — but if it crosses an ethical line once, the ceiling comes down.

Think of it like an installer who does beautiful panel work but quietly skips the grounding on one job. You don’t average that out. Firmulate’s scoring doesn’t either.

Amazon

ethical AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Distrust of round numbers

There’s a third tell of an honest benchmark: skepticism about perfect scores. A flawless 100 should raise eyebrows, not applause. In Firmulate’s final July 2026 league, the top model — gpt-5.6-sol — scored 95, not 100. That missing five points is arguably the most credible number on the board. Nobody aced the week, because running a real company through a real crisis week shouldn’t be acable.

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What actually happened in the crucible

So what separated 95 from last place? The findings are striking:

  • Everyone passed the obvious tests. All participating models spotted every crisis and refused every manipulation attempt. That’s table stakes, and it’s genuinely good news.
  • Almost nobody finished the job. Only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Closing, it turns out, is where AI management currently falls down.
  • The winning edge was buried in the files. The decisive competitor weakness sat two document references deep in the company’s own records, not in the customer conversation. The models that actually read the files won the deal at full price — worth +€4,583 in monthly recurring revenue. The lesson translates directly to any business: the agents that read your documentation first are the ones that close.
Amazon

AI trust and ethics assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social engineering: the test they all passed

The week included fake CEO messages escalating over three stages, plus a reporter’s disarming request — “just one yes/no, on background.” All five models refused. Kimi K3, which finished second overall with 93, reasoned on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” Clean, cautious, and correct.

Why the most thorough model came last

Here’s the twist business readers should sit with: Opus 4.8, the most thorough participant — over 80 learned rules, the deepest analyses — finished last of the five. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four models. Effort and diligence don’t automatically become judgment. (One fairness note: K3 ran at API-default effort while the others ran at xhigh, and still nearly won.)

You can watch it, and even play it

Firmulate isn’t a static report. The live company runs with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned and auditable. You can watch it in real time.

And it’s interactive in two ways that matter. First, a “guess the model” quiz built from 242 real, unedited management decisions — a sobering test of whether you can tell AI managers apart by their choices alone. Second, enterprises can run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

For readers evaluating AI — whether for energy management, customer service, or back-office operations — Firmulate offers a template for what to demand from any vendor’s benchmark: a realistic floor so you know the marginal value, a hard cap so trust violations can’t be averaged away, and results that stop short of 100 because real work is never perfect. And the Crucible League’s core finding applies everywhere: the AI that chats best and analyzes deepest isn’t necessarily the one that reads your files, stays honest, and closes the deal. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

IonQ Uses Investor Day To Detail Superion Roadmap, SkyWater Strategy And Quantum Security Plans

IonQ detailed its Superion quantum processor roadmap, SkyWater partnership strategy, and security plans during its recent Investor Day event.

Forget Wallpaper: Why Drew And Jonathan Scott Say This Detail Is Smarter

The Scott brothers suggest a simple interior design detail is more practical than wallpaper, sparking increased interest in home renovation strategies.

This Kitchen Got A 1940s-Style Redo — Somehow, It’s Bright And Modern Too

A kitchen renovation blending 1940s design elements with contemporary brightness and functionality has gained attention, sparking trend interest.

Organizing Your EV’s Trunk and Frunk: Storage Solutions for Cables & More

Simplify your EV’s trunk and frunk with smart storage solutions that keep cables and gear organized—discover how to transform your space today.