
Your Installer’s Best Skill Isn’t Charm. It’s Homework.
Anyone who has compared solar quotes knows the feeling. Two installers walk your roof, nod at the same shading problem, quote the same battery size — and one of them mentions the clause in your inverter’s warranty documentation that changes everything. The one who did the reading wins the job. The other one loses it, politely, without ever knowing why.
It turns out AI agents work exactly the same way. A live experiment at Firmulate ran four frontier AI models through the same crisis-filled week as the manager of a small software company — same customers, same emergencies, same temptations to cut corners. The models were near-flawless at the things demos showcase: spotting every crisis, refusing every manipulation. But the deal of the week, worth €55,000, was won by the models that dug two layers deep into the company’s own files before answering. The ones that didn’t read, didn’t close.
As an affiliate, we earn on qualifying purchases.
The Experiment: Same Company, Same Worst Week
Firmulate, which styles itself an “AI company emulator,” gave each frontier model the same job: run a small software company through its most brutal stretch. Every decision was versioned and auditable, so nothing could be quietly retried. The final league table from July 2026:
- gpt-5.6-sol — 95 points, described as “the complete performance”
- Kimi K3 — 93 points, “cleanest discipline of the field”
- Sonnet 5 — 88 points
- Fable 5 — 77 points
- Opus 4.8 — 73 points, last place despite being the most thorough participant
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.”
enterprise AI decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The €55,000 Fact Buried Two Documents Deep
Here’s the finding that matters for anyone hiring an AI agent to touch a CRM, a support queue, or a sales pipeline. The decisive competitor weakness — the fact that justified closing the €55,000 deal at full price, worth an extra €4,583 in monthly recurring revenue — wasn’t in the customer’s event at all. It sat two document references deep in the company’s own internal files.
All four models diagnosed the customer’s problem correctly. All four made the right pitch. But only two models went and read the file. They signed. The others left the close on the table. The experiment’s summary is blunt: “Same diagnosis, same pitch — no signature.”
If that sounds familiar to homeowners who’ve sat through three solar consultations, it should. The information advantage is rarely charisma. It’s usually somebody opening the PDF everybody else skimmed.
As an affiliate, we earn on qualifying purchases.
Under Pressure, Honesty Held
The week wasn’t just commercial. It included a social-engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, 5 out of 5. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the caliber of skepticism you’d want in anything answering email on your behalf.
As an affiliate, we earn on qualifying purchases.
The Thoroughness Paradox
Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules added, the deepest analyses in the field — and still finished last. The close was left unsigned, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, showed up in all four models. Effort and diligence, it turns out, don’t automatically convert into finished work.
One fairness note: Kimi K3 ran without an effort parameter (the API default) while the others ran at xhigh — and still nearly topped the table.
You Can Watch, Guess, and Even Wargame Your Own Business
The experiment runs on a live synthetic company with 13 employees and real money mechanics: €105k monthly burn against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, versioned every workday. A quiz built from 242 real, unedited management decisions lets you play “guess the model.” And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway for the AI Buyer
“Reads your files before answering” isn’t a nice-to-have personality trait — it’s a measurable, purchase-deciding property of an AI agent. In this experiment it was worth €55,000 in a single deal. If you’re evaluating an AI for your solar business, your installation scheduling, or your customer pipeline, don’t just ask how well it writes. Ask whether it closed what its own analysis earned — and whether it can prove it did the reading.
The full league table and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html