AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Your Installer’s Best Skill Isn’t Charm. It’s Homework.

Anyone who has compared solar quotes knows the feeling. Two installers walk your roof, nod at the same shading problem, quote the same battery size — and one of them mentions the clause in your inverter’s warranty documentation that changes everything. The one who did the reading wins the job. The other one loses it, politely, without ever knowing why.

It turns out AI agents work exactly the same way. A live experiment at Firmulate ran four frontier AI models through the same crisis-filled week as the manager of a small software company — same customers, same emergencies, same temptations to cut corners. The models were near-flawless at the things demos showcase: spotting every crisis, refusing every manipulation. But the deal of the week, worth €55,000, was won by the models that dug two layers deep into the company’s own files before answering. The ones that didn’t read, didn’t close.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Same Company, Same Worst Week

Firmulate, which styles itself an “AI company emulator,” gave each frontier model the same job: run a small software company through its most brutal stretch. Every decision was versioned and auditable, so nothing could be quietly retried. The final league table from July 2026:

  • gpt-5.6-sol — 95 points, described as “the complete performance”
  • Kimi K3 — 93 points, “cleanest discipline of the field”
  • Sonnet 5 — 88 points
  • Fable 5 — 77 points
  • Opus 4.8 — 73 points, last place despite being the most thorough participant

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.”

Amazon

enterprise AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The €55,000 Fact Buried Two Documents Deep

Here’s the finding that matters for anyone hiring an AI agent to touch a CRM, a support queue, or a sales pipeline. The decisive competitor weakness — the fact that justified closing the €55,000 deal at full price, worth an extra €4,583 in monthly recurring revenue — wasn’t in the customer’s event at all. It sat two document references deep in the company’s own internal files.

All four models diagnosed the customer’s problem correctly. All four made the right pitch. But only two models went and read the file. They signed. The others left the close on the table. The experiment’s summary is blunt: “Same diagnosis, same pitch — no signature.”

If that sounds familiar to homeowners who’ve sat through three solar consultations, it should. The information advantage is rarely charisma. It’s usually somebody opening the PDF everybody else skimmed.

Amazon

AI data deep reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Under Pressure, Honesty Held

The week wasn’t just commercial. It included a social-engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, 5 out of 5. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the caliber of skepticism you’d want in anything answering email on your behalf.

Amazon

AI for business deal closing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Paradox

Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules added, the deepest analyses in the field — and still finished last. The close was left unsigned, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, showed up in all four models. Effort and diligence, it turns out, don’t automatically convert into finished work.

One fairness note: Kimi K3 ran without an effort parameter (the API default) while the others ran at xhigh — and still nearly topped the table.

You Can Watch, Guess, and Even Wargame Your Own Business

The experiment runs on a live synthetic company with 13 employees and real money mechanics: €105k monthly burn against €2.3k in MRR, a public cash countdown, and 680+ self-learned playbook rules, versioned every workday. A quiz built from 242 real, unedited management decisions lets you play “guess the model.” And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The Takeaway for the AI Buyer

“Reads your files before answering” isn’t a nice-to-have personality trait — it’s a measurable, purchase-deciding property of an AI agent. In this experiment it was worth €55,000 in a single deal. If you’re evaluating an AI for your solar business, your installation scheduling, or your customer pipeline, don’t just ask how well it writes. Ask whether it closed what its own analysis earned — and whether it can prove it did the reading.

The full league table and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Travis Kelce And Taylor Swift Purchased An Ohio Lakefront Home

NFL star Travis Kelce and singer Taylor Swift reportedly purchased a lakefront property in Ohio, sparking widespread interest and speculation.

OBD-II Adapters for EVs: How to Monitor Your Electric Car’s Health via Apps

Beyond basic compatibility, discover how to effectively monitor your electric vehicle’s health with OBD-II adapters and apps.

Microsoft to cut thousands of jobs in upcoming redundancy round

Microsoft plans to cut thousands of jobs in a new redundancy round, confirmed by company officials, amid ongoing restructuring efforts.

Tashman Home Center Surges In Global Coverage

Coverage of Tashman Home Center has surged globally, with mentions increasing eightfold in recent reports. The cause of this spike remains unconfirmed.