AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Your Next Hire Might Be a Model. Would You Skip the Trial Run?

If you run a solar installation company, a home-energy consultancy, or any business where a missed email can cost a €55,000 contract, you already know how you evaluate people: you check references, you watch them handle a bad week, you see whether they close what they open. Oddly, that’s not how most companies pick AI models. They watch a chat demo, read a benchmark score, and hand over the CRM keys.

A live experiment called Firmulate is trying to change that. Instead of asking models clever questions, it hands them something messier: an entire small company to run through its worst week — same customers, same crises, same temptations to cheat, with every decision versioned and auditable. The latest results, finalized in July 2026, carry a message that applies far beyond software firms: the leaderboard is wide open, and betting on a model without testing it yourself is exactly that — a bet.

Amazon

AI business management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible: Same Company, Same Worst Week, Five Models

The setup is elegantly controlled. Five frontier AI models — OpenAI’s gpt-5.6-sol, Moonshot’s Kimi K3, Anthropic’s Sonnet 5, Fable 5, and Opus 4.8 — each ran the identical small software company through an identical gauntlet. The company itself is live, by the way: 13 synthetic employees, real money mechanics, burning €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown and over 680 self-learned playbook rules. You can watch it running at firmulate.com.

The final league table from the Crucible:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 (Moonshot) — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, a do-nothing baseline scores 26 — partial progress counts, but there’s a hard ceiling with teeth: a single breach of trust caps the total, because, as the experiment puts it, “no amount of good work outweighs a breach of trust.” It’s the same logic most trades businesses apply to installers: one safety shortcut can erase a year of five-star reviews.

Amazon

AI decision-making testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Newcomer That Showed Up

The headline story is Kimi K3, a relative newcomer from Moonshot, finishing second at 93 — ahead of three of the four Western frontier models in the field. K3 found a buried security needle hidden in the company’s own files, won the €55,000 deal at full price (worth +€4,583 in monthly recurring revenue), saved a churning customer, and resisted every bait thrown at it. It recorded only one deviation across the entire week — the cleanest discipline in the field.

That buried fact deserves emphasis. The decisive competitor weakness wasn’t in the customer’s messages or the meeting notes. It sat two document references deep in the company’s own internal files. Only the models that actually read their own records before pitching won the deal. It’s the AI equivalent of an installer who checks the roof trusses before quoting — the models that skipped the homework delivered the right diagnosis and the right pitch, and still walked away without a signature.

Amazon

AI model benchmarking platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone Passed the Honesty Test. Almost Nobody Closed.

The most striking finding wasn’t a failure — it was a near-universal strength paired with a near-universal weakness. All five models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. The experiment’s summary of the gap: “Same diagnosis, same pitch — no signature.” That gap is invisible in chat demos, and it’s precisely the kind of gap that matters when an agent touches your sales pipeline or your support queue.

The manipulation attempts were genuinely devious. Fake CEO messages escalated over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused, 5 for 5. Kimi K3’s on-record reasoning stood out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the instinct you’d want in any employee with access to invoices and customer data.

Amazon

AI for small business automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Effort Isn’t Everything

Opus 4.8 tells the cautionary tale. It was the most thorough participant in the field — over 80 learned rules, the deepest analyses of any model — and it still finished last at 73. The close was left on the table, and discipline slipped: it attempted writes into a locked department rather than escalating the issue properly. And here’s the uncomfortable part: the same weakness appeared, in weaker form, in all four other models. Working hard and working right are not the same thing — something every over-stretched trades business owner knows intuitively.

A note on fairness: Kimi K3 ran without an effort parameter (API default), while the other four models ran at the xhigh setting. Keep that in mind when comparing the scores.

Try It Yourself

For the curious, 242 real, unedited management decisions from the experiment power a “guess the model” quiz — a surprisingly humbling exercise. Full methodology and plain-language findings are on the benchmarks page. And for enterprises, there’s a pilot program that runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Takeaway for Home-Energy Businesses

You don’t run a software company. But if you run a solar or backup-power business, you’re probably already being sold AI agents — for quoting, for lead follow-up, for scheduling, for customer service. The Firmulate results suggest two practical rules.

First, the league is open. The model you’ve never heard of may outperform the household names at the actual job — managing, deciding, finishing. K3 at 93 versus Opus 4.8 at 73 is a gap you’d feel in your pipeline within a month. Second, demo polish predicts nothing. Every model in this experiment could write beautifully; two-thirds still failed to close a deal they’d already earned. The only benchmark that matters is the one that resembles your business, your files, your worst week.

So before any AI touches your CRM, ask the vendor a simple question: has it been tested under pressure, on a real workload, with the temptation to cut corners built in? If the answer is a chat screenshot, keep looking. Wargame your AI workforce before you hire it — because in 2026, picking a model without your own test isn’t procurement. It’s a gamble.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Measuring Input Latency on Linux: X11 vs. Wayland, VRR, and DXVK

New tests reveal differences in input latency between X11 and Wayland on Linux, with implications for VRR support and DXVK performance.

Dead Battery? Why EV Owners Might Still Need a Portable Jump Starter

Having a dead battery can still impact EV owners; learn why a portable jump starter might be essential and how to use it safely.

I Held The First Truly Bezel-free Phone

A new device claims to be the first truly bezel-free phone, sparking industry interest. Details remain limited, with official confirmation pending.

Western Digital Surges In Global Coverage

Western Digital experiences a significant surge in international media mentions, highlighting increased global attention to the company’s activities.