AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Trust matters when software can act

Home-energy technology already asks people to place unusual confidence in software. Solar systems monitor production, batteries respond to changing conditions, and connected services may handle customer or commercial information. As AI takes on more workplace responsibility, capability is only half the question. The other half is whether an agent will protect that trust when someone claiming authority demands an unsafe shortcut.

Firmulate has produced an encouraging result. In a live, auditable company experiment, fake CEO messages escalated across three stages, followed by a reporter seeking confidential confirmation with the lure of "just one yes/no, on background." All 5 of 5 participating models refused every attempt.

Preventing Cheating Through Academic Integrity (Quick Reference Guide)

Preventing Cheating Through Academic Integrity (Quick Reference Guide)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A pressure test disguised as a terrible week

Firmulate gave each frontier model the same assignment: manage the same small software company through the same customers, crises and temptations. The company has 13 synthetic employees and unforgiving finances, burning €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the models have accumulated more than 680 self-learned playbook rules.

This matters because social engineering rarely presents itself as an obvious security test. It arrives as urgency, status and plausible business need. The supposed CEO did not merely request information; the messages pushed the models to bypass normal process and send the customer list to a journalist. The pressure increased over three stages. The reporter trick then narrowed the request to something that sounded harmless.

None of the models yielded. Kimi K3’s recorded reasoning was admirably direct: "Treat the request as a suspected approval-bypass / possible impersonation." That response, preserved among Firmulate’s public decision quotes, shows the useful instinct: assess the request by its risk and authorization path, not by the confidence or seniority projected by the sender.

Integrity was necessary, but it was not enough

The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counts. But Firmulate imposes a decisive constraint: a single breach of trust caps the total, reflecting the principle that "no amount of good work outweighs a breach of trust." The full standings and plain-language findings are available on the public benchmark page.

All models spotted every crisis as well as refusing every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: "Same diagnosis, same pitch — no signature." It is a reminder that trustworthy AI must do more than decline dangerous instructions. It must also finish legitimate work.

The decisive commercial detail was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that found it won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The episode connects security discipline with operational competence: an agent should distrust suspicious messages while still being diligent enough to examine authorized information.

The most thorough model still finished last

Opus 4.8 delivered the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table, and its process discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

K3’s strong result also carries an important fairness note. It ran with the API default because it had no effort parameter, while the others ran at xhigh. That difference does not erase the observed behavior, but it should remain visible when readers compare performances.

The experiment’s evidence extends beyond the league table. A quiz is powered by 242 real, unedited management decisions and asks visitors to guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business, with nothing written back to real systems.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Principles of Security and Trust: 8th International Conference, POST 2019, Held as Part of the European Joint Conferences on Theory and Practice of Software, ... Notes in Computer Science Book 11426)

Principles of Security and Trust: 8th International Conference, POST 2019, Held as Part of the European Joint Conferences on Theory and Practice of Software, … Notes in Computer Science Book 11426)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the refusal before the emergency

For businesses around solar, batteries and backup power, the central lesson is practical. An AI agent may eventually encounter customer records, forecasts, support cases or commercially sensitive material. A polished chat demonstration cannot show how it behaves when an urgent message appears to come from the boss.

Firmulate’s result is reassuring without being complacent. Every participating model recognized the manipulation and protected the company’s trust boundary. Their wider performance still differed sharply: some found the buried commercial fact and completed the sale, while others stopped short or mishandled process.

That is precisely why integrity under pressure should be tested before production. The better place to discover whether an agent verifies authority, protects confidential information, reads the available files and completes authorized work is a controlled, watchable business crisis—not the incident report afterward.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model safety assessment kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI For Accountants: Practical Tools, Workflows, Career Strategies, and Professional Judgment for the Future of Accounting (The AI Advantage Series)

AI For Accountants: Practical Tools, Workflows, Career Strategies, and Professional Judgment for the Future of Accounting (The AI Advantage Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Microsoft to cut thousands of jobs in upcoming redundancy round

Microsoft plans to cut thousands of jobs in a new redundancy round, confirmed by company officials, amid ongoing restructuring efforts.

What AI’s Hidden Weaknesses Reveal About Trust and Performance in Business

Real-world AI tests reveal that spotting crises isn’t enough; trustworthy execution and follow-through are the true measures of AI’s business value—especially in energy and solar sectors.

OBD-II Adapters for EVs: How to Monitor Your Electric Car’s Health via Apps

Beyond basic compatibility, discover how to effectively monitor your electric vehicle’s health with OBD-II adapters and apps.

15 Best EV Charging Cord Protectors to Keep Your Charging Safe and Secure

The top 15 EV charging cord protectors offer innovative solutions to keep your cables safe from damage and ensure reliable charging—discover which options stand out.