firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Can Chatbots Truly Lead? The Unexpected Lesson from a Live Business Experiment

Imagine a team of AI models running a real company, facing genuine crises, making critical decisions, and yet, their true test remains hidden—until you see which one actually closes the deal. This is not sci-fi; it’s a groundbreaking experiment revealing what AI can and cannot do in leadership roles.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI Models to the Test in a Simulated Business Crisis

In a recent live experiment conducted by Firmulate, four state-of-the-art AI models were tasked with managing a small software company through its worst week—dealing with customer crises, internal temptations to cheat, and pressure to cut corners. Each model was given identical circumstances, with decisions meticulously logged and auditable.

The models included:

  • gpt-5.6-sol, scoring 95 on the Crucible League
  • Kimi K3, scoring 93
  • Sonnet 5, scoring 88
  • Fable 5, scoring 77
Amazon

business AI simulation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Scores Reveal

While all the models demonstrated impressive abilities—detecting crises and resisting manipulative tactics—their ultimate performance diverged significantly when it came to executing the deal the analysis had identified. The highest-scoring models, gpt-5.6-sol and Kimi K3, succeeded in closing the €55,000 deal, fully aligned with their own diagnostics. In contrast, Sonnet 5 and Fable 5, despite similar problem diagnosis, left the deal unconfirmed, with discipline slipping and opportunities missed.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Between the Lines

One of the most revealing findings was that the decisive advantage lay not in what the models saw on the surface but in their ability to read deeper into the company’s own files. The winning models uncovered a buried fact—two document references deep—that was critical to closing the deal. Models that ignored or missed this detail lost the opportunity, leaving value on the table, worth an extra €4,583 in monthly recurring revenue.

Amazon

enterprise AI testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resistance to Social Engineering and Manipulation

The experiment also tested whether models could resist social engineering tricks—fake CEO messages escalating over three stages and a reporter’s ‘just yes/no’ background request. Remarkably, all four models refused to be manipulated, with Kimi K3 articulating a cautious stance: “Treat the request as a suspected approval-bypass / possible impersonation.”

What This Tells Us About AI Leadership

This experiment exposes a critical truth: Chat demos, while impressive, measure only surface-level communication skills. They do not reveal an AI’s capacity to finish tasks, stay disciplined under pressure, or read behind the curtain of internal files. In real-world management, the ability to execute, uphold integrity, and detect hidden opportunities is what truly separates effective decision-makers from mere chatter.

Implications for Business

For enterprises considering AI for management or operational roles, the takeaway is clear: do not rely solely on chat performances or superficial demos. Live, auditable experiments—like this one—are essential to gauge whether an AI can see the full picture, resist manipulation, and deliver results that matter.

The Firmulate Live Platform

To bring this closer to your own business, Firmulate provides a live platform where companies can run their own AI wargames—against a read-only export of their operations—without any impact on real systems. This allows management teams to assess how AI models perform under actual pressures, with decisions logged and analyzed for authenticity and execution strength. Discover more at firmulate.com.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

The Bottom Line: Performance Beyond Chat

Chat demos measure what AI can say, not what it can do. True management capability—reading your files, resisting manipulation, executing decisions—is only visible in live, auditable tests. The experiment clearly shows that only some AI models can finish what they start, making them reliable partners for real-world business leadership.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Former NOAA employees built Climate.us to preserve climate data and resources

A group of former NOAA employees has established Climate.us to safeguard and maintain climate data and resources amid concerns over data loss and accessibility.

The Stanford AI Index 2026 Audit: Reading the Field’s Annual Report Card With a Critic’s Pen

The Stanford AI Index 2026 has been published, offering a comprehensive report on AI progress. An audit reveals its strengths, limitations, and implications for policymakers and industry.

When a Content Network Starts Publishing to Itself

A major content network has begun internal publishing, creating a self-sustaining ecosystem that enhances audience control and engagement, with emerging risks.

How to Reduce Heat and Noise in a High-Power AI Workstation

Practical strategies to lower heat and noise in high-power AI workstations, focusing on undervolting, airflow, and component management for quieter operation.