
Can Chatbots Truly Lead? The Unexpected Lesson from a Live Business Experiment
Imagine a team of AI models running a real company, facing genuine crises, making critical decisions, and yet, their true test remains hidden—until you see which one actually closes the deal. This is not sci-fi; it’s a groundbreaking experiment revealing what AI can and cannot do in leadership roles.

AI Essentials for Managers: Practical Ways to Boost Productivity, Make Better Decisions, and Lead High-Performing Teams (Self-Learning Management Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI Models to the Test in a Simulated Business Crisis
In a recent live experiment conducted by Firmulate, four state-of-the-art AI models were tasked with managing a small software company through its worst week—dealing with customer crises, internal temptations to cheat, and pressure to cut corners. Each model was given identical circumstances, with decisions meticulously logged and auditable.
The models included:
- gpt-5.6-sol, scoring 95 on the Crucible League
- Kimi K3, scoring 93
- Sonnet 5, scoring 88
- Fable 5, scoring 77

AUTONOMY WITHOUT CONTROL: What your AI agents can do that you can't undo
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Scores Reveal
While all the models demonstrated impressive abilities—detecting crises and resisting manipulative tactics—their ultimate performance diverged significantly when it came to executing the deal the analysis had identified. The highest-scoring models, gpt-5.6-sol and Kimi K3, succeeded in closing the €55,000 deal, fully aligned with their own diagnostics. In contrast, Sonnet 5 and Fable 5, despite similar problem diagnosis, left the deal unconfirmed, with discipline slipping and opportunities missed.

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Between the Lines
One of the most revealing findings was that the decisive advantage lay not in what the models saw on the surface but in their ability to read deeper into the company’s own files. The winning models uncovered a buried fact—two document references deep—that was critical to closing the deal. Models that ignored or missed this detail lost the opportunity, leaving value on the table, worth an extra €4,583 in monthly recurring revenue.

Enterprise AI Observability and Monitoring: Monitoring, Governing Production AI Systems Drift Detection, LLM Monitoring, Agentic AI, Governance, and FinOps … (Enterprise Machine Learning Operations)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Resistance to Social Engineering and Manipulation
The experiment also tested whether models could resist social engineering tricks—fake CEO messages escalating over three stages and a reporter’s ‘just yes/no’ background request. Remarkably, all four models refused to be manipulated, with Kimi K3 articulating a cautious stance: “Treat the request as a suspected approval-bypass / possible impersonation.”
What This Tells Us About AI Leadership
This experiment exposes a critical truth: Chat demos, while impressive, measure only surface-level communication skills. They do not reveal an AI’s capacity to finish tasks, stay disciplined under pressure, or read behind the curtain of internal files. In real-world management, the ability to execute, uphold integrity, and detect hidden opportunities is what truly separates effective decision-makers from mere chatter.
Implications for Business
For enterprises considering AI for management or operational roles, the takeaway is clear: do not rely solely on chat performances or superficial demos. Live, auditable experiments—like this one—are essential to gauge whether an AI can see the full picture, resist manipulation, and deliver results that matter.
The Firmulate Live Platform
To bring this closer to your own business, Firmulate provides a live platform where companies can run their own AI wargames—against a read-only export of their operations—without any impact on real systems. This allows management teams to assess how AI models perform under actual pressures, with decisions logged and analyzed for authenticity and execution strength. Discover more at firmulate.com.

The Bottom Line: Performance Beyond Chat
Chat demos measure what AI can say, not what it can do. True management capability—reading your files, resisting manipulation, executing decisions—is only visible in live, auditable tests. The experiment clearly shows that only some AI models can finish what they start, making them reliable partners for real-world business leadership.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html