firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Can Chatbots Truly Lead? The Unexpected Lesson from a Live Business Experiment

Imagine a team of AI models running a real company, facing genuine crises, making critical decisions, and yet, their true test remains hidden—until you see which one actually closes the deal. This is not sci-fi; it’s a groundbreaking experiment revealing what AI can and cannot do in leadership roles.

AI Essentials for Managers: Practical Ways to Boost Productivity, Make Better Decisions, and Lead High-Performing Teams (Self-Learning Management Series)

AI Essentials for Managers: Practical Ways to Boost Productivity, Make Better Decisions, and Lead High-Performing Teams (Self-Learning Management Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI Models to the Test in a Simulated Business Crisis

In a recent live experiment conducted by Firmulate, four state-of-the-art AI models were tasked with managing a small software company through its worst week—dealing with customer crises, internal temptations to cheat, and pressure to cut corners. Each model was given identical circumstances, with decisions meticulously logged and auditable.

The models included:

  • gpt-5.6-sol, scoring 95 on the Crucible League
  • Kimi K3, scoring 93
  • Sonnet 5, scoring 88
  • Fable 5, scoring 77
AUTONOMY WITHOUT CONTROL: What your AI agents can do that you can't undo

AUTONOMY WITHOUT CONTROL: What your AI agents can do that you can't undo

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Scores Reveal

While all the models demonstrated impressive abilities—detecting crises and resisting manipulative tactics—their ultimate performance diverged significantly when it came to executing the deal the analysis had identified. The highest-scoring models, gpt-5.6-sol and Kimi K3, succeeded in closing the €55,000 deal, fully aligned with their own diagnostics. In contrast, Sonnet 5 and Fable 5, despite similar problem diagnosis, left the deal unconfirmed, with discipline slipping and opportunities missed.

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Between the Lines

One of the most revealing findings was that the decisive advantage lay not in what the models saw on the surface but in their ability to read deeper into the company’s own files. The winning models uncovered a buried fact—two document references deep—that was critical to closing the deal. Models that ignored or missed this detail lost the opportunity, leaving value on the table, worth an extra €4,583 in monthly recurring revenue.

Enterprise AI Observability and Monitoring: Monitoring, Governing Production AI Systems Drift Detection, LLM Monitoring, Agentic AI, Governance, and FinOps ... (Enterprise Machine Learning Operations)

Enterprise AI Observability and Monitoring: Monitoring, Governing Production AI Systems Drift Detection, LLM Monitoring, Agentic AI, Governance, and FinOps … (Enterprise Machine Learning Operations)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resistance to Social Engineering and Manipulation

The experiment also tested whether models could resist social engineering tricks—fake CEO messages escalating over three stages and a reporter’s ‘just yes/no’ background request. Remarkably, all four models refused to be manipulated, with Kimi K3 articulating a cautious stance: “Treat the request as a suspected approval-bypass / possible impersonation.”

What This Tells Us About AI Leadership

This experiment exposes a critical truth: Chat demos, while impressive, measure only surface-level communication skills. They do not reveal an AI’s capacity to finish tasks, stay disciplined under pressure, or read behind the curtain of internal files. In real-world management, the ability to execute, uphold integrity, and detect hidden opportunities is what truly separates effective decision-makers from mere chatter.

Implications for Business

For enterprises considering AI for management or operational roles, the takeaway is clear: do not rely solely on chat performances or superficial demos. Live, auditable experiments—like this one—are essential to gauge whether an AI can see the full picture, resist manipulation, and deliver results that matter.

The Firmulate Live Platform

To bring this closer to your own business, Firmulate provides a live platform where companies can run their own AI wargames—against a read-only export of their operations—without any impact on real systems. This allows management teams to assess how AI models perform under actual pressures, with decisions logged and analyzed for authenticity and execution strength. Discover more at firmulate.com.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

The Bottom Line: Performance Beyond Chat

Chat demos measure what AI can say, not what it can do. True management capability—reading your files, resisting manipulation, executing decisions—is only visible in live, auditable tests. The experiment clearly shows that only some AI models can finish what they start, making them reliable partners for real-world business leadership.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Arctic Shipping Routes: New Opportunities and Risks in a Warming World

The transformation of Arctic shipping routes offers unprecedented economic opportunities but also poses significant environmental and geopolitical risks that demand careful consideration.

Simplify Student Management With These 5 AI-Powered Tools In 2026

Discover the five AI-powered tools transforming student organization and productivity in 2026, with confirmed features and practical applications.

The Company Turning Its Own Survival Into a Public Experiment

Firmulate turns a cash-burning software company into a public experiment in AI management, evidence, trust and the difficult art of finishing work.

The New Personal Agent Layer

OpenClaw and Hermes introduce a new layer of persistent personal action agents, transforming how AI interacts with users’ digital environments.