The Ultimate AI Leaderboard Starts Once The Demo Concludes
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

The final phase of the Firmulate experiment has concluded, revealing a management-focused AI leaderboard. The models were tested on real-world management tasks under crisis conditions, emphasizing trust and decision quality. The results challenge traditional AI benchmarks by prioritizing management capabilities.

The final results of the Firmulate management AI experiment have been announced, revealing how different models perform when tasked with running a simulated company during its most challenging week. The original analysis can be explored in this detailed report. The experiment tested five frontier models on their ability to diagnose crises, make decisions, and maintain trust under pressure. This development underscores a shift in AI evaluation from simple response quality to management and trustworthiness, with significant implications for enterprise deployment.

The Firmulate experiment placed five AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—against each other in a simulated management crisis. The models were scored on their ability to identify issues, communicate effectively, and execute decisions, with the top model, gpt-5.6-sol, earning a score of 95. The evaluation was rigorous, including a trust standard: a single breach of trust capped the overall score. Despite all models detecting crises and resisting manipulation attempts, only two successfully secured a key €55,000 deal, highlighting a gap between diagnosis and action. This emphasizes the importance of comprehensive AI evaluation, as discussed in the original analysis.

For example, models that read the company’s files more thoroughly were more successful in closing deals, but some still failed to present the critical fact needed for the sale, illustrating that effective management involves more than just technical competence. The experiment also tested responses to social engineering, with all models refusing to disclose sensitive information, demonstrating strong safety posture. However, even the most comprehensive model struggled with completing managerial tasks that required disciplined escalation and follow-through, exposing limitations in current AI management capabilities.

At a glance
reportWhen: concluded July 2026
The developmentThe Firmulate live experiment has finished its management evaluation, producing a leaderboard based on AI models’ ability to handle a simulated company’s worst week.

Why Management Skills Are the Next Benchmark in AI

This experiment highlights that management quality—not just chat or coding proficiency—should become a core measure for AI evaluation in enterprise contexts. The ability to diagnose, decide, communicate, and trustworthiness under pressure directly impacts real-world outcomes, such as closing deals or maintaining trust. The results suggest that AI models must be assessed on their capacity to handle consequences, prioritize organizational goals, and operate reliably over time, which are critical for deploying AI at scale in complex business environments.

Amazon

AI management simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Evolution from Response Benchmarks to Management Testing

Traditional AI benchmarks focus on isolated tasks like coding or conversational accuracy, which do not reflect the realities of managing organizations or crises. The Firmulate experiment is a pioneering effort to simulate a company’s worst week, forcing models to navigate multiple crises, make strategic decisions, and maintain trust. This approach builds on prior developments in AI safety and reliability, emphasizing that effective management involves reading organizational context, resisting shortcuts, and completing tasks through proper channels. The experiment’s design, involving real money mechanics and versioned decision logs, provides a more authentic test of AI’s readiness for enterprise management.

Previous efforts have shown that models can perform well on narrow benchmarks, but the challenge remains whether they can handle the messy, unpredictable nature of real-world management. This experiment bridges that gap by creating a controlled but realistic environment where AI must demonstrate operational competence and trustworthiness over sustained periods.

“Management quality, not just chat quality, deserves its own category of AI evaluation. Our experiment shows that managing a crisis requires discipline, trust, and decision-making under pressure.”

— Thorsten Meyer, lead researcher at Firmulate

Amazon

business crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About AI Management Performance

While the experiment provides valuable insights, several questions remain. It is not yet clear whether these models can consistently perform in diverse, real-world organizational contexts beyond the simulated environment. The long-term reliability of AI management, especially under unpredictable or novel crises, remains to be tested. Additionally, the impact of different organizational structures and cultures on AI decision-making is still unknown. The experiment also does not fully address how models handle ethical dilemmas or complex human relationships in management scenarios.

Amazon

AI decision-making training programs

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Steps for AI Management Benchmarks and Deployment

Following these results, researchers and enterprises are likely to develop more sophisticated management benchmarks that include long-term decision tracking, trust audits, and multi-stakeholder interactions. Companies considering AI assistants for management roles should begin testing models in simulated environments tailored to their specific operations, focusing on trust, escalation protocols, and decision accountability. The next phase involves integrating these models into real organizational workflows under close supervision, with ongoing evaluation of their management quality and safety. The ultimate goal is to establish management as a formal category in AI assessment, ensuring models can reliably handle the complexities of running a business.

Amazon

enterprise AI safety tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main purpose of the Firmulate experiment?

The experiment aims to evaluate AI models’ ability to manage a company’s crises, make decisions, and maintain trust during a simulated worst week, shifting focus from response quality to management skills.

Why is management performance important in AI evaluation?

Management performance reflects an AI’s capacity to handle real-world organizational tasks, including decision-making, trustworthiness, and accountability, which are critical for enterprise deployment.

Which model performed the best in the final leaderboard?

gpt-5.6-sol finished first with a score of 95, demonstrating strong crisis management and decision-making capabilities.

What are the limitations of the current experiment?

It remains uncertain whether these models can handle diverse, unpredictable real-world scenarios, and how they will perform over longer periods or in different organizational cultures.

What are the next steps for AI in management roles?

Future efforts will focus on developing more comprehensive benchmarks, testing models in enterprise environments, and establishing management as a formal AI evaluation category.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Operational SOP drift detector for franchise operators

A new SOP drift detection tool for multi-location franchise operators is being tested to monitor and address procedural deviations across locations.

Grant deadline radar for arts nonprofits

A new grant calendar tool for small arts nonprofits is being tested to reduce missed deadlines and streamline funding workflows, with initial validation underway.

A Skill Is A Folder, Not A Prompt: What Anthropic Learned Running Hundreds Of Them

Anthropic reveals that effective AI Skills are structured as folders containing instructions, scripts, and assets, transforming how organizations build and reuse AI capabilities.

Why Office Audio Quality Matters More Than Camera Quality in Meetings

Keen focus on office audio quality can make or break your meetings, as clear sound ensures your message is understood—discover why it matters more than camera quality.