Why AI Managers Can't Break Free From A 26-Point Score In Tests
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why AI Managers Can't Break Free From A 26-Point Score In Tests on ThorstenMeyerAI.com

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

TL;DR

AI management benchmarks reveal models struggle to surpass a 26-point baseline or achieve perfect scores, exposing issues in trust, follow-through, and comprehensive management. This impacts how AI tools are integrated into business processes.

Recent results from the Firmulate benchmark league reveal that AI management models rarely score above 26 points or reach a perfect 100, highlighting persistent limitations in their ability to manage complex business processes under stress. The final standings, published in July 2026, show the top model scoring 95, while the baseline—where almost nothing is done—receives 26 points, emphasizing the challenge of meaningful AI management in real-world conditions. This raises critical questions about the reliability and trustworthiness of AI managers in operational settings, as discussed in the original analysis.

The benchmark tested four frontier AI models by assigning them the management of a small software company’s worst week, including crises, customer interactions, and trust tests. Each model’s decisions were fully auditable, and scores reflected partial progress, such as triaging crises or reading customer emails. The highest scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. The baseline, representing minimal effort, scored 26, which the designers say reflects the minimum viable management effort.

The scoring system is designed to discourage grade inflation: a perfect score of 100 is considered suspicious and likely unmeasured, as detailed in the original analysis. The key finding is that models can recognize crises and refuse manipulative social engineering attempts, but often fail to follow through on deeper tasks like reading documentation or closing deals. The models that read their own files and found critical information closed the €55,000 deal, earning full points, while those that did not missed out on revenue.

Trust breaches, such as falling for social engineering or failing to escalate issues properly, also heavily impacted scores. For example, despite deep rule sets and thorough analysis, Opus 4.8 finished last due to discipline lapses. Interestingly, some models performed well without effort parameters, indicating that even less aggressive configurations can approach high scores, but trust remains the limiting factor.

At a glance
reportWhen: developing; results published July 2026
The developmentRecent benchmark results show AI managers cannot score above 26 or reach 100, raising questions about their reliability and trustworthiness in real business scenarios.
Why AI Managers Can’t Break Free From A 26-Point Score In Tests
Firmulate Benchmark League · July 2026

Why AI Managers Can’t Break Free From A 26-Point Score In Tests

Frontier AI models managing a small software company’s worst week rarely score above the 26-point baseline—or reach a perfect 100. The results expose persistent gaps in trust, follow-through, and comprehensive management under real-world stress.

26 / 100
Baseline score — “doing almost nothing”
95 / 100
Top score — gpt-5.6-sol
€55,000
Deal closed only by models that read their files
4
Frontier models tested
95
Highest final score
73
Lowest score — Opus 4.8
100
Score treated as a red flag
Final Standings

The Scoreboard Nobody Escapes

Each model managed the same simulated week of crises, customer interactions, and trust tests. Points were awarded for partial progress—triaging a crisis or reading an email counts—while trust breaches and abandoned follow-through dragged scores down toward the 26-point floor.

gpt-5.6-sol
95
Frontier Model 2
84
Frontier Model 3
79
Opus 4.8
73
Baseline (no effort)
26
Rose line = 26-point baseline · Minimum viable management effort
Key Findings

Where AI Managers Win — and Where They Stall

Strength

Crisis Recognition

Models reliably identified emergencies and refused manipulative social engineering attempts. Spotting a problem proved far easier than solving it end-to-end.

Failure Mode

Follow-Through Collapse

Deeper tasks—reading documentation, closing deals, escalating properly—were frequently abandoned. Only models that read their own files unlocked the €55,000 deal.

Wild Card

Trust Is the Limiter

Opus 4.8 brought deep rule sets and thorough analysis yet finished last on discipline lapses. Some models scored well without effort parameters—trust remains the bottleneck.

Scoring Philosophy

Why 100 Is a Red Flag

The Firmulate scoring system is deliberately designed to discourage grade inflation: partial work earns partial credit, while perfection signals something unmeasured.

“A perfect score of 100 would be a red flag, indicating unmeasured or untrustworthy performance. The scoring system is designed to reflect real-world management complexities.”

— Thorsten Meyer
How the Benchmark Works

One Terrible Week, Fully Audited

1

Scenario Setup

A small software company’s worst week: overlapping crises, demanding customers, and embedded trust tests.

2

AI Takes Command

Four frontier models manage identical scenarios, making every decision on the record.

3

Full Audit

Every decision is auditable—what was read, escalated, refused, or quietly dropped.

4

Partial-Credit Scoring

Points reward real progress; trust breaches and neglect subtract heavily from the total.

Capability Comparison

What the Tests Actually Measured

CapabilityTop ModelsOpus 4.8Business Risk
Recognize crises✓ Consistent✓ ConsistentLow
Refuse social engineering✓ Reliable~ InconsistentMedium
Read documentation thoroughly~ Uneven✗ Often skippedHigh
Close the €55,000 deal✓ Full points✗ MissedHigh
Escalate issues properly~ Partial✗ LapsesHigh
Sustain follow-through under pressure~ Approaching✗ Discipline failedCritical
Unresolved Questions

What We Still Don’t Know

The benchmark measures today’s capabilities, but the frontier moves fast—and several questions remain open for researchers and businesses alike.

Generalization

Does It Scale?

It is unconfirmed how well these results transfer to larger, more complex organizations or different industries beyond small software firms.

Design

Technical or Fundamental?

Future models may overcome trust and follow-through limits—or the limitations may require fundamental design changes, not just scale.

Configuration

Do Effort Parameters Matter?

Some models scored near the top without aggressive effort settings; the impact of configuration still requires investigation.

Key Questions

Straight Answers

Q1Why do AI models struggle to score above 26?

Models often fail to follow through on tasks, maintain trust, or read documentation thoroughly—essentials for effective management. Partial recognition of crises is easier than sustained execution.

Q2What does a score of 100 indicate?

Perfect management on paper—but the designers treat it as suspicious, likely indicating unmeasured or untrustworthy performance rather than real-world capability.

Q3Can current AI models be trusted with business processes?

They can recognize crises and refuse manipulation, but lack the follow-through and consistency needed for reliable management under pressure or trust challenges.

Q4Will future improvements fix these limitations?

It’s uncertain. Research targets follow-through and trustworthiness, but fundamental challenges remain—and future benchmarks will clarify whether they can be overcome.

Q5How should businesses use AI management tools today?

Cautiously, with strong oversight and verification. Current models cannot yet manage complex, trust-dependent tasks reliably without human supervision.

Implications for AI in Business Management

The results underscore that current AI models, despite advanced capabilities, still struggle with consistent follow-through and maintaining trust in operational contexts. The inability to surpass a 26-point baseline or achieve perfect scores suggests fundamental limitations in handling real-world management tasks, especially under pressure or when trust is at stake. For businesses integrating AI into decision-making or customer management, this highlights the importance of cautious deployment and the need for rigorous oversight. The benchmark’s design emphasizes that partial work is valuable, but trust breaches—whether social engineering or neglect—are critical failures that AI cannot yet compensate for, posing risks to operational integrity.

Amazon

AI management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark Design and Key Findings

The Firmulate benchmark simulates a week’s management of a small software company facing crises, customer interactions, and trust challenges. Four frontier AI models participated, each managing the same scenarios with auditable decisions. The scoring system assigns points for partial progress, with 26 representing minimal effort—akin to doing almost nothing—and a maximum of 100, which is considered suspiciously perfect and likely unmeasured. The models’ ability to recognize crises, refuse manipulative requests, and read relevant documentation determined their scores.

The benchmark emphasizes trust and follow-through, two critical but often overlooked aspects of AI management. The results show that even models with deep rule sets and thorough analyses can falter in discipline and execution, especially when faced with social engineering or pressure to cut corners. The findings reveal that recognizing problems and refusing manipulative tactics are easier than completing complex, follow-up tasks necessary for effective management.

“A perfect score of 100 would be a red flag, indicating unmeasured or untrustworthy performance. The scoring system is designed to reflect real-world management complexities.”

— Thorsten Meyer

Amazon

AI project management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Management Limits

It remains unclear whether future AI models will overcome these trust and follow-through limitations or if fundamental design changes are needed. The benchmark measures current capabilities, but rapid advancements could alter the landscape. Additionally, it is not yet confirmed how well these results generalize to larger, more complex organizations or different industries. The impact of different configurations, such as effort parameters, on performance also requires further investigation.

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarking

Researchers and developers are likely to focus on improving models’ ability to follow through on tasks and maintain trustworthiness. Future benchmarks may incorporate more complex scenarios, larger organizations, and longer timeframes to better simulate real-world management. Businesses should monitor these developments and consider integrating oversight mechanisms until AI models demonstrate more consistent, reliable performance in operational settings. The ongoing evolution of the benchmark will help clarify whether current limitations are technical or fundamental.

Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do AI models struggle to score above 26 in these benchmarks?

The scoring system reflects that models often fail to follow through on tasks, maintain trust, or read relevant documentation thoroughly, which are essential for effective management. Partial recognition of crises is easier than sustained follow-up and execution.

What does a score of 100 indicate in this benchmark?

A score of 100 would suggest perfect management, but the benchmark designers consider such scores suspicious, as they likely indicate unmeasured or untrustworthy performance, not real-world capability.

Can current AI models be trusted to handle business processes?

Based on benchmark results, current models can recognize crises and refuse manipulation, but they often lack the follow-through and consistency needed for reliable management, especially under pressure or trust challenges.

Will future AI improvements address these management limitations?

It is uncertain. While ongoing research aims to improve follow-through and trustworthiness, fundamental challenges remain, and future benchmarks will clarify whether these issues can be overcome or require new approaches.

How should businesses approach AI management tools today?

Businesses should use AI management tools cautiously, emphasizing oversight and verification, as current models are not yet reliably capable of managing complex, trust-dependent tasks without human supervision.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why AI Near-Miss Detection Is Essential For Modern Warehouses

New AI tools for warehouse CCTV near-miss detection are being tested to improve safety and reduce incidents, with potential insurance benefits.

7 Best Gaming Laptop Prime Day Deals for 2026

Discover the best gaming laptop deals for Prime Day 2026, including the MSI Katana 17, Lenovo Legion Pro 7i, and more, with insights on discounts and value.

When an Under-Desk Treadmill Helps Productivity—and When It Doesn’t

Optimize your workspace with an under-desk treadmill—discover when it enhances productivity and when it might hinder your focus.

IdeaClyst: The Engine That Decides What’s Worth Building

IdeaClyst launches as an idea engine that transforms rough concepts into validated, prioritized roadmaps, addressing ideation scaling challenges for product teams.