🔍 Read the full analysis: Why AI Managers Can't Break Free From A 26-Point Score In Tests on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
AI management benchmarks reveal models struggle to surpass a 26-point baseline or achieve perfect scores, exposing issues in trust, follow-through, and comprehensive management. This impacts how AI tools are integrated into business processes.
Recent results from the Firmulate benchmark league reveal that AI management models rarely score above 26 points or reach a perfect 100, highlighting persistent limitations in their ability to manage complex business processes under stress. The final standings, published in July 2026, show the top model scoring 95, while the baseline—where almost nothing is done—receives 26 points, emphasizing the challenge of meaningful AI management in real-world conditions. This raises critical questions about the reliability and trustworthiness of AI managers in operational settings, as discussed in the original analysis.
The benchmark tested four frontier AI models by assigning them the management of a small software company’s worst week, including crises, customer interactions, and trust tests. Each model’s decisions were fully auditable, and scores reflected partial progress, such as triaging crises or reading customer emails. The highest scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. The baseline, representing minimal effort, scored 26, which the designers say reflects the minimum viable management effort.
The scoring system is designed to discourage grade inflation: a perfect score of 100 is considered suspicious and likely unmeasured, as detailed in the original analysis. The key finding is that models can recognize crises and refuse manipulative social engineering attempts, but often fail to follow through on deeper tasks like reading documentation or closing deals. The models that read their own files and found critical information closed the €55,000 deal, earning full points, while those that did not missed out on revenue.
Trust breaches, such as falling for social engineering or failing to escalate issues properly, also heavily impacted scores. For example, despite deep rule sets and thorough analysis, Opus 4.8 finished last due to discipline lapses. Interestingly, some models performed well without effort parameters, indicating that even less aggressive configurations can approach high scores, but trust remains the limiting factor.
Why AI Managers Can’t Break Free From A 26-Point Score In Tests
Frontier AI models managing a small software company’s worst week rarely score above the 26-point baseline—or reach a perfect 100. The results expose persistent gaps in trust, follow-through, and comprehensive management under real-world stress.
The Scoreboard Nobody Escapes
Each model managed the same simulated week of crises, customer interactions, and trust tests. Points were awarded for partial progress—triaging a crisis or reading an email counts—while trust breaches and abandoned follow-through dragged scores down toward the 26-point floor.
Where AI Managers Win — and Where They Stall
Crisis Recognition
Models reliably identified emergencies and refused manipulative social engineering attempts. Spotting a problem proved far easier than solving it end-to-end.
Follow-Through Collapse
Deeper tasks—reading documentation, closing deals, escalating properly—were frequently abandoned. Only models that read their own files unlocked the €55,000 deal.
Trust Is the Limiter
Opus 4.8 brought deep rule sets and thorough analysis yet finished last on discipline lapses. Some models scored well without effort parameters—trust remains the bottleneck.
Why 100 Is a Red Flag
The Firmulate scoring system is deliberately designed to discourage grade inflation: partial work earns partial credit, while perfection signals something unmeasured.
“A perfect score of 100 would be a red flag, indicating unmeasured or untrustworthy performance. The scoring system is designed to reflect real-world management complexities.”
— Thorsten MeyerOne Terrible Week, Fully Audited
Scenario Setup
A small software company’s worst week: overlapping crises, demanding customers, and embedded trust tests.
AI Takes Command
Four frontier models manage identical scenarios, making every decision on the record.
Full Audit
Every decision is auditable—what was read, escalated, refused, or quietly dropped.
Partial-Credit Scoring
Points reward real progress; trust breaches and neglect subtract heavily from the total.
What the Tests Actually Measured
| Capability | Top Models | Opus 4.8 | Business Risk |
|---|---|---|---|
| Recognize crises | ✓ Consistent | ✓ Consistent | Low |
| Refuse social engineering | ✓ Reliable | ~ Inconsistent | Medium |
| Read documentation thoroughly | ~ Uneven | ✗ Often skipped | High |
| Close the €55,000 deal | ✓ Full points | ✗ Missed | High |
| Escalate issues properly | ~ Partial | ✗ Lapses | High |
| Sustain follow-through under pressure | ~ Approaching | ✗ Discipline failed | Critical |
What We Still Don’t Know
The benchmark measures today’s capabilities, but the frontier moves fast—and several questions remain open for researchers and businesses alike.
Does It Scale?
It is unconfirmed how well these results transfer to larger, more complex organizations or different industries beyond small software firms.
Technical or Fundamental?
Future models may overcome trust and follow-through limits—or the limitations may require fundamental design changes, not just scale.
Do Effort Parameters Matter?
Some models scored near the top without aggressive effort settings; the impact of configuration still requires investigation.
Straight Answers
Q1Why do AI models struggle to score above 26?
Models often fail to follow through on tasks, maintain trust, or read documentation thoroughly—essentials for effective management. Partial recognition of crises is easier than sustained execution.
Q2What does a score of 100 indicate?
Perfect management on paper—but the designers treat it as suspicious, likely indicating unmeasured or untrustworthy performance rather than real-world capability.
Q3Can current AI models be trusted with business processes?
They can recognize crises and refuse manipulation, but lack the follow-through and consistency needed for reliable management under pressure or trust challenges.
Q4Will future improvements fix these limitations?
It’s uncertain. Research targets follow-through and trustworthiness, but fundamental challenges remain—and future benchmarks will clarify whether they can be overcome.
Q5How should businesses use AI management tools today?
Cautiously, with strong oversight and verification. Current models cannot yet manage complex, trust-dependent tasks reliably without human supervision.
Implications for AI in Business Management
The results underscore that current AI models, despite advanced capabilities, still struggle with consistent follow-through and maintaining trust in operational contexts. The inability to surpass a 26-point baseline or achieve perfect scores suggests fundamental limitations in handling real-world management tasks, especially under pressure or when trust is at stake. For businesses integrating AI into decision-making or customer management, this highlights the importance of cautious deployment and the need for rigorous oversight. The benchmark’s design emphasizes that partial work is valuable, but trust breaches—whether social engineering or neglect—are critical failures that AI cannot yet compensate for, posing risks to operational integrity.
As an affiliate, we earn on qualifying purchases.
Benchmark Design and Key Findings
The Firmulate benchmark simulates a week’s management of a small software company facing crises, customer interactions, and trust challenges. Four frontier AI models participated, each managing the same scenarios with auditable decisions. The scoring system assigns points for partial progress, with 26 representing minimal effort—akin to doing almost nothing—and a maximum of 100, which is considered suspiciously perfect and likely unmeasured. The models’ ability to recognize crises, refuse manipulative requests, and read relevant documentation determined their scores.
The benchmark emphasizes trust and follow-through, two critical but often overlooked aspects of AI management. The results show that even models with deep rule sets and thorough analyses can falter in discipline and execution, especially when faced with social engineering or pressure to cut corners. The findings reveal that recognizing problems and refusing manipulative tactics are easier than completing complex, follow-up tasks necessary for effective management.
“A perfect score of 100 would be a red flag, indicating unmeasured or untrustworthy performance. The scoring system is designed to reflect real-world management complexities.”
— Thorsten Meyer
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Management Limits
It remains unclear whether future AI models will overcome these trust and follow-through limitations or if fundamental design changes are needed. The benchmark measures current capabilities, but rapid advancements could alter the landscape. Additionally, it is not yet confirmed how well these results generalize to larger, more complex organizations or different industries. The impact of different configurations, such as effort parameters, on performance also requires further investigation.
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarking
Researchers and developers are likely to focus on improving models’ ability to follow through on tasks and maintain trustworthiness. Future benchmarks may incorporate more complex scenarios, larger organizations, and longer timeframes to better simulate real-world management. Businesses should monitor these developments and consider integrating oversight mechanisms until AI models demonstrate more consistent, reliable performance in operational settings. The ongoing evolution of the benchmark will help clarify whether current limitations are technical or fundamental.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do AI models struggle to score above 26 in these benchmarks?
The scoring system reflects that models often fail to follow through on tasks, maintain trust, or read relevant documentation thoroughly, which are essential for effective management. Partial recognition of crises is easier than sustained follow-up and execution.
What does a score of 100 indicate in this benchmark?
A score of 100 would suggest perfect management, but the benchmark designers consider such scores suspicious, as they likely indicate unmeasured or untrustworthy performance, not real-world capability.
Can current AI models be trusted to handle business processes?
Based on benchmark results, current models can recognize crises and refuse manipulation, but they often lack the follow-through and consistency needed for reliable management, especially under pressure or trust challenges.
Will future AI improvements address these management limitations?
It is uncertain. While ongoing research aims to improve follow-through and trustworthiness, fundamental challenges remain, and future benchmarks will clarify whether these issues can be overcome or require new approaches.
How should businesses approach AI management tools today?
Businesses should use AI management tools cautiously, emphasizing oversight and verification, as current models are not yet reliably capable of managing complex, trust-dependent tasks without human supervision.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
