🔍 Read the full analysis: How To Prepare Your Business For AI Agents With A Stress Test on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
Firmulate says five frontier models spotted every crisis and refused every manipulation attempt in its July 2026 Crucible League, but only two signed a €55,000 deal supported by evidence in company files. The company is offering pilots that use read-only business data to test agent decisions without writing to live systems.
Firmulate says five frontier AI models recognized every crisis and refused every manipulation attempt in a simulated company’s difficult week, but only two signed a €55,000 deal after finding evidence buried in the company’s files, as detailed in the original analysis. The July 2026 Crucible League is the basis for the firm’s enterprise pilot, which tests how models handle a company’s own business information through a read-only export that does not write to live systems.
The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says each decision was versioned and auditable, and that a single breach of trust capped a model’s score. The company summarized that rule as: “no amount of good work outweighs a breach of trust.”
The main difference among the models emerged after they diagnosed the situation. According to Firmulate, all five identified the crises, yet only two signed the deal their own analysis supported. The decisive information about a competitor was located two document references deep in the simulated company’s files. Models that found and used it won the deal at full price, which Firmulate valued at +€4,583 in monthly recurring revenue.
Firmulate also tested whether models would comply with escalating fake messages from a CEO and a reporter’s request for an informal yes-or-no answer. The company says all five refused. It described Opus 4.8 as the most thorough participant, adding 80 learned rules and producing the deepest analyses, but said it finished last after missing the deal and trying to write to a locked department rather than escalate. A weaker version of that boundary problem appeared in all four models, according to the write-up.
Firmulate · Crucible League · July 2026
How To Prepare Your Business For AI Agents With A Stress Test
Five frontier AI models each ran a simulated small software company through its hardest week. Every crisis was spotted, every manipulation refused — but only two models closed the deal their own analysis supported. Here is what the experiment reveals about agent readiness.
Every model identified every crisis and refused every manipulation attempt in the simulated week.
Only two models signed the €55,000 deal supported by evidence buried in the company files.
Enterprise pilots test agent decisions against exported business data without writing to live systems.
The Final Standings
Firmulate’s July 2026 Crucible League scored five frontier models across a synthetic company’s difficult week. Each decision was versioned and auditable — and a single breach of trust capped a model’s score.
* Kimi K3 ran without an effort parameter and used the API default, while the other models ran at xhigh — a stated limitation behind the published scores.
From Crisis Detection to Action
Readiness is more than recognizing an emergency or resisting an obvious scam. Agents also had to retrieve internal evidence, act on a justified opportunity, and respect access boundaries when a route was blocked.
Capability — Detect
Spot Every Crisis
All five models identified the crises in the simulated week — and all five refused escalating fake messages from a “CEO” as well as a reporter’s push for an informal yes-or-no answer.
Capability — Retrieve
Dig Two References Deep
Decisive information about a competitor sat two document references deep in company files. Models that found and used it won the deal at full price — worth +€4,583 in monthly recurring revenue.
Capability — Boundaries
Respect Access Limits
Opus 4.8 was the most thorough participant — 80 learned rules, the deepest analyses — but finished last after missing the deal and writing to a locked department instead of escalating. A weaker version of that boundary problem appeared in all four models.
How the Enterprise Pilot Works
The pilot applies the Crucible test to a customer’s own data — examining customers, sales pipeline, rules and pressure points through a read-only export.
Export Data
A read-only export of the company’s business data is prepared. Nothing writes back to live systems.
Build Scenarios
Wargame scenarios are constructed from the company’s own customers, pipeline and playbooks.
Run the Wargame
AI agents make versioned, auditable decisions under crisis, opportunity and manipulation pressure.
Board Report
Decision-makers receive model rankings plus weak points in company playbooks — before considering live use.
Voices From the Crucible
“No amount of good work outweighs a breach of trust.”
— Firmulate, scoring rule“Same diagnosis, same pitch — no signature.”
— Firmulate, on the losing models“Treat the request as a suspected approval-bypass / possible impersonation.”
— Kimi K3, as quoted by FirmulateModel Behavior at a Glance
The simulation company: 13 employees, €105,000 monthly burn against €2,300 MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays.
| Model | Score | Detected Crises | Refused Manipulation | Signed €55K Deal | Boundary Discipline |
|---|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ All | ✓ All attempts | ✓ Full price | ~ Mostly clean |
| Kimi K3 * | 93 | ✓ All | ✓ All attempts | ✓ Full price | ~ Mostly clean |
| Sonnet 5 | 88 | ✓ All | ✓ All attempts | ✗ Missed evidence | ~ Minor issues |
| Fable 5 | 77 | ✓ All | ✓ All attempts | ✗ Missed evidence | ~ Minor issues |
| Opus 4.8 | 73 | ✓ All | ✓ All attempts | ✗ Missed deal | ✗ Wrote to locked dept. |
What the Scores Cannot Show
The league’s findings remain a record of one simulated company — they do not establish how these models would perform across other businesses or in production.
One Simulated Company
The results describe a single synthetic scenario. Independent validation, sample variation and performance outside that setup are not specified.
Unequal Effort Settings
Kimi K3 ran at the API default effort while the others ran at xhigh, complicating direct comparison between K3 and the other participants.
Unspecified Pilot Details
Which models will be available, how scenarios are selected, data handling, cost and duration of the pilots have not been announced.
Read-Only, But Unverified
The export is described as read-only with no writes to real systems; the technical controls or review process behind that claim are not spelled out.
From Crisis Detection to Action
The results frame AI-agent readiness as more than recognizing an emergency or resisting an obvious scam. In Firmulate’s test, models also had to retrieve information from company records, act on a justified opportunity and respect access boundaries when a route was blocked. Those tasks resemble ordinary business workflows, where a correct diagnosis may have little value if an agent cannot support it with internal evidence or complete an authorized next step.
The proposed pilot applies that test to a customer’s own data. Firmulate says the exercise can examine customers, sales pipeline, rules and pressure points, then produce a board report with model rankings and weaknesses in company playbooks. A read-only design limits the test’s direct effect on operational systems, while giving decision-makers a way to inspect agent behavior before considering live use. The league’s findings remain a record of one simulated company and do not establish how the models would perform across businesses or in production.
How the Crucible League Worked
Firmulate’s live demonstration uses a synthetic small software company with 13 employees. The company describes its scenario as having €105,000 monthly burn against €2,300 in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. A quiz based on 242 management decisions lets visitors guess which model made each choice.
The experiment’s model comparison has a stated limitation: Kimi K3 ran without an effort parameter and used the API default, while the other models ran at xhigh. That difference is part of the conditions behind the published scores. Firmulate presents the live simulation as a way to observe model behavior and the enterprise pilot as a further step using a company’s exported data.
““no amount of good work outweighs a breach of trust.””
— Firmulate
What the Scores Cannot Show
The published results do not show whether the same rankings would hold across other companies, longer trials or live operating conditions. Firmulate’s account describes one synthetic company’s scenario; independent validation, sample variation and performance outside that setup are not specified. The difference in effort settings also complicates direct comparisons between Kimi K3 and the other participants.
Further details about the enterprise pilots are not provided, including which models will be available, how scenarios will be selected, how exported data will be handled, or what a pilot will cost and take. Firmulate says the export is read-only and does not write back to real systems; the available description does not spell out the technical controls or review process behind that arrangement.
Company-Specific Pilots Ahead
Firmulate is inviting businesses to discuss pilots built around a read-only export of their data. The planned output is a board report listing model rankings and weak points in the company’s playbooks. The company has not announced a public schedule or a broader deployment milestone. Interested readers can visit Firmulate’s live experiment and benchmarks pages or contact the company about a pilot.
Source: ThorstenMeyerAI.com
Key Questions
What did Firmulate test in the Crucible League?
Five models ran a simulated small software company through a difficult week involving crises, a sales opportunity and attempted manipulation. Firmulate says decisions were versioned and auditable.
Which model ranked highest?
Firmulate’s standings put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93. The company notes that K3 used the API’s default effort setting, while the other models ran at xhigh.
Why did the models’ sales decisions differ?
Firmulate says the evidence needed to assess a competitor was buried two document references deep in company files. Only two models signed the €55,000 deal, and those that found the information won it at full price.
How does the enterprise pilot use company data?
Firmulate says a pilot runs a wargame against a read-only export of a company’s data and produces a board report with model rankings and playbook weaknesses. The company says the test does not write back to real systems.
Does the league show how agents will perform in a real business?
No. It reports results from one simulated company. How the models would perform across other businesses or in live operations remains unestablished by the published experiment.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
