firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Education has long relied on tests that ask whether someone knows the answer. But in business, knowing what to do and actually doing it can be separated by a signature line. Firmulate’s live experiment puts AI models through a small company’s worst week to see whether their decisions hold up when the stakes are concrete.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A shared test, with real consequences inside the simulation

Each frontier model ran the same small software company through the same customers, crises and temptations. Decisions were versioned and auditable. Firmulate presents the experiment as a test of management quality: not just whether a model can explain a good response, but whether it carries that response through.

The final Crucible League, published in July 2026, placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Diagnosis was not the same as execution

All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The finding, in Firmulate’s words: “Same diagnosis, same pitch — no signature.” For anyone assessing AI at work, that is a useful distinction: a polished recommendation does not itself show that a system will complete the task.

The deal turned on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That result makes the test feel less like a quiz about general knowledge and more like a workplace exercise in finding relevant evidence and acting on it.

Trust and discipline under pressure

The social-engineering test used fake CEO messages escalating over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Performance still had rough edges. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and tried to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a qualification to keep in mind when comparing the results.

A company you can watch, then test against your own

The live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and workdays that are versioned. The experiment is watchable at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and invites visitors to guess which model made each one.

For businesses, the next step is a pilot against their own company. Firmulate says an enterprise can provide a read-only export, run crisis scenarios and receive a board report with a model ranking and weak points in its playbooks. Nothing writes back to real systems. That shifts the question from watching how models handle a simulated company to examining how they might respond to yours.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The experiment suggests that spotting a crisis, protecting trust and completing a commercially valuable decision are distinct tests. A company considering AI agents can move from observing the public experiment to trying a wargame against its own business. Explore a Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Kill-Switch-Proof: How to Build So Washington Can’t Take Your AI Stack Down

A detailed guide on how organizations can design AI stacks resistant to government shutdowns, emphasizing control and flexibility.

The SSD Squeeze: Why Storage Joined the Party

Enterprise and consumer SSD prices soar as AI’s storage needs strain supply, with industry tightness driven by wafer competition and rising demand.

The Orchestration Layer Arrives: What Anthropic’s Finance Agents Mean for Bloomberg, FactSet, and Wall Street

Anthropic launches ten finance agent templates with Claude integration, positioning as an orchestration layer over major data providers, challenging Bloomberg’s UI moat.

Build Custom Browser Tools Without Coding Knowledge

A new web app enables users without coding skills to create Chrome extensions via natural language prompts, simplifying browser automation.