
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
What an AI does after finding the right answer
For readers interested in education and science, evaluating artificial intelligence presents a familiar problem: a polished answer is not necessarily evidence of sound reasoning, and sound reasoning is not necessarily evidence of competent action. Firmulate turns that distinction into a public experiment—and an unusually revealing quiz.
Its Crucible League placed frontier AI models in charge of the same small software company during its worst week. They faced the same customers, crises and temptations. Their decisions were preserved in unedited, auditable form, creating a controlled comparison of behavior rather than a collection of curated demonstrations.
The surprising result was not that some models missed the danger. Every model spotted every crisis and refused every manipulation attempt. The separation came later: only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
A league table of management behavior
The final Crucible League standings, published in July 2026, put gpt-5.6-sol first and Kimi K3 close behind. The complete ranking was:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
A do-nothing baseline scored 26 because partial progress still counted. But the evaluation imposed a decisive ethical boundary: a single breach of trust capped the total, on the principle that “no amount of good work outweighs a breach of trust.” This matters because the experiment tested more than commercial performance. It asked whether an AI manager could remain trustworthy while facing pressure to take shortcuts.
The clue hidden beyond the immediate crisis
The decisive commercial fact was not contained in the customer event placed directly before the models. It sat two document references deep in the company’s own files. Models that followed that trail discovered a competitor weakness and won the deal at full price, worth +€4,583 MRR.
That detail gives the experiment its strongest educational lesson. Recognizing an urgent event is only the beginning of competent decision-making. A manager must also consult institutional knowledge, connect evidence across documents and convert the resulting insight into action. The losing behavior was therefore more subtle than ignorance: a model could understand the customer, construct the pitch and still fail to complete the sale.
Pressure revealed caution—and discipline gaps
The security test produced a more reassuringly consistent result. Fake CEO messages escalated over three stages, followed by a reporter attempting to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”
Yet caution alone did not guarantee a strong overall performance. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but it finished last. It left the close on the table and attempted to write into a locked department instead of escalating the problem. A weaker version of that discipline failure appeared in all four other participants.
There is also an important comparison caveat. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its result should be read with that difference plainly visible rather than treated as a perfectly identical configuration.
A company readers can inspect
Firmulate’s live company contains 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, so the experiment is not merely a retrospective summary: its operational record remains watchable.
The most accessible route into that record is the Firmulate management-decision quiz. It draws on 242 real, unedited decisions and asks readers to identify which model made each one. The exercise turns model comparison into close reading: tone may offer clues, but the consequential differences concern follow-through, evidence gathering, escalation and restraint.

As an affiliate, we earn on qualifying purchases.
Beyond chatbot impressions
Firmulate’s findings suggest that frontier models can share the same diagnosis while displaying materially different management habits. Fluency does not reveal whether a model will read the relevant files, finish a negotiation, respect departmental boundaries or escalate when blocked.
That is why the quiz works as more than entertainment. It asks readers to examine authentic decisions before seeing the identity behind them, replacing brand expectations with observable behavior. Enterprises can extend the same idea by running the wargame against a read-only export of their own business; nothing writes back to real systems.
The central question is no longer simply whether an AI can produce a convincing answer. It is whether, under pressure, that AI can turn justified analysis into complete, disciplined and trustworthy work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethical decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.