firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

What an AI does after finding the right answer

For readers interested in education and science, evaluating artificial intelligence presents a familiar problem: a polished answer is not necessarily evidence of sound reasoning, and sound reasoning is not necessarily evidence of competent action. Firmulate turns that distinction into a public experiment—and an unusually revealing quiz.

Its Crucible League placed frontier AI models in charge of the same small software company during its worst week. They faced the same customers, crises and temptations. Their decisions were preserved in unedited, auditable form, creating a controlled comparison of behavior rather than a collection of curated demonstrations.

The surprising result was not that some models missed the danger. Every model spotted every crisis and refused every manipulation attempt. The separation came later: only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A league table of management behavior

The final Crucible League standings, published in July 2026, put gpt-5.6-sol first and Kimi K3 close behind. The complete ranking was:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

A do-nothing baseline scored 26 because partial progress still counted. But the evaluation imposed a decisive ethical boundary: a single breach of trust capped the total, on the principle that “no amount of good work outweighs a breach of trust.” This matters because the experiment tested more than commercial performance. It asked whether an AI manager could remain trustworthy while facing pressure to take shortcuts.

The clue hidden beyond the immediate crisis

The decisive commercial fact was not contained in the customer event placed directly before the models. It sat two document references deep in the company’s own files. Models that followed that trail discovered a competitor weakness and won the deal at full price, worth +€4,583 MRR.

That detail gives the experiment its strongest educational lesson. Recognizing an urgent event is only the beginning of competent decision-making. A manager must also consult institutional knowledge, connect evidence across documents and convert the resulting insight into action. The losing behavior was therefore more subtle than ignorance: a model could understand the customer, construct the pitch and still fail to complete the sale.

Pressure revealed caution—and discipline gaps

The security test produced a more reassuringly consistent result. Fake CEO messages escalated over three stages, followed by a reporter attempting to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”

Yet caution alone did not guarantee a strong overall performance. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, but it finished last. It left the close on the table and attempted to write into a locked department instead of escalating the problem. A weaker version of that discipline failure appeared in all four other participants.

There is also an important comparison caveat. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its result should be read with that difference plainly visible rather than treated as a perfectly identical configuration.

A company readers can inspect

Firmulate’s live company contains 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, so the experiment is not merely a retrospective summary: its operational record remains watchable.

The most accessible route into that record is the Firmulate management-decision quiz. It draws on 242 real, unedited decisions and asks readers to identify which model made each one. The exercise turns model comparison into close reading: tone may offer clues, but the consequential differences concern follow-through, evidence gathering, escalation and restraint.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Beyond chatbot impressions

Firmulate’s findings suggest that frontier models can share the same diagnosis while displaying materially different management habits. Fluency does not reveal whether a model will read the relevant files, finish a negotiation, respect departmental boundaries or escalate when blocked.

That is why the quiz works as more than entertainment. It asks readers to examine authentic decisions before seeing the identity behind them, replacing brand expectations with observable behavior. Enterprises can extend the same idea by running the wargame against a read-only export of their own business; nothing writes back to real systems.

The central question is no longer simply whether an AI can produce a convincing answer. It is whether, under pressure, that AI can turn justified analysis into complete, disciplined and trustworthy work.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI ethical decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The early History of the Singular Value Decomposition (1993) [pdf]

Analysis of the 1993 publication on the origins of Singular Value Decomposition, highlighting confirmed facts and ongoing questions about its development.

Candor as a Moat: A Critical Reading of Dario Amodei and Anthropic

Analysis of Dario Amodei’s candid writings and their implications for AI regulation and industry power dynamics, culminating in recent US government actions.

Technology Operations Signal Monitoring Made Easy With OpenWrt One

OpenWrt One offers a streamlined way for small software companies to monitor platform updates like hardware routers, aiding faster decision-making.

Outcome-First Decisions: The Friction Is the Feature

A new decision-making approach prioritizes testing and evidence, helping businesses avoid costly mistakes by focusing on tangible outcomes.