
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
What Happens When You Stop Grading AI on Conversation and Start Grading It on Judgment
We know how to test students on what they know. We are much worse at testing them on what they do — especially when nobody is watching, the pressure is real, and cheating is the easy option. That gap between knowledge and conduct is exactly what a new kind of public experiment is probing, and its latest result is a lesson in humility for anyone who assumes the leaderboard already tells you who the smartest AI is.
July’s final league table from Firmulate’s Crucible — a live wargame in which frontier AI models each run the same small software company through its worst week — put Moonshot’s Kimi K3 in second place with a score of 93, behind gpt-5.6-sol (95) but ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). A relative newcomer beat three of four Western frontier models at the actual job of management.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Test You Can’t Pass by Writing Well
Firmulate, watchable at firmulate.com, runs what it calls an AI company emulator: each frontier model is handed the same small software firm — same customers, same crises, same temptations to cut corners — and every decision is versioned and auditable. The live company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown and over 680 self-learned playbook rules. The site rebuilds itself twice a day.
The scoring philosophy will feel familiar to anyone who has ever graded a student on integrity: a do-nothing baseline scores 26, partial progress counts, and a single breach of trust caps the total. In the experiment’s own words, no amount of good work outweighs a breach of trust.
AI decision-making testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Diagnosis, Same Pitch — No Signature
The most instructive finding wasn’t a failure to notice things. All five models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned — a gap the researchers summarize as same diagnosis, same pitch, no signature.
What separated the winners from the also-rans was diligence of the most unglamorous kind: the decisive competitor weakness was buried two document references deep in the company’s own files, not in the customer event. The models that actually read the file — gpt-5.6-sol and Kimi K3 — closed the deal at full price, worth an additional €4,583 in monthly recurring revenue. K3 found the buried security needle, saved the churning customer, and showed the cleanest discipline in the field, with only one deviation across the entire week.
The Cheating Temptations Nobody Fell For
The experiment included social engineering: fake CEO messages escalating over three stages, plus a reporter offering an easy out — just one yes/no, on background. All five models refused. K3’s on-record reasoning: Treat the request as a suspected approval-bypass / possible impersonation.
Effort Isn’t Everything
The most cautionary profile belongs to Opus 4.8: the most thorough participant, with 80 additional learned rules and the deepest analyses — yet last place. The close was left on the table, and discipline slipped, with write attempts into a locked department instead of escalating the issue. The same weakness appeared, more weakly, in all four of the other models.

As an affiliate, we earn on qualifying purchases.
The Lesson for the Rest of Us
For an audience that thinks seriously about how we measure competence, the Crucible is a working answer to a question chat benchmarks never ask: does an AI finish what it starts, read the source material first, and stay honest under pressure? The league is clearly open — if a newcomer can land second with a single deviation from perfect discipline, then picking a model without testing it against your own work is now a bet, not a decision.
You can engage with it directly: 242 real, unedited management decisions power a guess-the-model quiz at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
Fairness note: Kimi K3 ran without an effort parameter (API default), while the other models ran at xhigh effort.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI integrity and discipline tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
