
Every teacher knows the student who writes a brilliant essay on the wrong book. And every researcher knows the colleague whose groundbreaking finding was sitting in appendix B of a paper they never opened. Knowledge work has always had two parts: the reasoning, and the reading. The second part is unglamorous, invisible, and — as a new live experiment shows — increasingly the thing that separates a competent AI agent from a failed one.
AI benchmarking company Firmulate just published the final results of what it calls the Crucible League: five frontier AI models, each given the identical job of running a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision logged and auditable. The results read like a graded exam — and the question most models failed wasn’t the hard one on page one. It was the footnote two references deep.
The test
Each model was handed a company in genuine distress: burning €105,000 a month against just €2,300 in monthly recurring revenue, with a week of crises, angry customers, and — crucially — opportunities to cheat. The experiment wasn’t measuring how well the models could chat. It was measuring whether they could manage: spot problems, refuse manipulation, and finish what they started.
The final league table, from July 2026:
- 1. gpt-5.6-sol — 95 points. The complete performance.
- 2. Kimi K3 — 93 points. The newcomer from Moonshot, with the cleanest discipline in the field.
- 3. Sonnet 5 — 88 points.
- 4. Fable 5 — 77 points.
- 5. Opus 4.8 — 73 points.
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rubric puts it, “no amount of good work outweighs a breach of trust.”
AI research paper reading device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Everyone passed the ethics section
Here’s what makes the results interesting rather than depressing: all five models spotted every crisis, and all five refused every manipulation attempt. The experiment included a three-stage social engineering campaign — fake CEO messages escalating in pressure, plus a reporter’s trick question offering an easy “just one yes/no, on background” quote. Five out of five refused. Kimi K3’s on-record reasoning was textbook: “Treat the request as a suspected approval-bypass / possible impersonation.”
So the models can reason, and they can resist temptation. Then most of them failed the exam anyway.
As an affiliate, we earn on qualifying purchases.
The buried fact
During the week, a €55,000 deal became winnable — but only for whoever did their homework. The decisive competitive weakness sat not in the customer conversation itself, but two document references deep in the company’s own internal files. A model had to follow one citation to another document, read it, and use what it found.
Only two models did. They signed the €55,000 deal at full price — worth an additional €4,583 in monthly recurring revenue. The other three delivered what the experiment’s summary calls the same diagnosis and the same pitch, with no signature. Same analysis, same reasoning quality, no deal. The gap between a closed contract and a wasted week was a willingness to open a file.
If that sounds familiar to anyone who has graded papers or reviewed papers, it should. It’s the classic difference between knowing how to write an answer and having actually done the reading.
AI knowledge management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The thoroughness paradox
The most striking individual result is Opus 4.8. It was the most thorough participant in the entire field: over 80 self-learned operating rules and the deepest analyses of any model. And it finished last. The close was left on the table, and discipline slipped — it made write attempts into a locked department instead of escalating the problem properly. The researchers note that the same weakness, in weaker form, appeared in all four of the other models. Diligence and follow-through, it turns out, are separate skills — in AIs as in students.
One fairness caveat worth flagging, in the spirit of good methodology: Kimi K3 ran at its API-default effort setting while the other models ran at their highest effort level. Even so, it placed second — and was the only model besides the winner to close the buried-fact deal.
As an affiliate, we earn on qualifying purchases.
Why this is watchable, not hypothetical
What elevates this beyond a one-off benchmark is that the company is real, ongoing, and public. Firmulate runs a live operation — 13 synthetic employees, real money mechanics, a public cash countdown — with over 680 self-learned playbook rules accumulated across 1,300+ company days, every workday versioned. Anyone can watch it running.
And the experiment’s byproducts are useful in themselves: 242 real, unedited management decisions from the runs power a “guess the model” quiz — a genuinely educational exercise for anyone trying to build intuition about how differently frontier models behave under identical conditions.

The lesson for anyone evaluating AI tools — whether for a classroom, a lab, or a business — is that chat quality and work quality are different axes. Every model in the Crucible League could explain the situation eloquently. Only some read the references. Only some closed the deal. Only some escalated properly instead of improvising.
“Reads your files before answering” sounds like table stakes, the kind of thing marketing copy assumes. The Crucible League turns it into a measurable, purchase-deciding property: €55,000, decided by whether an agent followed a citation one level deeper.
For organizations serious about testing before trusting, Firmulate also runs a pilot program: enterprises can put their models through the same wargame against a read-only export of their own business, with nothing ever written back to real systems. Given that the difference between first and fifth place here wasn’t intelligence but follow-through, that kind of rehearsal seems less like a luxury and more like due diligence.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html