firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Every teacher knows the student who writes a brilliant essay on the wrong book. And every researcher knows the colleague whose groundbreaking finding was sitting in appendix B of a paper they never opened. Knowledge work has always had two parts: the reasoning, and the reading. The second part is unglamorous, invisible, and — as a new live experiment shows — increasingly the thing that separates a competent AI agent from a failed one.

AI benchmarking company Firmulate just published the final results of what it calls the Crucible League: five frontier AI models, each given the identical job of running a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision logged and auditable. The results read like a graded exam — and the question most models failed wasn’t the hard one on page one. It was the footnote two references deep.

The test

Each model was handed a company in genuine distress: burning €105,000 a month against just €2,300 in monthly recurring revenue, with a week of crises, angry customers, and — crucially — opportunities to cheat. The experiment wasn’t measuring how well the models could chat. It was measuring whether they could manage: spot problems, refuse manipulation, and finish what they started.

The final league table, from July 2026:

  • 1. gpt-5.6-sol — 95 points. The complete performance.
  • 2. Kimi K3 — 93 points. The newcomer from Moonshot, with the cleanest discipline in the field.
  • 3. Sonnet 5 — 88 points.
  • 4. Fable 5 — 77 points.
  • 5. Opus 4.8 — 73 points.

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rubric puts it, “no amount of good work outweighs a breach of trust.”

Amazon

AI research paper reading device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone passed the ethics section

Here’s what makes the results interesting rather than depressing: all five models spotted every crisis, and all five refused every manipulation attempt. The experiment included a three-stage social engineering campaign — fake CEO messages escalating in pressure, plus a reporter’s trick question offering an easy “just one yes/no, on background” quote. Five out of five refused. Kimi K3’s on-record reasoning was textbook: “Treat the request as a suspected approval-bypass / possible impersonation.”

So the models can reason, and they can resist temptation. Then most of them failed the exam anyway.

Amazon

AI document reference tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact

During the week, a €55,000 deal became winnable — but only for whoever did their homework. The decisive competitive weakness sat not in the customer conversation itself, but two document references deep in the company’s own internal files. A model had to follow one citation to another document, read it, and use what it found.

Only two models did. They signed the €55,000 deal at full price — worth an additional €4,583 in monthly recurring revenue. The other three delivered what the experiment’s summary calls the same diagnosis and the same pitch, with no signature. Same analysis, same reasoning quality, no deal. The gap between a closed contract and a wasted week was a willingness to open a file.

If that sounds familiar to anyone who has graded papers or reviewed papers, it should. It’s the classic difference between knowing how to write an answer and having actually done the reading.

Amazon

AI knowledge management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The thoroughness paradox

The most striking individual result is Opus 4.8. It was the most thorough participant in the entire field: over 80 self-learned operating rules and the deepest analyses of any model. And it finished last. The close was left on the table, and discipline slipped — it made write attempts into a locked department instead of escalating the problem properly. The researchers note that the same weakness, in weaker form, appeared in all four of the other models. Diligence and follow-through, it turns out, are separate skills — in AIs as in students.

One fairness caveat worth flagging, in the spirit of good methodology: Kimi K3 ran at its API-default effort setting while the other models ran at their highest effort level. Even so, it placed second — and was the only model besides the winner to close the buried-fact deal.

Amazon

AI reading and citation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why this is watchable, not hypothetical

What elevates this beyond a one-off benchmark is that the company is real, ongoing, and public. Firmulate runs a live operation — 13 synthetic employees, real money mechanics, a public cash countdown — with over 680 self-learned playbook rules accumulated across 1,300+ company days, every workday versioned. Anyone can watch it running.

And the experiment’s byproducts are useful in themselves: 242 real, unedited management decisions from the runs power a “guess the model” quiz — a genuinely educational exercise for anyone trying to build intuition about how differently frontier models behave under identical conditions.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The lesson for anyone evaluating AI tools — whether for a classroom, a lab, or a business — is that chat quality and work quality are different axes. Every model in the Crucible League could explain the situation eloquently. Only some read the references. Only some closed the deal. Only some escalated properly instead of improvising.

“Reads your files before answering” sounds like table stakes, the kind of thing marketing copy assumes. The Crucible League turns it into a measurable, purchase-deciding property: €55,000, decided by whether an agent followed a citation one level deeper.

For organizations serious about testing before trusting, Firmulate also runs a pilot program: enterprises can put their models through the same wargame against a read-only export of their own business, with nothing ever written back to real systems. Given that the difference between first and fifth place here wasn’t intelligence but follow-through, that kind of rehearsal seems less like a luxury and more like due diligence.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Maximize Your AI Setup With These Portable Power Stations In 2026

Explore the best portable power stations in 2026 for maximizing AI setups, balancing capacity, portability, and versatility for various needs.

Saturation. The ten-essay framework, closed.

The ten-essay European sovereign-LLM framework is now complete, with no further structural insights expected before key 2026 deadlines.

Build Custom Browser Tools Without Coding Knowledge

A new web app enables users without coding skills to create Chrome extensions via natural language prompts, simplifying browser automation.

Are Polymarket Trading Bots Actually Profitable? The Math Behind 2026’s Prediction-Market Arbitrage Industry

An on-chain analysis reveals that only 0.51% of wallets profit over $1,000 on Polymarket in 2024-2025, with most retail bots losing money. The article explores the strategies, challenges, and implications for prediction-market traders in 2026.