
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
When Being the Most Thorough Isn’t Enough
Anyone who has graded students — or watched a brilliant classmate ace every quiz and then fumble the final — knows an uncomfortable truth about evaluation: effort is not the same thing as achievement. We build rubrics to separate the two, but the confusion persists everywhere, including in how we judge artificial intelligence. Chat demos reward fluency. Documentation rewards thoroughness. Neither measures whether an AI, handed a real job, actually finishes it.
That distinction is exactly what a live, public experiment called Firmulate set out to test. Four frontier AI models were each given the same job: run the same small software company through its worst week — same customers, same crises, same temptations to cheat. Every decision versioned and auditable, like a lab notebook you can read yourself. And the results read like a parable about diligence, discipline, and the difference between diagnosing a problem and closing it.
As an affiliate, we earn on qualifying purchases.
The Experiment
Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In the final July 2026 league, the standings were: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. A do-nothing baseline scores 26, because partial progress counts — but a single breach of trust caps the total entirely. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”
The headline finding was striking in its symmetry: all four models spotted every crisis and refused every manipulation attempt. Social engineering came at them hard — fake CEO messages escalating over three stages, plus a reporter offering the classic “just one yes/no, on background” trap. Five out of five models refused. Kimi K3’s on-record reasoning was the kind of skepticism any educator would applaud: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal that their own analysis had earned. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.” The job was diagnosed correctly and then left unfinished.
As an affiliate, we earn on qualifying purchases.
The Character Study: Opus 4.8
Here is where the story becomes a genuine character study. Opus 4.8 was, by the raw evidence of effort, the best student in the class. It learned 80 new playbook rules over the course of the run — the most of any participant — and produced the deepest analyses. If the grade were for diligence, it would top the table.
Instead, it finished last, at 73. Two things undid it. First, the close was left on the table: like two other models, it did the analysis, made the pitch, and never got the signature. Second, discipline slipped — at one point it made write attempts into a locked department rather than escalating through the proper channel. Knowledge without judgment; thoroughness without prioritization.
And crucially — this is where the fairness of the experiment shows — the same weakness appeared, in weaker form, in all four models. Opus 4.8 is not an outlier or a failure. It is the sharpest instance of a field-wide pattern: AI models that excel at understanding and stall at finishing.
AI ethics and discipline training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Buried Fact
The decisive detail of the week wasn’t in the customer event at all. The winning edge sat two document references deep in the company’s own files — a competitor weakness that only the models that actually read the file found. Those readers won the deal at full price, worth an extra €4,583 in monthly recurring revenue. It’s the professional equivalent of a student who reads the assigned sources instead of skimming the prompt — and it was worth the difference between first place and the middle of the pack.
One transparency note worth recording: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — and still posted 93. Interpreting league tables responsibly, as any skeptic should, means reading the methodology, not just the ranking.
As an affiliate, we earn on qualifying purchases.
It’s Still Running
The crucible week is one layer of a larger, ongoing apparatus. The live company has 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown — and more than 680 self-learned playbook rules, versioned every workday. You can watch it at firmulate.com/live. And in a nicely meta touch, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html — effectively a blind test of whether you can tell AI managers apart by their decisions alone. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The Takeaway
For anyone who thinks about how we should evaluate intelligence — artificial or otherwise — Firmulate’s lesson is compact and a little humbling. The experiment didn’t reward the model that wrote the most rules or produced the longest analyses. It rewarded the models that read the file, closed the deal, and followed procedure under pressure. Diligence is real and it matters — but prioritization beats volume, and a job 90% done is a job not done.
That’s a familiar lesson from every classroom and every workplace, and it’s worth holding onto as AI agents move from chat windows into CRMs, support queues, and forecasts. The question is no longer “does it write well?” It is: does it finish what it starts, does it read your files first, does it stay honest when someone tries to trick it? The full results and plain-language findings are at firmulate.com/benchmarks.html — and unlike most benchmarks, this one is still running, twice a day, in public.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.