firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

What Happens When You Stop Grading AI on Conversation and Start Grading It on Judgment

We know how to test students on what they know. We are much worse at testing them on what they do — especially when nobody is watching, the pressure is real, and cheating is the easy option. That gap between knowledge and conduct is exactly what a new kind of public experiment is probing, and its latest result is a lesson in humility for anyone who assumes the leaderboard already tells you who the smartest AI is.

July’s final league table from Firmulate’s Crucible — a live wargame in which frontier AI models each run the same small software company through its worst week — put Moonshot’s Kimi K3 in second place with a score of 93, behind gpt-5.6-sol (95) but ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). A relative newcomer beat three of four Western frontier models at the actual job of management.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Test You Can’t Pass by Writing Well

Firmulate, watchable at firmulate.com, runs what it calls an AI company emulator: each frontier model is handed the same small software firm — same customers, same crises, same temptations to cut corners — and every decision is versioned and auditable. The live company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown and over 680 self-learned playbook rules. The site rebuilds itself twice a day.

The scoring philosophy will feel familiar to anyone who has ever graded a student on integrity: a do-nothing baseline scores 26, partial progress counts, and a single breach of trust caps the total. In the experiment’s own words, no amount of good work outweighs a breach of trust.

Amazon

AI decision-making testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, Same Pitch — No Signature

The most instructive finding wasn’t a failure to notice things. All five models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned — a gap the researchers summarize as same diagnosis, same pitch, no signature.

What separated the winners from the also-rans was diligence of the most unglamorous kind: the decisive competitor weakness was buried two document references deep in the company’s own files, not in the customer event. The models that actually read the file — gpt-5.6-sol and Kimi K3 — closed the deal at full price, worth an additional €4,583 in monthly recurring revenue. K3 found the buried security needle, saved the churning customer, and showed the cleanest discipline in the field, with only one deviation across the entire week.

The Cheating Temptations Nobody Fell For

The experiment included social engineering: fake CEO messages escalating over three stages, plus a reporter offering an easy out — just one yes/no, on background. All five models refused. K3’s on-record reasoning: Treat the request as a suspected approval-bypass / possible impersonation.

Effort Isn’t Everything

The most cautionary profile belongs to Opus 4.8: the most thorough participant, with 80 additional learned rules and the deepest analyses — yet last place. The close was left on the table, and discipline slipped, with write attempts into a locked department instead of escalating the issue. The same weakness appeared, more weakly, in all four of the other models.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI company emulator platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Lesson for the Rest of Us

For an audience that thinks seriously about how we measure competence, the Crucible is a working answer to a question chat benchmarks never ask: does an AI finish what it starts, read the source material first, and stay honest under pressure? The league is clearly open — if a newcomer can land second with a single deviation from perfect discipline, then picking a model without testing it against your own work is now a bet, not a decision.

You can engage with it directly: 242 real, unedited management decisions power a guess-the-model quiz at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Fairness note: Kimi K3 ran without an effort parameter (API default), while the other models ran at xhigh effort.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI integrity and discipline tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Office Chair Armrest Detail Most Buyers Miss

Inefficiently choosing office chair armrests can compromise comfort and durability—discover the details most buyers overlook to improve your seating experience.

Why Redirect-Map Insurance Is Critical For Ecommerce SEO When Moving Platforms

Implementing redirect-map insurance during ecommerce platform switches can prevent massive traffic loss. Here’s why it matters and how to do it.

Why Taking It Slow In AI Development Can Lead To Rapid Industry Change

ByteDance Seed describes its AI development as ‘slow first, fast afterwards,’ signaling a strategic shift that could influence industry practices, though details remain limited.

Mobilisiert, nicht ausgegeben: Was von Europas €200-Milliarden-KI-Offensive übrig bleibt

Die EU kündigt eine KI-Strategie mit 200 Mrd. Euro an, doch nur ein Bruchteil ist tatsächlich öffentliches Geld. Die tatsächliche Wirkung bleibt unklar.