The Referee Problem In An Age Of Cheap AI Production
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Referee Problem In An Age Of Cheap AI Production on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source essay argues that AI is making it far cheaper to produce work than to verify it, citing mathematical manuscripts, software pull requests and contract tasks. The examples point to growing pressure on human reviewers, though several figures come from vendors and the full evidence and methods are not provided in the source material.

An analysis published this week by ThorstenMeyerAI.com argues that AI systems are producing research, code and professional drafts faster than people can verify them, creating a growing human-review bottleneck. The essay points to OpenAI’s mathematical manuscripts, software pull-request data and a legal-workflow evaluation as examples, while some of its figures come from commercial vendors and cannot be independently assessed from the material provided.

The essay says OpenAI’s mathematical research programme generated 722 manuscripts grouped into 372 families after its model was given about 4,000 problems. It reports an average of roughly three hours of compute per result. Some work was formally checked in Lean, a proof assistant, while OpenAI cautioned that results without formal verification could contain issues. The source does not provide the manuscripts themselves or a full account of how the output was selected and assessed.

To illustrate the effort involved in review, the essay contrasts that volume with the response to an earlier result from the programme: a counterexample to an Erdős conjecture that it says was carefully checked by five leading mathematicians. That example illustrates the essay’s central distinction: automated tools can check whether a proof follows from a stated claim, but people still need to assess whether the claim is relevant, correctly framed and meaningful.

In software, the essay cites Faros AI data associating high-AI-adoption periods with 98% more merged pull requests and a 91% rise in review time. It also cites LinearB’s analysis of 8.1 million pull requests across 4,800 organisations: AI-generated changes reportedly waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. These are reported findings, not a controlled comparison established by the source material.

At a glance
analysisWhen: Published this week, according to the s…
The developmentA recent essay from ThorstenMeyerAI.com argues that rapid AI production is creating a review bottleneck across mathematics, software and professional workflows.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Is Becoming a Constraint

If the pattern described in the essay holds, organisations may be able to generate more work than they can responsibly use. Review capacity then becomes a practical limit on the value of AI: faster drafting does not automatically produce faster decisions, safer software or dependable research. Work can pile up awaiting inspection, or be approved with less scrutiny than its consequences warrant.

The essay identifies three possible responses: reviewers may approve work too quickly, give AI-generated submissions lower priority, or leave producers to decide which outputs deserve attention. Each carries a cost. Rubber-stamping can let errors through; suspicion can delay sound work; and relying on the producer’s own selection leaves less independent scrutiny. The source presents these as risks and examples, not as a measured account of how often each occurs across industries.

There is also a workforce concern. The essay argues that junior employees often develop judgement by doing the drafting, coding and proof work that AI tools can now take on. If entry-level workers mainly review machine-generated material, employers could weaken the experience pipeline that produces senior reviewers. That is a concern about future training and staffing, rather than a demonstrated outcome in the data supplied.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence Across Three Workflows

The analysis links three areas rather than reporting a single new scientific finding. In mathematics, it frames the issue as a gap between formal verification and expert adjudication. A proof checker can test whether a formal proof establishes its stated theorem; it cannot, by itself, decide whether the theorem addresses the intended question or merits attention.

For software, the source cites Faros AI and LinearB, both of which sell products related to software development or code review. That commercial interest does not invalidate their findings, but it is a reason to examine methods, definitions and comparison groups before treating the figures as representative of the whole industry. The essay also refers to a peer-reviewed 2026 study reporting that 61% of AI-agent pull requests received no human review before being merged or closed. The supplied material does not identify the study’s authors, sample or publication venue.

For professional work, the essay describes an OpenAI partnership with contract-software company Ironclad. It says GPT-6 Astra met an average of 55% of evaluation criteria across 11 contracting tasks, an improvement over a previous model. The source does not give the earlier score, explain the evaluation design or establish how well those tasks represent real legal work. The result should be read as a reported benchmark, not proof that the model is ready to handle contracts without oversight.

Amazon

software pull request review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Broad Is the Review Gap?

The supplied source does not include links, underlying datasets or complete methods for the cited studies and vendor analyses. It is therefore unclear how well their samples represent the wider software sector, whether the reported changes result from AI adoption itself, and how review time and acceptance were defined. The stated percentages should retain their original comparisons; they do not establish a universal effect for every team.

Several other details remain open. The source does not say how many of the 722 mathematical manuscripts were independently reviewed, how many were formally verified, or what happened to the remaining outputs. It also does not provide enough information to judge the Ironclad evaluation’s design or real-world performance. The essay’s claim that review costs may rise is a broader interpretation; the examples show reported patterns, not a single comprehensive measure of verification capacity across professions.

Amazon

formal verification tools for mathematics

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Methods and Oversight to Watch

The next useful evidence would include publication of the mathematical work and a clear accounting of which results received formal checks and independent expert review. For software, further reporting should show study methods, comparison periods and review outcomes across organisations, including whether AI-generated changes differ in complexity or risk. Those details would help distinguish a general review bottleneck from effects specific to the datasets cited.

For contract work, the evaluation criteria and task-level results would clarify what the reported 55% score means in practice. Meanwhile, organisations adopting AI can track how review queues, acceptance rates and defect rates change, and whether junior staff still get opportunities to build drafting and coding skills. The central question is not only how much AI can produce, but how institutions will verify that output and develop people able to take responsibility for it.

Amazon

AI research verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development described?

A recent analysis argues that AI is lowering the cost and time required to produce work faster than it is lowering the effort needed for people to verify that work.

Did OpenAI publish 722 mathematical manuscripts?

The supplied source says OpenAI’s programme produced 722 manuscripts from about 4,000 problems. It also says some results were formally checked in Lean and warns that unformalized results could have issues. The material provided does not independently document the manuscripts or their review status.

What do the software figures show?

The cited analyses report longer review waits or times for AI-generated work, alongside increased output in some teams. The figures come partly from commercial vendors, and the supplied source does not provide enough methodological detail to establish that they apply across software organisations.

Can AI check its own output?

Automated tools can check specific properties, such as whether a formal proof follows from a stated theorem or whether code passes defined tests. Those checks do not necessarily establish that the theorem, tests or task specification capture the intended problem. Human review remains relevant to those judgements and to accountability.

What remains uncertain?

The source does not provide full methods for the cited studies, the status of every mathematical manuscript, or enough detail to judge the contract-task evaluation. It also does not establish how widespread the reported review pressures are across industries.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Unlocking AI’s Secrets: Inside The Engine Room Of Twelve Machines

A detailed look into how twelve core AI models operate, revealing the mechanics behind chatbot responses and what this means for AI development.

What Makes a Turntable Setup Feel Ritualistic in the Best Way

Unlock the secrets to transforming your turntable setup into a mindful, ritualistic experience that elevates your listening journey—discover more inside.

Fable 5.1 Solves The Cyphral Distich, A 370-Year-old Cipher

Fable 5.1 has successfully solved the Cyphral Distich, a cipher dating back 370 years, marking a breakthrough in historical cryptography.

10 Things: Movie Night

NASA Science’s “10 Things: Movie Night” points viewers to free, on-demand NASA+ films and videos, including lunar imagery and music.