AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Real Bottleneck In AI: Checking The Results on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source report describes a growing gap between AI-generated work and the human capacity to verify it, citing examples in mathematics, software and contract work. Its figures come from a mix of company statements, vendor analyses and a peer-reviewed study; the claims point to review as a constraint, but do not establish a single economy-wide measure of the problem.

A report on AI-assisted work argues that checking results is becoming a bottleneck as systems generate more mathematical manuscripts, software changes and professional drafts. The report cites OpenAI’s publication of 722 mathematical manuscripts and data on code review, but its evidence comes from sources with different methods and incentives, so it does not establish how large the problem is across the economy.

According to the source, OpenAI’s model was given about 4,000 mathematics problems and produced 722 manuscripts grouped into 372 families. Some results were formally checked in Lean, a proof assistant. OpenAI cautioned that unformalized results could have issues. The report contrasts that output with the careful examination of an earlier result from the same programme: a proposed counterexample to an Erdős conjecture was verified by five leading mathematicians.

The report also cites software-industry measurements. Faros AI said teams in periods of high AI adoption merged 98% more pull requests, while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. Those figures are vendor analyses, not a single independent measure of all software teams.

A 2026 peer-reviewed study cited by the source found that 61% of AI-agent pull requests received no human review before they were merged or closed. The report also describes an OpenAI partnership with contract-software company Ironclad, saying GPT-6 Astra averaged 55% of evaluation criteria across 11 tasks. The source presents that as an improvement over a prior model; the remaining criteria illustrate the need for review, but do not by themselves measure the severity of errors in deployed work.

At a glance
reportWhen: Described as developments from this wee…
The developmentA source report brings together recent examples and data to argue that verification, rather than generation, is becoming a constraint as AI produces more work.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Why Review Capacity Matters

If AI makes drafts and code cheaper to produce, organisations still need people who can decide whether an output is correct, useful and appropriate. That review can determine whether faster generation leads to usable work or simply moves effort downstream. The source’s examples suggest that some teams face longer queues, skipped reviews or a greater need to sort machine-generated work.

The report’s labor-market argument is an interpretation, not a measured forecast: it says experienced reviewers could become more valuable if demand for checking rises faster than the supply of qualified people. That possibility matters for hiring and training. Expert judgment depends on experience, and routine work has often provided the practice through which junior staff develop it.

There are also consequences for accountability. A person or institution may need to stand behind a contract, a bridge design, a scientific paper or a software release. Automated checks can catch defined errors, but responsibility for what was asked, what was missed and whether the result is fit for use remains a separate question.

Amazon

AI result verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Generated Work to Review

The source frames the issue through three fields rather than a single new benchmark. In mathematics, formal proof systems can verify that a proof follows from its stated assumptions, while expert review also considers whether the theorem is the intended one and whether the result is meaningful. The source says some of OpenAI’s manuscripts were checked in Lean; it does not give a complete breakdown of which results passed formal verification.

In software, tests and code-review tools can automate parts of checking, but their conclusions depend on what the tests cover and what reviewers inspect. Faros AI and LinearB sell products related to software engineering and review, a commercial interest the source itself flags. Their figures should be read with attention to their samples and definitions, even as the reported measures point in a similar direction.

The source’s contract example makes a related point about professional work: a model can meet many evaluation criteria without meeting all of them. The report argues that human review remains necessary where a missed rule or clause could matter. It does not provide details on the evaluation rubric, the kinds of errors found, or how often such work is used without revision.

Amazon

software code review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Available Evidence

The cited numbers do not share a common sample or definition. The source does not provide the underlying study details for every figure, and some software metrics come from companies that sell review-related tools. The available material therefore cannot show whether the reported patterns hold across sectors, organisation sizes or types of AI use.

Several specifics also remain unstated: how many of the 722 manuscripts were formally verified, what errors appeared in the contract evaluation, and how the review-time and acceptance measures were calculated. The source’s claim that checking is becoming a broad economic bottleneck is an argument drawn from these examples, not a quantified economy-wide finding.

It is also unclear whether AI review systems will reduce the human workload enough to change the balance. Automated systems may catch some errors, but the source argues that they cannot settle questions about intent, relevance or accountability. The scale of those limits will vary by task and by the consequences of an incorrect result.

Amazon

mathematical proof assistants

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Tracking Review and Training

The next useful evidence will be comparable, independently reported measures of how long AI-generated work takes to review, how often it is corrected, and what happens when human review is skipped. Follow-up results from software studies and clearer reporting on the mathematics and contract evaluations could show whether the patterns in the source persist.

For employers, the immediate question is how to use AI without removing the work through which less-experienced staff learn to assess quality. The source recommends protecting routes into expert judgment, though it provides no tested training plan. Organisations will also need to decide who is responsible for approving outputs and what level of checking is appropriate for each task.

Until better comparative evidence is available, the strongest conclusion is limited: AI can increase the volume of work produced, but the cited examples do not show that verification has become equally fast or automatic. How much that gap constrains adoption remains an open question.

Amazon

AI validation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the report mean by a verification bottleneck?

It means AI may produce work faster than people can check whether it is correct and fit for use. The source illustrates the concern with mathematics, software review and contract tasks.

Were all 722 mathematical manuscripts formally verified?

No such claim is made. The source says some results were checked in Lean and quotes OpenAI warning that unformalized results could have issues. It does not give a full count of formally checked manuscripts.

How reliable are the software review figures?

They come from analyses by Faros AI and LinearB, companies with commercial interests in software engineering or review tools, plus a peer-reviewed study cited by the source. Their methods and samples differ, so the figures should not be treated as one universal measure.

Does the report prove that AI will eliminate junior jobs?

No. It raises a concern that automating routine work could reduce opportunities for junior staff to build judgment, but it provides no employment forecast proving that outcome.

Can AI systems check other AI systems?

They can automate some checks, such as testing code against defined criteria or verifying formal proofs. The source argues that those checks may not determine whether the original task was framed correctly or who should be accountable for the result.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Anthropic Launches In-country Claude AI Inference In India Via AWS – The Times Of India

A Times of India headline reports Claude inference in India through AWS, but availability, regions and data-handling terms remain unconfirmed.

How to Choose Code Review Tools For Developers

A step-by-step guide to selecting, configuring, and rolling out code review tools that fit your team, workflow, and codebase.

Grok Trails Anthropic By Years, Musk Says, While Agentic AI Expands 100% Monthly

A Wccftech headline attributes two claims to Elon Musk, but the available material lacks the remarks, comparison details and growth data.

Small Businesses And AI: Turning Ideas Into Action

An OpenAI page is titled “Helping small businesses put AI to work,” but available details do not confirm an announcement, product or program.