Can AI Help Solve the Peer-Review Crisis? Promises and Pitfalls Are Emerging

Key Points
Scientific publishing has surged, with more than 8 million scientific publications in 2025, according to Science, putting additional pressure on an already strained peer-review system.
Editors needed an average of 4.5 reviewer invitations to secure one completed review in 2025, roughly twice the number required in 2018, according to Silverchair.
A global Frontiers survey of 1,645 researchers found that 53% of peer reviewers used AI tools in 2025, despite policies at many journals restricting such use.
Experiments suggest AI can be useful for routine checks, methodological feedback, consistency, and improving reviewer comments, but current systems remain unreliable for subjective judgments such as novelty, significance, fairness, and scientific importance.
Researchers increasingly see the most practical path as AI-assisted peer review with humans retaining responsibility for important decisions, rather than fully automated reviewing.
advertisement
The scientific publishing system is facing a difficult balancing act: researchers are producing more papers than ever, while editors are finding it increasingly difficult to recruit enough qualified reviewers to evaluate them. At the same time, the rapid spread of generative artificial intelligence is changing the workload itself, with large language models able to produce drafts, analyze manuscripts and assist with review tasks in minutes.
A new report published by Science on September 3, 2026, examines whether AI could help ease that pressure — and where relying on machines could create new problems of its own.
The scale of the challenge is significant. More than 8 million scientific publications were produced in 2025, according to Science, with the number of publications doubling in just five years. Some publishers have reported even faster growth in submissions, partly as generative AI makes it easier to produce manuscripts.
The result is a peer-review system under increasing pressure. Editors can spend weeks searching for willing experts, while authors may wait months or, in some cases, more than a year for publication decisions. Data from Silverchair's 2026 Future of Peer Review report shows that the number of reviewers in its dataset grew substantially between 2018 and 2025, yet acceptance rates declined enough that editors needed an average of 4.5 invitations to secure a single review in 2025, nearly double the level in 2018.
That pressure has helped make AI-assisted reviewing increasingly attractive.
A Frontiers survey of 1,645 researchers worldwide found that 53% of peer reviewers said they used AI tools for review tasks in 2025, up sharply from the previous year. Researchers reported using AI for activities including drafting review comments, summarizing material and, in some cases, examining methods and statistics.
Yet the rapid adoption has outpaced policy. Many journals still restrict or prohibit reviewers from using generative AI, particularly when confidential manuscripts are involved. That has created a gap between what researchers are already doing and what publishers formally permit.
Joachim Baumann, a postdoctoral researcher at Stanford University who studies the societal effects of AI, argues that the underlying peer-review problem is genuine but warns that automating scientific judgment too quickly could make matters worse.
The central question is therefore not simply whether AI can review a paper. It is which parts of peer review AI can perform reliably, and which parts still require human expertise.
Human peer review has weaknesses of its own. Reviewers can be rushed, inconsistent or overly subjective. They may focus on different aspects of a manuscript, provide uneven levels of detail or pay insufficient attention to routine errors. Research has also raised concerns about biases in how scientific work is evaluated, including advantages associated with prominent researchers.
For some scientists, those weaknesses make AI assistance attractive.
Agnieszka Swiatecka-Urban, a pediatric nephrologist and researcher at the University of Virginia, told Science that AI tools had helped her identify issues such as whether conclusions followed from results and whether numerical values remained consistent between text, tables and figures. She said the apparent objectivity was particularly impressive.
Research also suggests that AI-generated assessments can be more consistent than human ones. That consistency could be helpful when human reviewers are distracted by workload or fatigue. But it creates another concern: too much consistency could reduce intellectual diversity.
Baumann and colleagues have warned that if machine-generated judgments cluster too closely together, AI-assisted reviewing could encourage what they describe as an “intellectual monoculture,” potentially disadvantaging unconventional research that falls outside dominant expectations.
That concern highlights one of the most important differences between human and machine review. Human reviewers often disagree because they notice different things. One may focus on methodology, another on theoretical implications and another on missing literature. Ruth Ley, a microbiologist at the Max Planck Institute for Biology, argues that this diversity is one of the strengths of conventional peer review.
Novelty is an especially difficult test.
Determining whether a scientific paper represents a genuinely new contribution is not a simple fact-checking exercise. AI systems can search the literature, compare claims with earlier work and identify similarities, but they may not recognize a conceptual breakthrough that does not resemble existing research.
Researchers including Tom Hope have nevertheless been testing structured AI approaches to novelty assessment. In a study involving 182 ICLR 2025 submissions, an LLM-based system was designed to extract novelty claims, find related literature and compare the new work against previous research. The study reported 86.5% alignment with human reasoning and 75.3% agreement on novelty conclusions, demonstrating substantial potential while also illustrating that the task remains difficult.
For narrower and more objective tasks, the case for AI can be stronger.
A study led by researchers including James Zou of Stanford University used GPT-5 to inspect 2,500 papers accepted by three major machine-learning conferences from 2018 through 2025. The system was asked to identify specific, verifiable problems, including mathematical errors and inconsistencies between text, tables and figures.
The researchers reported that the AI identified at least one potentially relevant error in roughly one-third of the papers examined. Human experts subsequently confirmed 83% of the potential errors they reviewed. The work supports the idea that AI could be particularly valuable for tedious checks that human reviewers may not have enough time to conduct line by line.
That does not mean AI is simply another reviewer that can replace a scientist. Rather, it suggests a different role: an automated second set of eyes.
Another large experiment points in the same direction.
At ICLR 2025, researchers tested an AI system designed to critique reviewer comments rather than replace the reviewers themselves. The system provided feedback on issues such as vague language, misunderstandings and unprofessional comments.
In a randomized study involving more than 20,000 reviews, 27% of reviewers who received AI feedback changed their reviews. The researchers reported that the revised reviews were judged to be more informative, while the reviews became longer among those who updated them. The experiment provides evidence that AI may improve the quality of human reviewing when it is used as a feedback tool rather than as the final decision-maker.
The picture was not uniformly positive. Interviews conducted with reviewers at the same conference found that some regarded AI suggestions as redundant or overly prescriptive. In other words, even potentially useful AI feedback can fail to add value when it simply repeats what an experienced reviewer already knows.
There are also early examples of journals incorporating AI into actual editorial workflows.
NEJM AI, an affiliate of the New England Journal of Medicine focused on artificial intelligence in medicine, has been testing a seven-day fast-track process that combines AI analysis with detailed human editorial review. The approach is limited to papers that editors believe would already have a strong chance of succeeding through the conventional pathway.
The important safeguard is that human editors remain responsible for the final judgment. According to the Science report, Arjun Manrai of Harvard Medical School said AI has been particularly useful for identifying methodological problems and improving clarity, while editors continue to verify the system's claims against the paper.
That principle — trust but verify — may ultimately define the most realistic role for AI in peer review.
AI can process large quantities of information quickly. It can compare numbers, inspect formulas, search literature, identify inconsistencies and suggest ways to improve the clarity of a review. Such capabilities could save reviewers time and potentially reduce some of the burden contributing to publication delays.
But scientific peer review is not simply a checklist.
The hardest questions are often about whether a result matters, whether a methodology is appropriate for a particular scientific context, whether an interpretation is convincing, whether an unexpected idea deserves serious consideration and whether the evidence is strong enough to justify a major claim.
Those are judgments that depend on expertise, context and sometimes disagreement among specialists.
There is also the question of bias and gaming. Any AI system used in a high-stakes evaluation process could reproduce patterns in its training data or favor familiar approaches. Researchers have warned that automated review would need rigorous testing before being trusted with decisions that could influence careers, funding and publication.
Current research therefore points toward a middle ground rather than a fully automated future. AI appears strongest when performing narrow, measurable and repetitive tasks, while humans remain better positioned to provide context, interpret unusual results and take responsibility for difficult scientific judgments.
The transition is already underway, however. Researchers are using AI despite uneven policies, publishers are testing new systems, and experiments are beginning to measure when machine assistance actually improves review quality.
The question facing scientific publishing is no longer whether AI will appear in peer review. It already has.
The more consequential question is how carefully the scientific community will decide where AI belongs.
For publishers, reviewers and researchers, the coming years are likely to focus less on replacing peer review than on designing systems in which machines handle appropriate tasks while humans retain control over the decisions that matter most.
That distinction could determine whether AI becomes a useful tool for reducing pressure on peer review — or another source of problems in a system that is already struggling to keep pace with modern science.
Key Points Summary
Scientific publication volume reached more than 8 million papers in 2025, increasing pressure on peer review.
Editors required an average of 4.5 reviewer invitations for one completed review in 2025 in Silverchair's dataset.
53% of surveyed peer reviewers reported using AI tools in 2025, according to Frontiers.
AI is showing particular promise for fact checking, calculations, consistency checks and review improvement.
Novelty, significance, context and scientific judgment remain much harder problems, strengthening the case for human oversight.
What This Means
AI could help reduce some of the workload behind scientific peer review, particularly by taking on repetitive checks that reviewers may not have time to perform.
The people most directly affected are researchers, journal editors, peer reviewers and publishers, but the implications extend to anyone who relies on published scientific literature.
The key issue to watch is whether publishers develop systems that are transparent, independently evaluated and human-supervised, rather than treating AI-generated judgments as authoritative.
The most important shift may therefore be from AI replacing reviewers to AI assisting reviewers.
advertisement
Frequently Asked Questions [FAQ]
Can AI replace human peer reviewers?
Current evidence does not establish that AI can reliably replace human scientific reviewers. The strongest results so far are generally associated with specific tasks rather than complete autonomous judgment.
What can AI do well in peer review?
AI appears useful for checking calculations, identifying inconsistencies, examining methods, searching related literature, and improving the clarity and specificity of reviewer comments.
How many peer reviewers are using AI?
A Frontiers survey of 1,645 researchers found that 53% of peer reviewers reported using AI tools in 2025.
Why is peer review under pressure?
The growth in scientific publications and submissions has increased demand for reviewers, while editors are facing lower response rates and more difficulty securing completed reviews. Silverchair reported an average of 4.5 invitations per completed review in 2025, nearly twice the 2018 level.
Can AI judge whether a paper is novel?
Researchers are making progress, but novelty remains difficult because scientific originality can involve ideas that do not closely resemble existing literature. A study of 182 ICLR 2025 submissions found substantial agreement between a structured AI approach and human novelty assessments, but the researchers also acknowledged important limitations.
Does AI-assisted reviewing actually improve human reviews?
Evidence suggests it can. In a randomized ICLR experiment involving more than 20,000 reviews, 27% of reviewers receiving AI feedback updated their reviews, and blinded evaluation found the revised reviews more informative.
What is the biggest risk of AI peer review?
One major concern is that AI systems could introduce or amplify bias, while excessive consistency could reduce the diversity of scientific perspectives. Researchers have also warned about hallucinations, gaming and inappropriate automation of high-stakes decisions.
Sources:
- Science — Jeffrey Brainard, “Ready for the robot reviewers? As AI’s capabilities grow, researchers and publishers are exploring how it can support peer review—and where it still falls short,” September 3, 2026.
https://www.science.org/content/article/can-ai-help-solve-peer-review-crisis-here-are-its-promises-and-pitfalls
Additional Verified Sources
- Silverchair — 2026 Future of Peer Review Report
The Numbers Behind the Noise: What Our 2026 Peer Review Report Shows (Silverchair) - Frontiers — 2025 Global Survey of Researchers
Most peer reviewers now use AI, and publishing policy must keep pace (Frontiers) - Nature Machine Intelligence — AI Feedback and Peer Review
A large-scale randomized study of large language model feedback in peer review - arXiv — AI Detection of Errors in Published Papers
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis (arXiv) - arXiv — LLM-Assisted Evaluation of Scientific Novelty
Beyond “Not Novel Enough”: Enriching Scholarly Critique with LLM-Assisted Feedback (arXiv) - OpenReview — Concerns About Automating Peer Review
Stop Automating Peer Review Without Rigorous Evaluation
Thank you !