Algorithms at the Gate: The Promise and Peril of Artificial Intelligence in Peer Review
Photo: Our World in Data, CC BY 4.0, via Wikimedia Commons
Opinion | Technology & Innovation
There is something almost paradoxical about using artificial intelligence to evaluate the products of human intellectual labor. Peer review, at its best, is an act of disciplinary empathy—a trained scholar reading another scholar's work with the dual aims of understanding it on its own terms and assessing its contribution to a shared body of knowledge. Can a machine replicate that? Should it try? In 2024, these are no longer hypothetical questions. They are operational ones, being answered in real time by journal editors, publishers, and technology developers who may not always agree on the answers.
From Screening Tool to Quality Arbiter
The integration of AI into peer review workflows has proceeded along a spectrum. At the most modest end, natural language processing tools now assist editors in screening manuscripts for plagiarism, statistical irregularities, image manipulation, and compliance with formatting requirements. These applications are largely uncontroversial. Detecting duplicated text or flagging improbable data distributions is a task well-suited to computational methods, and offloading it to automated systems frees human reviewers to focus on substantive intellectual evaluation.
More ambitious deployments move further into the review process itself. Several major publishers—including Elsevier, Springer Nature, and a number of independent US-based journal operations—have begun piloting AI tools that analyze manuscript quality, predict reviewer recommendations, and match submissions to appropriate expert reviewers based on semantic analysis of the text. Some systems now generate structured reviewer reports that human editors can use as a baseline, editing and supplementing rather than composing from scratch.
Dr. Priya Sundarajan, an editor at a prominent biomedical journal based in Boston, describes the practical appeal of these tools candidly. "We receive several hundred submissions a month. The volume alone creates pressure to make faster decisions with less deliberation. AI screening gives us a way to handle that volume without simply lowering our standards—or at least, that's the aspiration."
The qualification at the end of her sentence is telling.
What AI Does Well—and What It Cannot
AI systems in peer review are genuinely effective at certain tasks. They can identify statistical anomalies that human reviewers frequently miss, particularly in fields like psychology and medicine where underpowered studies and selective reporting have contributed to well-documented replication crises. Tools such as Statcheck, which automatically verifies statistical calculations in manuscripts, have already been adopted widely and have identified errors—some innocent, some not—at rates that manual review rarely achieves.
Similarly, AI-assisted reviewer matching has the potential to reduce the chronic problem of reviewer fatigue. By analyzing the semantic content of a manuscript and cross-referencing it against a database of published research and reviewer profiles, these systems can identify qualified reviewers with greater precision and less reliance on editors' personal networks. For journals seeking to diversify their reviewer pools—a goal that aligns with broader equity imperatives in the field—this capability is genuinely valuable.
But the limitations are equally significant. Current large language models, however sophisticated, do not understand scientific claims in any meaningful sense. They process patterns in language. A manuscript that is stylistically polished, internally consistent, and methodologically conventional will perform well under AI evaluation even if its central hypothesis is flawed, its experimental design subtly compromised, or its conclusions overstated. Novelty, theoretical risk, and genuine intellectual contribution—the qualities that distinguish transformative research from competent incremental work—are precisely the things that AI systems are least equipped to recognize.
Dr. Marcus Webb, a computational linguist who has collaborated with several technology developers building peer review tools, puts it bluntly: "These systems are trained on the corpus of published science. That means they are, by design, biased toward work that looks like what has already been accepted. If you want to use AI to find the next paradigm shift, you are pointing the tool in exactly the wrong direction."
The Ethical Dimensions
The ethical concerns surrounding AI in peer review extend beyond technical limitations. Several issues deserve direct attention from professional organizations and their members.
Transparency and disclosure. When AI tools contribute to editorial decisions, authors have a legitimate interest in knowing this. Current practice is inconsistent. Some journals disclose the use of automated screening tools; others do not. PRR Society holds that meaningful transparency requires explicit disclosure of AI involvement at every stage of the review process, including the nature of the tools used and the weight given to their outputs.
Accountability gaps. When a human reviewer makes a flawed or biased recommendation, there are—at least in principle—mechanisms for appeal, correction, and accountability. When an algorithm produces a flawed output that shapes an editorial decision, accountability becomes diffuse. Who is responsible: the editor who relied on the tool, the developer who built it, or the publisher who licensed it? These questions are not yet resolved, and the absence of clear answers creates real risks for authors and for the integrity of the scholarly record.
The commodification of expertise. Perhaps the most philosophically significant concern is whether the integration of AI into peer review subtly devalues expert judgment by treating it as a scarce resource to be minimized rather than a professional contribution to be cultivated. If AI-assisted review becomes the norm, the incentive structures that motivate scholars to invest time and care in evaluating their peers' work may erode. The result could be a system that is faster and cheaper but shallower—one that processes manuscripts efficiently without genuinely engaging with them.
Perspectives from the Field
The technology developers building these tools tend to emphasize augmentation over replacement. "We are not trying to make editorial decisions," says one product lead at a US-based academic technology firm who asked not to be named. "We are trying to give editors better information faster so that they can make better decisions themselves."
That framing is reasonable, but it depends on editors having both the time and the institutional support to engage critically with AI outputs rather than simply ratifying them. In an environment where editorial workloads are already substantial and institutional resources for journal management are under pressure, the practical distance between augmentation and delegation may be narrower than developers acknowledge.
Researchers who have studied AI adoption in publishing contexts note that the organizational context matters enormously. Journals with strong editorial cultures, robust reviewer training programs, and genuine commitment to rigorous evaluation are better positioned to use AI tools responsibly than those deploying them primarily as cost-reduction measures.
A Path Forward
AI is not going away, nor should the research community wish it to. Used thoughtfully, these technologies can address real inefficiencies in peer review workflows and support—rather than supplant—the expert judgment that gives the process its legitimacy. But thoughtful use requires deliberate governance.
PRR Society recommends that member organizations adopt formal policies governing AI use in review processes, including mandatory disclosure to authors, regular audits of AI-assisted decision outcomes for accuracy and equity, and investment in reviewer training that prepares scholars to engage critically with algorithmically generated assessments rather than simply accepting them.
The integrity of peer review rests ultimately on the intellectual seriousness with which scholars engage each other's work. Algorithms can support that engagement. They cannot substitute for it. Keeping that distinction clear—and building it into the policies and practices of professional organizations—is the task before us in 2024 and beyond.