Peer Review in Psychology: Why Reviewers Barely Agree

Peer review is supposed to be science’s quality filter, but the psychology running beneath it is messy, subjective, and sometimes counterproductive. A meta-analysis of inter-rater reliability across journal peer reviews found that reviewers agree with each other at levels barely above chance, with an average agreement score so low it would be considered unacceptable in virtually any other professional evaluation context. The process is shaped by cognitive biases, status hierarchies, emotional reactions, and social incentives that have surprisingly little to do with the scientific merit of the work being judged. Understanding these psychological forces matters for anyone who publishes, reviews, or simply reads research.

Reviewers Barely Agree With Each Other

The single most striking finding about peer review psychology is how little consensus it produces. A large multilevel meta-analysis of inter-rater reliability across journal peer reviews found a mean agreement level of just 0.34 on standard reliability measures, with an even lower figure of 0.17 on Cohen’s Kappa, a stricter metric that adjusts for chance agreement.1PLoS ONE. A Reliability-Generalization Study of Journal Peer Reviews: A Multilevel Meta-Analysis of Inter-Rater Reliability and Its Determinants To put that in perspective, a Kappa of 0.17 means reviewers reading the same manuscript arrive at similar conclusions only slightly more often than two people flipping coins would. The same paper one reviewer praises for its rigor, another may criticize for its methods.

This inconsistency is not a quirk of one field or one journal. It shows up across disciplines and journal tiers. Reviewers have also been shown to overlook methodological flaws and statistical errors, avoid reporting suspected instances of fraud, and commonly reach a level of agreement barely exceeding what would be expected by chance.2PubMed Central. Reimagining peer review as an expert elicitation process The implication is unsettling: whether a paper gets published often depends less on its objective quality than on which two or three people happen to read it. This randomness is baked into the system, and most researchers know it from personal experience even if they rarely discuss it openly.

The Weight of a Famous Name

Prestige bias is one of the best-documented psychological effects in peer review. When reviewers know, or can guess, that a manuscript comes from a prestigious institution, their evaluations shift. A recent audit found that institutional affiliation is the dominant cue influencing quality scores: papers attributed to low-prestige affiliations received lower scores across every field tested, a penalty that survived rigorous statistical correction even at the editorial stage.3arXiv. Prestige over merit: An adapted audit of LLM bias in peer review In other words, the same manuscript is treated as better science when a Harvard logo sits in the header than when a lesser-known university’s name appears there.

This effect is amplified at the most selective journals. An analysis of over 110,000 manuscript submissions across five years at two elite general science journals found strong selective effects associated with higher institutional prestige, larger team size, and certain countries. The study also found that corresponding authors who are men had a small but significant advantage at one of the journals, while authors based in China faced a significant disadvantage.4Science Advances. Editorial and peer review dynamics at elite general science journals These associations were generally stronger at the editorial review stage, where editors decide whether to even send a paper out for review, than during peer review itself. That detail matters: the first gatekeeping decision, made quickly and often by a single person, is where prestige bias does the most damage.

Gender, Language, and Other Filters

Gender bias in peer review is a subject of intense debate, and the evidence is more nuanced than many assume. A large-scale study across 145 scholarly journals found that manuscripts written by women as solo authors or coauthored by women were actually treated somewhat more favorably by referees and editors, suggesting that peer review and editorial processes do not penalize manuscripts by women.5Science Advances. Peer review and gender bias: A study on 145 scholarly journals That finding may surprise people who assume gender bias runs uniformly against women in academia. The picture varies by field and journal, though, and the study measured how submitted manuscripts were treated once in the system, not whether systemic barriers earlier in the pipeline affect which manuscripts women submit and where.

Linguistic bias, by contrast, appears to be more consistently documented. An experimental study found that abstracts written in “non-standard English” were much more likely to be rated as poor quality compared to abstracts in “standard” international academic English, even when the underlying research content was controlled for. The effect was statistically significant, with non-standard English abstracts receiving meaningfully lower ratings.6Journal of English for Academic Purposes. Preliminary evidence of linguistic bias in academic reviewing For the millions of researchers worldwide who write in English as a second or third language, this bias adds an invisible hurdle. Reviewers may genuinely believe they are evaluating the science, but their perception of writing quality bleeds into their perception of research quality in ways they may not recognize.

What Harsh Reviews Do to People

Peer review is not just a system for filtering manuscripts. It is also an experience that leaves a mark on the people involved, and the psychological costs can be substantial. A study on the social and psychological costs of peer review found that researchers commonly experience professional failure, cognitive dissonance, and awareness of discriminatory practices within the review process.7Journal of Management Inquiry. The Social and Psychological Costs of Peer Review Getting a paper rejected is not like getting a bad grade on a test. For many academics, manuscripts represent months or years of work and are tightly bound up with professional identity. A dismissive review can feel like an attack on competence itself.

And some reviews are genuinely damaging beyond ordinary criticism. Research on inappropriate and hostile peer review has found that in its worst forms, flawed referee input and indifferent or misdirected journal leadership can damage the quality of published material, harm professional relationships, and derail careers.8Taylor & Francis Online. Dealing with inappropriate-, low-quality-, and other forms of challenging peer review “Reviewer 2,” the semi-joking archetype of the unreasonable referee, is a running joke in academic circles precisely because so many researchers have encountered a review that felt personal, contradictory, or motivated by something other than scientific rigor. The anonymity that protects reviewers from retaliation also protects reviewers who are careless, vindictive, or simply having a bad day.

Early-career researchers are especially vulnerable. A scathing review of a first or second paper can shape how a young scientist thinks about their own abilities for years, and some leave academia partly because of these experiences. The emotional labor of revising a manuscript in response to hostile feedback, while maintaining a deferential tone in a response letter, is a skill no graduate program explicitly teaches but every researcher must learn.

When the System Misses the Best Work

One of the more damaging psychological patterns in peer review is conservatism, the tendency to favor work that fits comfortably within established paradigms and to resist ideas that challenge them. A study of gatekeeping effectiveness at three major medical journals tracked over 800 manuscripts and found a troubling pattern: while papers with lower reviewer scores generally did receive fewer citations once published elsewhere, the journals had rejected many of their most-cited manuscripts, including all 14 of the most popular papers eventually published from the dataset. Of those 14 highly influential articles, 12 had been desk-rejected, meaning they were turned away before even reaching peer review.9PubMed Central. Measuring the effectiveness of scientific gatekeeping

The psychology behind this is straightforward. Reviewers and editors are experts in their fields, and expertise comes with assumptions about what constitutes valid methods, interesting questions, and plausible findings. A truly novel piece of research, by definition, pushes against those assumptions. It looks wrong. It does not fit the mental model that the reviewer has spent a career building. So the reviewer, acting in good faith, recommends rejection, and the editor, scanning hundreds of submissions a month, trusts that judgment. The paper finds a home elsewhere, accumulates citations, and the original journal quietly misses the boat. This does not mean peer review is useless at filtering quality. It does filter, and lower-scored papers do tend to be less influential. But the system has a blind spot at the top end, precisely where the stakes are highest.

Why People Review at All

Given how much effort it requires, the motivations behind peer reviewing are psychologically interesting in their own right. A qualitative study of reviewer motivations identified several overlapping drivers: contributing to science, personal scientific development, career advancement, personal satisfaction, and financial incentives, though the last of these is vanishingly rare since most journals do not pay reviewers.10PubMed Central. Motivations and barriers to engaging in peer review: a qualitative study For most reviewers, the work is genuinely prosocial: they believe in the system and want to support their community. But altruism alone does not sustain effort over time, and the barriers, including time pressure, lack of recognition, and the thanklessness of the task, eventually catch up.

Attempts to solve this with recognition programs have produced counterintuitive results. One large-scale study found that after receiving a peer review accolade award, reviewers’ subsequent motivation actually decreased. The explanation draws on diminishing marginal utility: the first award feels meaningful, but the prospect of a second one does not carry the same psychological weight. Researchers who were motivated by self-achievement or prosocial goals experienced a natural decline in reviewing effort after being recognized, partly because they felt they had already exceeded expectations.11Humanities and Social Sciences Communications. Can peer review accolade awards motivate reviewers? A large-scale quasi-natural experiment

A separate experiment tested whether non-monetary rewards could attract new reviewers and found the opposite of what was expected: making the reward contingent on performance reduced the number of scientists willing to review by about 60 percent compared to a no-reward setting, with the effect especially pronounced among highly productive researchers and those at private universities.12PubMed Central. Are non-monetary rewards effective in attracting peer reviewers? A natural experiment The psychological reading here is that introducing a performance-based reward reframed reviewing from a voluntary professional duty into a competitive evaluation, and many researchers found that framing unappealing. The intrinsic motivation that had sustained their reviewing was crowded out by the extrinsic incentive. This is a well-known phenomenon in motivation psychology, but seeing it play out in peer review specifically is a reminder that fixing the system with carrots and sticks is harder than it sounds.

Does Blinding Actually Work?

Double-blind review, where neither authors nor reviewers know each other’s identities, is the most commonly proposed structural fix for prestige and affiliation bias. The logic is obvious: if reviewers do not know who wrote the paper, they cannot be influenced by the author’s name or institution. In practice, though, the evidence for its superiority over other models is surprisingly thin.

An evolutionary game theory model simulating author and reviewer behavior under double-blind and open review systems found no reliable difference between the two systems in terms of incentivizing reviewer effort. Under some conditions, open review, where everyone knows everyone’s identity, performed comparably. The model also revealed a paradox: higher payoffs for good reviewing could lead to less author effort under open review, not more.13arXiv. Double blind vs. open review: an evolutionary game logit-simulating the behavior of authors and reviewers Blinding faces a practical problem too: in many fields, especially small ones, reviewers can often identify authors from the topic, dataset, writing style, or self-citations. True anonymity is difficult to achieve, and partial anonymity may not produce the benefits the system promises.

Open review has its own psychological dynamics. When reviewers know their names will be attached to their comments, they tend to be more polite and constructive, but they may also be more reluctant to deliver negative evaluations, especially of senior colleagues who might later review their own work or sit on hiring committees. The choice between blinding models is not just a procedural question. It is a question about which set of psychological biases you would rather deal with.

Structured Rubrics and Training

If reviewers are inconsistent and biased, can you train them to be better? The evidence is cautiously encouraging. A peer-review training program designed for a global community of researchers found that all participants who completed the evaluation agreed that the training helped them understand their role as reviewers, and most reported improvements in both their reviewing and their writing skills.14PubMed Central. Lessons learnt from a scientific peer-review training programme designed to support research capacity and professional development in a global community Forty-nine participants from six continents completed the program. That is a small sample, and self-reported improvement is not the same as measured improvement, but it points in a useful direction: reviewing is a skill, and like most skills, it can be taught.

Structured rubrics are another promising approach. Research on peer assessment in educational contexts has found that providing clear, detailed rubrics significantly improved the consistency and reliability of evaluations.15IGI Global Scientific Publishing. Evaluating Peer Contributions by Rubrics, Feedback Loops, and Fair Assessment in Online Courses Additional research confirmed that strategies such as maintaining anonymity, using structured rubrics, providing evaluative training, and incorporating supervisor oversight each produced significant improvements in fairness and accuracy across diverse assessment contexts.16AL-ISHLAH Jurnal Pendidikan. Bias in Peer Assessment: Challenges, Solutions, and Best Practices for Fair Student Evaluation Much of this research comes from educational peer assessment rather than journal peer review, but the psychological principles transfer: when evaluators have explicit criteria to anchor their judgments, they drift less toward personal taste, ideological preferences, and unconscious biases.

The challenge is implementation. Academic journals vary enormously in how much structure they provide. Some send reviewers a detailed evaluation form with specific questions about methodology, novelty, and presentation. Others send a manuscript and a blank text box. The journals that invest in structured guidance tend to get more consistent reviews, but building and maintaining those structures costs editorial time and money that many journals, especially smaller ones, do not have.

AI as Reviewer Assistant

Artificial intelligence tools are beginning to enter the peer review process, and the psychological dynamics of that shift are worth watching. A study of MetaWriter, an AI writing support tool for meta-reviewers (the senior reviewers who synthesize multiple reviews into a single recommendation), found that the tool significantly sped up the authoring process and improved the coverage of meta-reviews as rated by experts. The study used 32 participants, each writing meta-reviews with and without the tool.17Proceedings of the ACM on Human-Computer Interaction. MetaWriter: Exploring the Potential and Perils of AI Writing Support in Scientific Peer Review

The efficiency gains were clear, but participants raised concerns about trust, over-reliance, and agency. These are not trivial worries. If reviewers begin outsourcing their critical judgment to AI tools, the process risks losing the one thing it does well: applying deep human domain expertise to evaluate whether a piece of research makes sense. An AI tool that helps a meta-reviewer organize and articulate their own assessment is genuinely useful. An AI tool that generates the assessment itself, which the reviewer then rubber-stamps, defeats the purpose. The line between assistance and replacement is psychologically blurry, and researchers who are already stretched thin on time may gravitate toward the labor-saving option even when they know they should not.

There is also an emerging concern about AI-generated reviews being submitted by reviewers themselves. Several journals have reported detecting reviews that appear to have been written entirely by large language models, raising questions about whether the reviewer actually read the manuscript. The technology is new enough that norms and detection methods are still catching up, but the psychological incentive structure is clear: reviewing is time-consuming, unrewarded, and increasingly demanded of researchers who are already overcommitted. If a tool can produce a plausible-sounding review in minutes, the temptation to use it is real, even among conscientious scientists who would not dream of fabricating data.

How Power Dynamics Shape the Process

Peer review does not happen in a social vacuum. Reviewers are embedded in professional networks, career hierarchies, and competitive landscapes that shape how they engage with manuscripts. A junior researcher reviewing a paper from a senior figure in their field faces a different psychological situation than a tenured professor reviewing work from a graduate student. The junior reviewer may pull punches out of self-preservation; the senior reviewer may dismiss the student’s work as insufficiently sophisticated without fully engaging with it.

These dynamics intersect with the prestige effects documented earlier. The data from elite journal submissions showed that selective effects associated with institutional prestige and team size were stronger at the editorial stage than during peer review.18Science Advances. Editorial and peer review dynamics at elite general science journals Editors at top journals are themselves senior scientists embedded in the same status hierarchies. Their quick judgments about which papers deserve external review inevitably reflect their mental map of who does important work, and that map is drawn in part by institutional reputation and professional visibility rather than by the contents of any given manuscript.

Reciprocity also plays a quiet role. Researchers review for journals where they publish and interact repeatedly with the same editors and communities. A reviewer who has had a bad experience with a particular research group, or who is working on a competing project, brings that history into their evaluation whether they intend to or not. Formal conflict-of-interest policies exist, but they catch only the most obvious cases, like a reviewer evaluating their own collaborator’s paper. The subtler conflicts, professional rivalry, theoretical disagreement, personal dislike, are invisible to the editor and sometimes invisible to the reviewer themselves.

When Reviewers Do Not Know What They Do Not Know

One of the more uncomfortable findings in peer review research is that reviewers are often poor judges of their own limitations. The same literature that documents low agreement also suggests that reviewers routinely miss methodological flaws and statistical errors.19PubMed Central. Reimagining peer review as an expert elicitation process A reviewer selected for expertise in a paper’s topic may know the literature well but have no particular training in the statistical methods the paper uses. They review what they understand and skim what they do not, often without flagging the gap.

This is not a personal failing. It is a structural one. Journals typically ask one person to evaluate an entire manuscript, covering theory, methodology, analysis, and conclusions. No single reviewer is equally qualified to assess all of those dimensions. The result is that reviews tend to be strong on the aspects the reviewer knows best and superficial on everything else. A paper with a novel theory but flawed statistics might sail through review by a theorist and be rejected by a methodologist, and which outcome occurs is determined by the semi-random process of reviewer assignment.

Some reformers have proposed breaking manuscripts into components and sending each component to a specialist: a methods reviewer, a statistics reviewer, a domain expert. This approach is more expensive and logistically complex, but it addresses the fundamental mismatch between the breadth of a manuscript and the depth of any single reviewer’s expertise. A few journals have experimented with statistical review specifically, and the results have been promising, catching errors that traditional reviewers missed. Whether this model can scale remains an open question, but it represents a psychologically sound response to a well-documented limitation of the current system.