Psychology has a well-documented problem with replicability: when independent researchers try to reproduce published findings, the results frequently fail to hold up. The most influential test of this problem found that only about 36% of replicated studies produced statistically significant results, compared with 97% of the originals, and the average effect size shrank by half. That gap has prompted a decade of reform, self-examination, and heated debate about what went wrong and how to fix it.
The Landmark Test That Quantified the Problem
In 2015, the Open Science Collaboration published results from one of the largest replication efforts ever attempted. Researchers re-ran 100 studies that had been published in three major psychology journals, using high-powered designs and, when possible, the original materials. The numbers were stark: while nearly all of the original studies had reported statistically significant results, only about a third of the replications did. Roughly 39% of the replication attempts were judged to have successfully reproduced the original finding, and replication effect sizes were, on average, about half as large as those originally reported.1PubMed. Estimating the reproducibility of psychological science
Those numbers landed hard, but they also invited more questions than they answered. Did the failures reflect genuine flaws in the original work, or were some of them artifacts of how the replications were run? Were certain areas of psychology worse off than others? And what, concretely, was driving the problem?
Some Corners of Psychology Fare Better Than Others
Replicability is not uniformly poor across the discipline. A large-scale analysis covering two decades of psychology papers estimated replication scores for multiple subfields, and the differences were meaningful. Personality psychology scored highest, with a mean replication score of 0.55 on the study’s metric, followed by organizational psychology at 0.50. Cognitive psychology landed at 0.42, while social psychology brought up the rear at 0.37.2PubMed Central. A discipline-wide investigation of the replicability of Psychology papers over the past two decades
The pattern makes intuitive sense when you consider the types of phenomena each subfield studies. Social psychology often investigates effects that are sensitive to the social and cultural context in which the experiment takes place. A priming study run in a lab at a U.S. university in 2008 might behave differently when replicated in a different country a decade later, not because the original study was fraudulent, but because the effect depends on conditions that shift. An analysis of the 100 studies from the Open Science Collaboration’s project found that the contextual sensitivity of the original research topic predicted replication success even after adjusting for differences in statistical power and effect size.3PubMed Central. Contextual sensitivity in scientific reproducibility This doesn’t let poor methods off the hook, but it does mean that a failed replication isn’t always a clean indictment of the original finding.
Why So Many Findings Don’t Hold Up
Several forces converge to inflate the number of published findings that later prove difficult to reproduce. Low statistical power is near the top of the list. A study is “underpowered” when it doesn’t have enough participants to reliably detect the effect it’s looking for, which makes both false positives and wildly imprecise estimates more likely. An analysis of statistical power in psychology from 1975 to 2017 found that, assuming a medium-sized true effect, the average power was only about 59%. For small effects, power dropped to roughly 23%.4PLOS ONE. Are most published research findings false? Trends in statistical power, publication selection bias, and the false discovery rate in psychology (1975–2017) In other words, for the kinds of subtle effects that many psychology studies are chasing, the typical study had less than a one-in-four chance of finding the effect even if it were real. Low power doesn’t just produce false negatives; when combined with the pressure to publish significant results, it selects for fluky, inflated findings that sail past the significance threshold by chance.
A systematic review of studies on a well-known implicit learning task illustrates how this plays out in a specific literature. Across 73 studies and 181 statistical tests, underpowered awareness tests repeatedly failed to detect a real but modest effect, leading researchers to conclude falsely that learning in their experiments was unconscious.5PubMed Central. Underpowered samples, false negatives, and unconscious learning The field had built a body of “converging evidence” for unconscious learning, but the convergence was an artifact of consistently weak tests.
Then there are the practices that actively tip the scales. Researchers who fiddle with their analyses until they find a significant result, a practice known as p-hacking, can dramatically inflate the false-positive rate. An extensive literature review catalogued a dozen distinct p-hacking strategies and used simulations to show how each one inflates the odds of a false alarm.6PubMed Central. Big little lies: a compendium and simulation of p-hacking strategies These range from selectively dropping outliers to testing multiple dependent variables and reporting only the one that “worked.”
Publication bias amplifies the damage. Journals overwhelmingly publish positive, significant results and ignore null findings, which means the published literature is a skewed sample of the research that was actually conducted. One analysis argued that when too many studies in a literature all report positive results, that pattern itself is evidence of publication bias, because in a world of honestly reported studies, some should fail by chance. In the worst case, the entire published set of experiments on a topic could consist of false positives.7PubMed. Publication bias and the failure of replication in experimental psychology
Behind all of these is an incentive system that rewards novelty and productivity over accuracy. Researchers build careers by publishing in high-profile journals, and those journals want surprising, clean results. The pressure to maximize publication metrics can push scholars toward practices that look good on a CV but undermine the reliability of the science.8PubMed Central. The misalignment of incentives in academic publishing and implications for journal reform
Ego Depletion and the Power of a High-Profile Failure
Some replication failures have become almost symbolic. Ego depletion, the idea that self-control draws on a limited mental resource that gets used up like fuel in a tank, was one of the most cited concepts in social psychology. Earlier meta-analyses had found a medium-sized effect, but questions about bias in that literature led to a large registered replication effort. Twenty-three laboratories tested a standardized ego-depletion protocol with over 2,100 participants. The result was an effect size of essentially zero, with confidence intervals that comfortably included no effect at all.9PubMed. A Multilab Preregistered Replication of the Ego-Depletion Effect
Cases like ego depletion did something that abstract statistical arguments couldn’t: they gave the crisis a concrete face. Textbook chapters had been written around this concept. Popularizers had built talks and books on it. When the effect evaporated under rigorous testing, it became difficult to dismiss the broader problem as academic infighting.
What Counts as a Replication, and Why It Matters
One complication in interpreting all of this is that “replication” isn’t a single thing. A direct (or exact) replication tries to follow the original procedure as closely as possible. A conceptual replication tests the same underlying idea but with different methods or materials. Philosophers of science have debated whether the distinction is even useful. Some have argued that the concept of exact replication is inherently flawed because no two studies can truly be identical, while others have pushed back, defending the distinction as genuinely important for identifying what went wrong when results don’t hold up.10PubMed Central. Explicating Exact versus Conceptual Replication
The tension matters practically. When a direct replication fails, did it fail because the original finding was wrong, or because some unnoticed feature of the original context was doing the heavy lifting? As noted earlier, context-sensitive research topics replicate at lower rates even when the replication methods are sound. This doesn’t mean that every failed replication should be waved away as a “hidden moderator” problem, but it does mean that interpreting replication failures requires judgment, not just statistics.
Reforms Gaining Traction
The most promising structural reform is the registered report, a publishing format in which a study’s introduction and methods are peer-reviewed and provisionally accepted before any data are collected. The idea is simple: if a journal commits to publishing the study regardless of how the results turn out, publication bias evaporates. And early evidence suggests it works, at least partly. In a study where 353 researchers evaluated pairs of papers, registered reports outperformed standard papers on nearly every quality metric. The improvements were largest for methodological rigor and analytic rigor, and registered reports were statistically indistinguishable from standard papers on novelty and creativity, countering the worry that the format would produce boring science.11Nature Human Behaviour. Initial evidence of research quality of registered reports compared with the standard publishing model
Preregistration, a lighter-weight version of the same idea where researchers publicly log their hypotheses and analysis plans before collecting data, has become far more common in recent years. Its track record, however, is more mixed. A comparison of preregistered and non-preregistered psychology studies found that preregistered studies tended to be better powered and more impactful, but they did not clearly have lower rates of positive results, smaller effect sizes, or fewer statistical errors.12PubMed Central. Preregistration in practice: A comparison of preregistered and non-preregistered studies in psychology The most generous interpretation is that preregistration is a useful norm that hasn’t yet fully changed behavior; the less generous interpretation is that researchers can game it the same way they game other safeguards.
Open Data Helps, but Not as Much as You’d Think
Sharing raw data is widely championed as a pillar of open science, and many journals now encourage or require it. But making data available is not the same as making it usable. When researchers tried to reproduce the statistical results of 25 articles from the journal Psychological Science that had been awarded open-data badges, only 36% of the articles were fully reproducible without help from the original authors. Another 24% required author involvement to get the numbers to match, and 28% could not be fully reproduced even with author assistance.13PubMed Central. Analytic reproducibility in articles receiving open data badges at the journal Psychological Science: an observational study
A similar audit at the journal Cognition, which implemented a mandatory open-data policy, found that sharing rates increased substantially (from about 25% of articles including data statements to 78% after the policy), but analytic reproducibility remained imperfect. For 13 out of 35 articles with reusable data, at least one target value could not be reproduced even with author help.14PubMed Central. Data availability, reusability, and analytic reproducibility: evaluating the impact of a mandatory open data policy at the journal Cognition The main culprits were unclear descriptions of analytic procedures and minor reporting errors, not deliberate misconduct. Open data is a necessary condition for scrutiny, but it’s far from sufficient if researchers don’t also document what they did clearly enough for someone else to follow.
Multi-Lab Projects and Distributed Science
One of the more creative responses to the crisis has been the rise of multi-lab replication projects, where a large number of independent laboratories run the same study simultaneously. This approach solves several problems at once: it generates huge sample sizes, provides built-in variation in context and population, and makes it very difficult for any single lab’s quirks to explain the result. The Many Labs projects have been among the most visible of these efforts, with the fifth installment testing original findings across a median of about six and a half labs per study, with median total samples exceeding 1,200 participants.15Advances in Methods and Practices in Psychological Science. Many Labs 5: Testing Pre-Data-Collection Peer Review as an Intervention to Increase Replicability
The Psychological Science Accelerator, a distributed network of hundreds of laboratories worldwide, has formalized this model. Its mission is to produce studies with the kind of large, diverse samples that most individual labs simply cannot achieve, whether the goal is to replicate existing findings or to run new studies on a global scale.16PubMed Central. The Psychological Science Accelerator: Advancing Psychology through a Distributed Collaborative Network These networks also help address a persistent criticism of psychology as a field: that its findings are based on narrow, unrepresentative samples.
The WEIRD Problem
Most psychological research draws participants from Western, educated, industrialized, rich, and democratic societies, and disproportionately from the United States.17PubMed Central. Beyond Western, Educated, Industrial, Rich, and Democratic (WEIRD) Psychology: Measuring and Mapping Scales of Cultural and Psychological Distance This creates a generalizability problem that overlaps with, but is distinct from, the replicability problem. A finding might replicate perfectly in another U.S. university lab but fail to hold in a different cultural context, not because the original study was poorly done, but because the phenomenon itself is culturally specific. Conversely, a finding that fails to replicate in one culture might actually be robust and real, just bounded.
Multi-lab networks like the Psychological Science Accelerator are beginning to chip away at this by running studies across dozens of countries simultaneously. But the broader problem of who gets studied, and who gets to do the studying, is a structural issue that no single reform can fix overnight.
A Measurement Problem Underneath the Replication Problem
There’s a deeper concern that reforms focused on statistical methods and transparency may be treating the symptoms rather than the disease. One critique argues that psychology faces a “validation crisis” in addition to its replication crisis. Many of the questionnaires, scales, and behavioral tasks used in psychological research have never been rigorously validated, meaning researchers often can’t be sure their measures actually capture the construct they’re supposed to capture. Without valid measures, even perfectly replicable results may not tell you what you think they do.18Meta-Psychology. The Validation Crisis in Psychology
Rating scales, the workhorses of psychological measurement, are a particular target of this critique. Their widespread use obscures fundamental questions about how people translate internal states into numbered responses. A separate analysis concluded that improving data analysis alone is not enough to overcome psychology’s intertwined crises of replication, confidence, and validity without also rethinking how data are generated in the first place.19Social and Personality Psychology Compass. What’s wrong with rating scales? Psychology’s replication and confidence crisis cannot be solved without transparency in data generation
Detecting Problems After the Fact
A growing toolkit of statistical methods has emerged to identify suspicious patterns in published data. The p-curve technique, for example, uses the distribution of significant p-values to estimate whether a body of evidence reflects a true underlying effect or is the product of selective reporting. It can arrive at conclusions that directly contradict traditional meta-analyses, as demonstrated when it was applied to the “choice overload” literature, a popular idea that having too many options paralyzes decision-making.20PubMed. p-Curve and Effect Size: Correcting for Publication Bias Using Only Significant Results However, evaluations of these tools have found that while p-curve and similar methods work reasonably well in idealized settings, they can perform poorly in more realistic scenarios where the assumptions they rely on are violated.21PubMed. Adjusting for Publication Bias in Meta-Analysis: An Evaluation of Selection Methods and Some Cautionary Notes
Beyond bias detection, there are also tools designed to flag potential data fabrication. While outright fabrication is rare, estimates suggest it happens often enough to warrant concern. Statistical forensic methods have been developed to detect anomalies in datasets that are unlikely to arise from genuine data collection, creating at least some deterrent against fraud.22PubMed Central. Tools of the data detective: A review of statistical methods to detect data and result anomalies in psychology
Not Just Psychology’s Problem
It is tempting to treat the replication crisis as psychology’s unique embarrassment, and the field has certainly been the most publicly self-critical. But a mixed-methods analysis of commentaries on replication across psychology and the biomedical sciences found that the conversations in both fields share a common core: concerns about the effectiveness of traditional peer review, the need for greater transparency, and the distorting effects of academic incentive structures. The near-simultaneous emergence of these concerns across multiple disciplines suggests shared institutional forces rather than a flaw specific to psychological science.23SAGE Journals (Review of General Psychology). Psychology Exceptionalism and the Multiple Discovery of the Replication Crisis
The published psychological literature also contains a high rate of straightforward errors: citation errors, methodological mistakes, statistical blunders, and misinterpretations. A review of standard correction mechanisms in psychology journals found them to be limited, with no robust system for flagging and fixing errors after publication.24PubMed Central. Inaccuracy in the Scientific Record and Open Postpublication Critique Self-correction is supposed to be one of science’s core virtues, but in practice, correcting a published paper is slow, stigmatized, and often ignored.
Teaching the Next Generation
One area where progress has been uneven is education. Undergraduates studying psychology are learning the field’s content from textbooks that still feature findings known to be unreliable. An early effort to develop a one-hour lecture on the replication crisis for undergraduates showed that students could be taught about the issues and that doing so shifted their attitudes in productive ways.25Teaching of Psychology. How (and Whether) to Teach Undergraduates About the Replication Crisis in Psychological Science Yet a more recent assessment found that discussions of replication and open research are still largely absent from social psychology classrooms, and textbooks continue to feature nonreplicable phenomena without adequate caveats.26Edward Elgar Publishing. What have we learned from the replication crisis? Integrating open research into social psychology teaching
This gap matters because today’s undergraduates are tomorrow’s researchers, clinicians, and policy-makers. If they leave university believing that every textbook finding is settled science, the same problems recur. If they leave overly cynical, convinced that nothing in psychology can be trusted, they may dismiss genuinely robust findings along with the flimsy ones. Getting the framing right in teaching is arguably as important as any methodological reform aimed at researchers, because it shapes the expectations and norms an entire generation brings to the field.

