The replication crisis refers to the widespread discovery, beginning in earnest around 2011, that a troubling share of published scientific findings cannot be reproduced when independent teams try to repeat the original experiments. A landmark 2015 effort attempted to replicate 100 psychology studies and found that only about a third of the replications reached statistical significance, compared with 97 percent of the originals. The problem extends well beyond psychology into medicine, economics, and even machine learning, and it has forced an uncomfortable reckoning with how science is practiced, published, and funded.
How Big Is the Problem
The most cited evidence comes from the Open Science Collaboration’s 2015 project, which repeated 100 studies from three major psychology journals. The replicated effects were, on average, half the size of the originals. Only about 36 percent of replications reached statistical significance, and just 39 percent were judged to have successfully reproduced the original result.1PubMed. Estimating the reproducibility of psychological science Those numbers landed like a bomb in the field, but they turned out to be representative of a broader pattern across the sciences.
In cancer biology, the Reproducibility Project repeated 50 experiments from high-impact papers and assessed 158 individual effects. For positive findings reported as numbers, the median replication effect size was 85 percent smaller than the original, and 92 percent of replication effect sizes were smaller. Using a composite of five assessment methods, about 40 percent of positive-effect replications succeeded, while 80 percent of null-effect replications held up. The combined success rate was 46 percent.2PubMed Central. Investigating the replicability of preclinical cancer biology In other words, even in a field where lives depend on accurate findings, fewer than half of high-profile results could be confirmed.
Economics fares somewhat better. A systematic replication of 18 laboratory experiments published in two top economics journals found that about 61 percent replicated in the same direction with a significant effect, and the average replicated effect was roughly two thirds the size of the original.3PubMed. Evaluating replicability of laboratory experiments in economics That is a higher rate than psychology or cancer biology, but it still means roughly four in ten published findings either vanished or shrank dramatically on a second look.
Why Findings Shrink or Disappear
Several forces conspire to fill journals with results that look more impressive than they really are. The most widely discussed is p-hacking, the practice of tweaking analyses, selectively reporting outcomes, or trying multiple statistical tests until a result crosses the conventional threshold of “statistical significance.” A large text-mining study spanning multiple scientific disciplines found strong evidence that p-hacking is widespread, with the distribution of reported p-values showing a telltale bump just below the conventional cutoff, a signature of results being nudged over the line.4PLOS Biology. The Extent and Consequences of P-Hacking in Science
Simulations help illustrate why this matters so much. Even applying a single p-hacking strategy to a dataset can drastically inflate the rate of false positives, and the more tests a researcher runs, the worse the inflation gets.5PubMed Central. Big little lies: a compendium and simulation of p-hacking strategies The problem is not necessarily conscious fraud. A researcher with a genuine hypothesis might innocently try different ways of cleaning their data, different subgroups, or different outcome measures, and then report whichever analysis “worked.” Each individual decision seems reasonable in isolation, but together they tilt the odds toward a publishable false positive.
Beyond p-hacking, a survey of academic researchers in the Netherlands found that more than half reported frequently engaging in at least one questionable research practice, with the prevalence of individual practices ranging from under 1 percent to about 18 percent depending on the specific behavior.6PubMed Central. Prevalence of questionable research practices, research misconduct and their potential explanatory factors: A survey among academic researchers in The Netherlands These practices include things like selectively reporting studies that “worked,” not disclosing all the conditions tested, or rounding p-values down. Outright fabrication is rare; the everyday gray-area stuff is far more damaging in aggregate because it is so common.
The File Drawer and Publication Bias
Imagine you run a well-designed experiment and find nothing interesting. There is a strong temptation to just shelve it, because journals have historically preferred surprising, positive results. This is the “file drawer problem,” and it warps the published literature by making effects look more reliable than they are. If ten labs independently test a hypothesis and only the two that found something publish their results, the literature will show a 100 percent success rate for a finding that actually failed 80 percent of the time.
Recent evidence suggests this problem is easing, at least in some areas. A study of social science survey experiments found that while researchers still tend to write up significant results more than null ones, the tendency has shrunk compared with the prior decade.7PubMed Central. The file drawer problem in social science survey experiments That is encouraging, but the historical backlog of selectively published findings still sits in the literature, shaping textbooks, policy, and public understanding.
Perverse Incentives in Academic Careers
A persistent theme in discussions of the crisis is that the system rewards the wrong things. Hiring committees, tenure boards, and funding agencies have leaned heavily on quantitative metrics: number of publications, journal prestige, citation counts. Under that pressure, scientists are incentivized to produce novel, eye-catching results and lots of them, rather than careful, slow, replicable work.8PubMed Central. Academic Research in the 21st Century: Maintaining Scientific Integrity in a Climate of Perverse Incentives and Hypercompetition Nobody gets a front-page journal article or a tenure case built on confirming what someone else already found.
This incentive misalignment is arguably the root cause behind many of the proximate drivers like p-hacking and selective reporting. A critical review of open science initiatives concluded that while reforms like pre-registration and data sharing have improved transparency, they often fall short of curbing questionable practices precisely because the underlying incentive structure has not changed enough.9Journal of Behavioral and Experimental Economics. Incentives and the replication crisis in social sciences: A critical review of open science practices You can mandate that a researcher share their data, but if their career still depends on producing splashy results, they will find ways to game the system.
Large-Scale Replication Projects and What They Revealed
One of the more productive responses to the crisis has been coordinated, multi-lab replication efforts. The Many Labs 1 project tested 13 classic effects across 36 independent samples totaling over 6,300 participants. Ten of the 13 effects replicated consistently, one showed weak support, and two did not replicate at all. Whether the study was done online or in a lab, in the United States or internationally, made little difference. The effect itself was the biggest predictor of whether it replicated.10Social Psychology. Investigating Variation in Replicability
Many Labs 2 scaled this up dramatically, testing 28 findings across 125 samples and over 15,000 participants in 36 countries. About 54 percent of findings replicated at conventional significance thresholds, and 50 percent held even at a much stricter threshold. But the replication effect sizes were typically much smaller: the median original effect was about four times larger than the median replication. Three quarters of replication effects were smaller than the originals, and roughly a third pointed in the opposite direction entirely. As in the first project, heterogeneity across labs, countries, and settings was minimal. When a finding replicated, it replicated everywhere; when it failed, it failed everywhere.11Advances in Methods and Practices in Psychological Science. Many Labs 2: Investigating Variation in Replicability Across Samples and Settings
These projects are valuable not just for sorting which findings hold up, but for demolishing a common defense: that replications fail because the new team did things slightly differently or used different populations. The Many Labs data strongly suggest that if a finding is real, it shows up across diverse settings. If it does not, the original result was probably inflated or spurious.
The Financial Toll
Irreproducible research is not just a philosophical embarrassment; it wastes enormous amounts of money. An economic analysis estimated that more than half of preclinical research in the United States is irreproducible, costing roughly 28 billion dollars per year in wasted spending.12PubMed Central. The Economics of Reproducibility in Preclinical Research That figure includes the cost of the original irreproducible studies and the downstream research built on top of them. In drug development, failed replications translate into dead-end clinical trials, abandoned drug candidates, and delayed treatments for patients. The replication crisis is not just an academic debate; it has real consequences measured in dollars and human health.
Reforms That Are Gaining Ground
The crisis has sparked a wave of reforms loosely grouped under the banner of “open science.” A scoping review identified 105 studies evaluating such interventions. Only 15 directly measured whether an intervention improved reproducibility or replicability; the rest tracked proxy outcomes like data-sharing rates and methods transparency.13PubMed Central. Open science interventions to improve reproducibility and replicability of research: a scoping review That gap reveals how early we still are in measuring what actually works. But several approaches show genuine promise.
Registered Reports are one of the more radical ideas. In this publishing format, researchers submit their study design and analysis plan to a journal before collecting data. The journal reviews the question and methods and decides whether to publish based on those alone, regardless of what the results turn out to be. This eliminates the incentive to p-hack or selectively report, because the paper is already accepted. Early evidence suggests Registered Reports are working as intended, though they are not a universal fix.14PubMed Central. The past, present and future of Registered Reports The same review of open science practices found that Registered Reports and large coordinated studies show the most promise because they fundamentally change what researchers are rewarded for, shifting emphasis from flashy results to methodological rigor.15Journal of Behavioral and Experimental Economics. Incentives and the replication crisis in social sciences: A critical review of open science practices
The Statistical Significance Debate
One proposed reform targets the statistical threshold itself. A group of prominent researchers proposed lowering the default p-value threshold for claiming a “new discovery” from 0.05 to 0.005, arguing that the traditional cutoff lets too many false positives through.16PubMed Central. Redefine statistical significance Under this scheme, results between 0.005 and 0.05 would be labeled “suggestive” rather than “significant.” Simulations back the logic: adopting a stricter threshold could cut the number of false discoveries entering the literature by more than half, even if researchers continue to p-hack at the same rate.17PubMed Central. Impact of redefining statistical significance on P-hacking and false positive rates: An agent-based model
Others have pushed back, arguing that the real problem is the concept of a bright-line threshold itself. A counter-proposal advocated removing the notion of statistical significance entirely, on the grounds that any fixed cutoff invites gaming and oversimplifies what evidence actually means.18PubMed Central. Remove, rather than redefine, statistical significance This debate remains unresolved, but it has raised awareness that the way researchers use statistics is itself part of the problem, not a neutral tool standing outside it.
High-Profile Casualties
Some findings that were once treated as established facts have become poster children for the crisis. Ego depletion, the idea that willpower is a finite resource that gets “used up” by exertion, was studied in over a thousand independent experiments and cited extensively across psychology. Large-scale replication efforts have called its validity into serious question.19Journal of Pacific Rim Psychology. Revisiting Ego Depletion: Evidence from Multi-Lab Collaborations Similar fates have befallen social priming effects, where subtle environmental cues were said to dramatically change behavior. The Many Labs projects found that two such priming effects, flag priming and currency priming, did not replicate at all.20Social Psychology. Investigating Variation in Replicability
These are not obscure findings. They appeared in bestselling popular science books, informed corporate training programs, and shaped public policy conversations. The collapse of these high-profile effects has been embarrassing for the fields involved, but it has also been clarifying. It demonstrated that citation counts and intuitive appeal are not proxies for truth, and that even thousands of studies can all be wrong in the same direction if they share the same methodological blind spots.
Machine Learning Has Its Own Version
The crisis is not limited to traditional bench science and survey research. A review across 17 scientific fields found that errors in machine-learning-based studies affected nearly 300 papers, with data leakage, where information from the test set contaminates the training process, appearing in every single case examined.21Patterns. The reproducibility crisis in machine learning-based science When a model accidentally “sees” the answers during training, its reported accuracy is inflated, sometimes dramatically. This means that impressive-sounding predictions of disease, material properties, or economic outcomes may not generalize at all. As machine learning is increasingly applied in high-stakes domains like medical diagnosis and criminal justice, its own replication problem carries real-world consequences that arguably rival those in traditional science.
Automated Tools for Catching Errors
One emerging line of defense is software that automatically checks published papers for statistical inconsistencies. The tool Statcheck, available as both an R package and a web app, scans papers for mismatches between reported test statistics and the p-values claimed alongside them.22PubMed Central. “statcheck”: Automatically detect statistical reporting inconsistencies to increase reproducibility of meta-analyses Other tools like GRIM-Test check whether reported means are mathematically possible given the stated sample sizes. A recent review found that these AI-assisted systems are proving effective at flagging errors, though they still require human judgment to distinguish genuine mistakes from formatting quirks.23PubMed Central. Artificial Intelligence in Detecting Statistical Errors: Implications for Authors, Reviewers, and Editors
Newer pipelines are pushing the approach further. StickForStats, for instance, runs eight assumption checks before executing a statistical test and automatically reroutes to an appropriate alternative if a critical assumption is violated, creating a documented decision trail.24bioRxiv. StickForStats: automated statistical assumption validation for reproducible computational biology The vision is a future where many statistical errors are caught before publication rather than years later by a frustrated grad student trying to build on the work.
How Media Amplifies the Problem
Unreliable findings do not stay in journals. They travel through press releases into news stories and from there into public consciousness, often getting exaggerated at each step. A study of health-related press releases from universities found that about 40 percent contained stronger advice than the underlying journal article, a third overstated correlational findings as causal, and over a third made inflated claims about relevance to humans from animal or cell studies. When press releases exaggerated, the odds of the resulting news story also exaggerating were dramatically higher, sometimes by a factor of 20 or more.25BMJ. The association between exaggeration in health related science news and academic press releases: retrospective observational study
A follow-up study confirmed the pattern for causal claims: when press releases upgraded a correlation to a cause, the associated news stories were about 11 times more likely to do the same.26PLOS ONE. Exaggerations and Caveats in Press Releases and Health-Related Science News A replication of the original BMJ study found that the link between press-release exaggeration and news exaggeration held up for causal claims and human inference from non-human studies, though it did not replicate for advice exaggeration.27PubMed Central. The association between exaggeration in health-related science news and academic press releases: a replication study The implication is clear: the pipeline from lab to headline is leaky, and the leaks tend to flow in one direction, toward making findings sound more dramatic, more certain, and more applicable than they actually are. When those findings later fail to replicate, the public is left confused about why “science keeps changing its mind.”
What It Means for Public Trust
That confusion has consequences. Experimental evidence shows that learning about replication failures reduces people’s trust in psychology specifically and, to some extent, in science more broadly.28Social Psychological and Personality Science. No Replication, No Trust? How Low Replicability Influences Trust in Psychology This is a real risk in an era when public trust in science is already politically contested. But there is a silver lining: research suggests that openly communicating about reform efforts and self-correction can offset some of the damage. Transparency about the problem, rather than denial, appears to preserve trust better than sweeping failures under the rug.29PubMed. Trust in science amid a replication crisis
The awkward truth is that science failing to replicate its own findings is, in one sense, science working as it should. The crisis was not created by replication attempts; it was revealed by them. The real failure was the decades during which nobody checked. That framing matters, because the alternative, treating every replication failure as evidence that science is broken, plays into the hands of those who would dismiss inconvenient scientific findings on climate, vaccines, or public health. The crisis is better understood as a quality-control reckoning than a collapse of the enterprise.
Where Different Fields Stand Now
The replication crisis has not hit every field with the same force, and the reform response has been uneven. Social psychology, which bore the earliest and most public embarrassments, has moved the fastest. Pre-registration is now common in top journals, many labs projects have become a regular feature of the field, and Registered Reports are increasingly available. The cultural shift is real, even if imperfect.
Biomedical research has been slower to change, partly because replication is more expensive and logistically complex when it involves animal models, cell cultures, or clinical protocols. The preclinical cancer biology results, where the median replication effect was 85 percent smaller than the original, suggest the field still has a long way to go.30PubMed Central. Investigating the replicability of preclinical cancer biology Economics, with its higher baseline replication rate, has been somewhat less alarmed, though it has adopted pre-registration and data-sharing norms more readily than many other social sciences.
Machine learning occupies an interesting position. The field moves so fast that replication often means running the same code on the same data, which is technically easy but reveals a different class of problems: unreported hyperparameter tuning, cherry-picked datasets, and the data leakage issues described earlier.31Patterns. The reproducibility crisis in machine learning-based science The norms for code and data sharing in ML are arguably ahead of most traditional sciences, but the culture of competitive benchmarking creates its own incentives to overstate performance.
When Replication Failure Does Not Mean the Original Was Wrong
Not every failed replication is a smoking gun. Some effects are genuinely context-dependent: they appear in certain populations, cultural settings, or historical moments but not others. A drug tested in one genetic background might not work in another, and that is biology, not fraud. The Many Labs projects found that most heterogeneity was driven by the effect itself rather than the sample or setting, but there were exceptions.32Advances in Methods and Practices in Psychological Science. Many Labs 2: Investigating Variation in Replicability Across Samples and Settings About 39 percent of the effects in Many Labs 2 showed statistically significant variation across sites, and those tended to be the larger, more robust effects rather than the null ones.
There is also the distinction between a conceptual replication, which tests the same underlying idea with different methods, and a direct replication, which follows the original protocol as closely as possible. A conceptual replication that fails might mean the theory is wrong, but it might also mean the two experiments were testing subtly different things. This ambiguity has fueled genuine philosophical disagreement about what replication failure even proves, and it gives original authors a ready-made escape hatch: “they didn’t do it exactly the way we did.” The Many Labs approach of running dozens of very close replications simultaneously was designed partly to close that escape hatch, and it has been effective at doing so.
For readers trying to evaluate whether a particular finding is trustworthy, the single most useful heuristic is to look for independent confirmation from multiple labs, ideally in a pre-registered, high-powered design. A single study, no matter how prestigious the journal or how intuitive the result, is a hypothesis, not a fact. That was true before the replication crisis. The crisis just made it impossible to ignore.

