Chasing the Threshold: The Enduring Grip of the P-Value on Scientific Judgment
In 1925, the British statistician Ronald Fisher proposed a rough heuristic for evaluating experimental data: if the probability of obtaining a result at least as extreme as the one observed, assuming no real effect, fell below one in twenty, the result might be worth a second look. Fisher was explicit that this was a pragmatic rule of thumb, not a philosophical criterion for truth. He did not intend it to serve as a universal boundary between knowledge and ignorance.
Nearly a century later, that boundary is precisely what it has become. Across biomedical science, psychology, nutrition research, and dozens of adjacent fields, a p-value below .05 remains the primary credential a finding must carry to be considered publishable, fundable, and real. The consequences of this categorical dependence — for the reliability of scientific literature, for the efficiency of research investment, and for the quality of decisions made downstream of published findings — have been extensively documented and persistently ignored.
A Metric Designed for a Different Purpose
To understand why the p-value endures, it helps to understand what it actually measures — and what it does not. The p-value quantifies the probability of observing data at least as extreme as those collected, conditional on a null hypothesis being true. It says nothing, directly, about the probability that a hypothesis is correct. It says nothing about the size of an effect, its practical significance, or the likelihood that a finding will replicate.
These limitations are not obscure. They are covered in introductory statistics courses and have been the subject of formal statements by the American Statistical Association, most recently in 2016 and again in a 2019 editorial that explicitly called for moving "beyond statistical significance" as an organizational principle for scientific inference. Major journals, including The American Statistician and Nature, have published prominent critiques. The critique is not new, not contested among statisticians, and not difficult to comprehend.
Yet in the laboratories and grant offices and editorial rooms where scientific decisions are made daily, the threshold holds.
The Economic Logic of a Bright Line
One reason for the p-value's durability is straightforwardly economic. Scientific publishing, grant review, and institutional evaluation all require efficient sorting mechanisms — ways of distinguishing, quickly and with apparent objectivity, between findings that warrant attention and those that do not. The .05 threshold provides exactly this: a single number that converts a continuous distribution of evidence into a binary judgment.
This binary is enormously convenient for actors throughout the research ecosystem. Journal editors can apply it as a preliminary filter. Grant reviewers can cite it as evidence of productivity. Department chairs can count significant findings as a proxy for research output. None of these applications requires understanding what the p-value actually measures; they require only that the threshold exist and that everyone agrees to treat it as meaningful.
The result is what statisticians have called p-hacking — the practice, sometimes deliberate and sometimes unconscious, of analyzing data in multiple ways until a p-value below .05 emerges, then reporting only that analysis. Surveys of researchers across disciplines suggest this practice is widespread. A 2011 study published in Psychological Science found that a majority of surveyed researchers admitted to at least one questionable analytical practice that inflates the probability of obtaining a significant result. The threshold does not merely measure scientific findings; it shapes the process by which they are produced.
Psychological Investment and the Significance Reflex
Beyond economic incentives, there is a psychological dimension to the p-value's persistence that deserves examination. Researchers who have spent months or years pursuing a hypothesis develop a strong motivational stake in its confirmation. The moment a p-value crosses .05 carries genuine emotional weight — a sense of vindication, of effort rewarded, of a question answered. This emotional valence is not trivial. It is reinforced by the social architecture of laboratory life, where significant results are celebrated and null results are received with disappointment.
The significance reflex — the automatic equation of p < .05 with "finding" and p ≥ .05 with "failure" — is acquired early in scientific training and reinforced continuously by the publication environment. Retraining researchers to think probabilistically, to treat effect sizes and confidence intervals as primary rather than supplementary, and to value well-conducted null results as genuine contributions, requires not just methodological instruction but a reorientation of the emotional rewards associated with scientific work.
Several researchers interviewed for this article described the difficulty of communicating non-significant findings to collaborators and mentors. "I had a result that was p equals .07," one clinical researcher recalled. "The effect size was clinically meaningful. The confidence interval was informative. But my co-investigator looked at the number and said, 'So we got nothing.' That's a deeply ingrained reflex."
What Alternatives Have Been Proposed
Leading statisticians have proposed several frameworks to supplement or replace the significance threshold, each with distinct methodological properties and practical tradeoffs.
Bayesian inference offers perhaps the most philosophically coherent alternative. Rather than asking how probable the data are under a null hypothesis, Bayesian analysis asks how much the data should update prior beliefs about a hypothesis — yielding a posterior probability that is, in principle, closer to what researchers actually want to know. The Bayes factor, which quantifies the relative support for competing hypotheses, provides a continuous measure of evidence that does not require a binary cutoff.
The practical obstacles to widespread Bayesian adoption are real. Bayesian analysis requires the specification of prior distributions — explicit quantifications of pre-existing belief — which introduce a form of subjectivity that frequentist methods formally avoid. In adversarial contexts such as regulatory review or litigation, this subjectivity can be exploited. Training requirements are also substantial; the majority of working researchers were trained in frequentist methods and would require significant retraining.
Estimation-based approaches, which foreground effect sizes and confidence intervals rather than significance tests, have gained traction in several disciplines, particularly in psychology following the replication crisis of the early 2010s. This approach does not require abandoning p-values entirely but treats them as one input among several rather than as the primary decision criterion. The New Statistics framework advocated by Geoff Cumming has been formally adopted by several journals in the behavioral sciences.
Pre-registration and registered reports address not the metric itself but the conditions of its application. By requiring researchers to specify hypotheses, sample sizes, and analytical plans before data collection begins, these mechanisms reduce the degrees of freedom available for p-hacking and separate confirmatory from exploratory analysis. Registered reports, in which journals commit to publication contingent on methodological quality rather than results, directly attack the publication bias that makes the .05 threshold so consequential.
Why Change Remains Slow
The proposals above are not new. Several have been formally endorsed by major statistical and scientific organizations. Some have been adopted at the journal level. None has achieved the kind of field-wide penetration that would displace the significance threshold as the dominant currency of scientific credibility.
The reason, ultimately, is coordination. The p-value's authority derives not from its methodological properties but from the fact that everyone uses it. Switching to a different framework imposes costs on early adopters — their findings become harder to compare with the existing literature, their manuscripts face editorial skepticism, their grant applications read as methodologically unconventional — while the benefits of a reformed system accrue only when adoption is widespread. This is a classic collective action problem, and collective action problems in large, decentralized institutions are notoriously resistant to voluntary resolution.
Formal intervention — through funding agency requirements, journal policy mandates, or accreditation standards for graduate statistical training — could break this coordination failure. The NIH has taken modest steps in this direction, including requirements for statistical power justification in grant applications. These steps have not, to date, materially altered the centrality of the significance threshold in how American science is conducted and evaluated.
The Cost of Inertia
The stakes of this methodological stasis extend well beyond academic dispute. Clinical trials that reach significance by narrow margins, and are then translated into treatment guidelines, may be doing so on the basis of findings that would not survive more rigorous evidentiary standards. Regulatory decisions informed by p-values at the margin of significance carry uncertainty that is not visible in the published record. Public health recommendations built on a literature shaped by publication bias inherit the distortions that bias produces.
The p-value is not the cause of these problems; it is the instrument through which deeper institutional failures are expressed. Reforming statistical practice without reforming the incentive structures that drive its misuse will accomplish little. But the reverse is also true: reforming incentives while leaving the significance threshold intact will simply redirect the same pressures through a different numerical channel.
Fisher's rule of thumb was never meant to carry this weight. A century of institutional accretion has made it load-bearing. Recognizing that it was never designed for the role it now plays is, at minimum, a necessary precondition for building something more adequate in its place.