A Tale of Two Disciplines: What Separates the Sciences That Replicate from Those That Don't
When the Open Science Collaboration published its landmark 2015 effort to reproduce 100 psychological studies, roughly 60 percent failed to replicate their original results. The scientific community took note. Yet during that same period, experimental physics laboratories around the world were routinely confirming quantum chromodynamics predictions to extraordinary precision, and synthetic chemists were reproducing reaction yields with margins of error measured in fractions of a percentage point. The contrast was stark enough to demand an explanation—and that explanation, it turns out, runs far deeper than any single field's carelessness or rigor.
The divergence in replication success across scientific disciplines is not a story about good scientists versus bad ones. It is a story about the underlying architecture of different fields: how they define variables, how they incentivize publication, how they construct experimental conditions, and how much natural variability their subject matter introduces from the outset.
The Structural Advantages of the Physical Sciences
Physics and chemistry occupy a privileged position in the replication landscape, and much of that privilege is earned through the nature of their subject matter. When a chemist synthesizes a compound under specified temperature, pressure, and reagent conditions, those conditions are, in principle, infinitely reproducible. The molecule does not have a mood. It does not respond differently depending on whether it was prepared on a Tuesday or in a laboratory in Boston versus one in Bangalore. The system under study is closed, or can be made sufficiently close to closed that variability is manageable.
This is not merely philosophical. Physical constants do not drift. Atomic behavior is governed by laws that hold regardless of cultural context, historical moment, or the expectations of the researcher. When replication fails in physics, it is typically traceable to equipment calibration, measurement error, or procedural deviation—problems that are identifiable, correctable, and, crucially, not intrinsic to the phenomenon itself.
Chemistry benefits from a similar structural clarity, reinforced by decades of standardization in reporting. Reaction protocols in peer-reviewed chemistry journals are expected to include sufficient procedural detail that an independent laboratory can reproduce them without contacting the original authors. That norm did not emerge by accident; it was cultivated deliberately by journal editors and professional societies who recognized reproducibility as a prerequisite for scientific progress.
Why Biology and Psychology Face a Different Problem
The life sciences and behavioral sciences operate in a fundamentally different epistemic environment. Human subjects carry histories. Laboratory animals respond to handling stress, housing conditions, and the microbiome of their facility. Cell lines drift genetically across passages. Psychological constructs—anxiety, attention, implicit bias—resist the kind of operational precision that a melting point or a bond angle affords.
This is not an excuse; it is a diagnosis. When a social psychology experiment measures "priming effects" on behavior, the construct being measured is itself contested, context-sensitive, and potentially moderated by dozens of variables the original study never considered. The replication attempt is not simply repeating a procedure; it is re-entering a complex system that may have shifted in ways neither team can fully account for.
Biology faces its own version of this problem. The irreproducibility crisis in preclinical cancer research—brought into sharp relief by the Amgen and Bayer replication efforts of the early 2010s, which found that fewer than a quarter of landmark studies could be reproduced internally—reflects a field in which biological systems are heterogeneous, antibody reagents are poorly validated, and statistical thresholds have historically been treated as flexible rather than binding.
The Incentive Layer: How Publication Norms Amplify the Gap
Structural differences in subject matter do not fully account for the replication divide. Incentive structures play an equally important role, and here the gap between fields becomes a self-reinforcing cycle.
In fields where replication is difficult and novelty is prized, the publication incentive systematically selects for surprising findings. A psychology study demonstrating that a brief written exercise improves student performance by a statistically significant margin is publishable. The failed replication of that study, conducted with a larger and more diverse sample, frequently is not—or at least was not, until preregistration and registered reports began shifting that calculus. The result is a literature that is systematically skewed toward effects that are either inflated, context-dependent, or both.
Physics and chemistry are not immune to publication pressure, but their replication infrastructure is more deeply embedded in the reward system. High-energy physics, for instance, has long required independent confirmation before claiming discovery—the five-sigma threshold for announcing a new particle is not merely a statistical convention but a cultural norm that reflects how the field defines credibility. No such norm has historically governed the announcement of a new behavioral effect in social psychology.
Sample Sizes, Statistical Power, and the Arithmetic of Irreproducibility
One of the most tractable contributors to differential replication rates is statistical power—the probability that a study will detect a true effect if one exists. Underpowered studies, those with sample sizes too small to reliably detect the effect sizes they are measuring, produce results that are both noisier and more likely to represent false positives.
The physical sciences, dealing with phenomena that can often be measured repeatedly and precisely within a single experiment, naturally accumulate the statistical weight needed for reliable inference. A chemist measuring reaction yield across fifty trials within a single experiment is, in effect, operating with a sample size that most psychology studies never approach. The behavioral and biomedical sciences, where each data point represents a human participant or an animal subject, face cost and logistical constraints that have historically pushed researchers toward underpowered designs—and toward the p-value gaming that underpowered designs encourage.
The emergence of power analysis requirements in grant applications and the growing adoption of preregistration in psychology and medicine represent genuine progress on this front. But adoption remains uneven, and the backlog of underpowered studies already embedded in the literature continues to shape meta-analyses, clinical guidelines, and research agendas.
What the Replicating Fields Can Teach the Struggling Ones
The lesson from physics and chemistry is not that the behavioral and life sciences should become something they are not. Human behavior is genuinely more variable than electron spin; that variability is scientifically interesting, not merely inconvenient. But there are transferable principles worth examining.
Standardization of measurement instruments, independent replication as a condition of high-stakes claims, transparent reporting of null results, and the cultural normalization of verification as a scientific contribution rather than a lesser form of scholarship—these are not unique to the physical sciences. They are practices that any field can adopt, and that several corners of psychology, medicine, and biology are actively working to implement.
The fields that replicate well did not arrive at their current reliability by chance. They built it, over generations, through methodological discipline and cultural expectations that treated reproducibility as a prerequisite rather than an afterthought. The fields currently struggling with replication have the advantage of learning from that example—and the urgency of a credibility crisis to motivate the effort.