UIM Journal All articles
Research Methodology

Lost in the Archive: The Quiet Disappearance of Scientific Raw Data and What It Costs Us

UIM Journal
Lost in the Archive: The Quiet Disappearance of Scientific Raw Data and What It Costs Us

Photo: NOIRLab/NSF/AURA/T. Slovinský, CC BY 4.0, via Wikimedia Commons

Science's promise of reproducibility rests on a foundational assumption: that the data underlying published findings remain available for scrutiny, reanalysis, and replication. That assumption is increasingly difficult to defend. A growing body of evidence suggests that roughly 60 percent of research datasets become practically inaccessible within five years of a study's publication — not because of deliberate concealment, but because of systemic neglect, institutional inertia, and a persistent cultural indifference to long-term data stewardship. The consequences for scientific integrity are difficult to overstate.

The Scope of the Problem

The disappearance of scientific data is rarely dramatic. There is no single moment of loss, no obvious point of failure. Instead, datasets erode gradually — stored on personal hard drives that fail, housed on university servers that are decommissioned, or linked to institutional email addresses that expire when a researcher moves to a new position. A 2014 study published in Current Biology tracked datasets from papers published between 1991 and 2011 and found that the odds of data being available declined by approximately 17 percent with each passing year. More recent audits of repositories across multiple disciplines have confirmed that the trend has not meaningfully reversed.

The problem is compounded by the structure of academic incentives. Researchers are evaluated and promoted based on publication output, grant acquisition, and citation metrics — not on the quality or longevity of their data management practices. Depositing data in a well-organized, thoroughly documented repository requires time and effort that generates no immediate professional reward. In a competitive funding environment, that calculus consistently loses.

Where Data Goes to Die

The pathways to data inaccessibility are varied but predictable. Personal storage — external hard drives, desktop computers, and institution-specific network drives — accounts for a disproportionate share of losses. When a principal investigator retires, relocates, or leaves academia entirely, the datasets they managed often have no designated custodian. Graduate students who collected the data have moved on. The lab that generated the findings may no longer exist.

Link rot compounds the problem for datasets nominally deposited in online repositories. Supplementary materials hosted on journal websites are particularly vulnerable; studies have documented dead URL rates in published papers that exceed 20 percent within a decade of publication. Even when a repository remains technically operational, datasets deposited without adequate metadata — variable definitions, collection protocols, software version information — are functionally useless to outside researchers attempting reanalysis.

Federal repositories maintained by agencies such as the National Institutes of Health and the National Science Foundation offer more durable infrastructure, but they are not uniformly required across disciplines, and compliance with existing data-sharing mandates is inconsistently enforced. A 2021 review of NIH-funded studies found that a meaningful proportion of researchers who agreed to data-sharing terms in their grant applications had not fulfilled those commitments at the time of publication.

Institutional Failures in Stewardship

Universities occupy a central position in this failure. While many institutions have adopted formal research data management policies in recent years — often in response to funder mandates — the operational infrastructure to support those policies frequently lags behind the stated intent. Research data librarians, where they exist at all, are often under-resourced and overextended. Training for graduate students and early-career researchers on data management best practices remains inconsistent and, in many programs, entirely absent from the curriculum.

The problem is not exclusively one of resources. Culture matters. In many research environments, raw data is understood implicitly as the property of the principal investigator rather than as a public good generated through public funding. This framing discourages the kind of transparent, standardized archiving that reproducibility requires. When data-sharing requests arrive — from journalists, independent researchers, or regulatory bodies — the response is too often delay, incomplete disclosure, or outright refusal, sometimes years after publication.

The Reproducibility Cost

The downstream consequences for scientific reproducibility are concrete. When independent researchers cannot access the datasets underlying published findings, they cannot assess whether analytical decisions were sound, whether preprocessing steps introduced artifacts, or whether results hold under alternative modeling assumptions. The replication crisis that has received sustained attention in psychology and biomedicine is, in part, a data-access crisis — not only a methodological one. Studies that cannot be examined cannot be trusted, and science that cannot be trusted cannot serve as a reliable foundation for policy or clinical practice.

For fields such as epidemiology and clinical medicine, where findings directly inform treatment guidelines and public health recommendations, the stakes are particularly high. The inability to reanalyze trial data has, in documented cases, delayed the identification of adverse drug effects and obscured the limitations of interventions already in widespread use.

What Reform Requires

Addressing the data accessibility crisis demands action at multiple levels simultaneously. Funding agencies — NIH, NSF, the Department of Defense, and private foundations — must move from aspirational data-sharing language toward enforceable requirements with real compliance verification. Mandates that lack audit mechanisms are, in practice, optional.

Journals bear responsibility as well. Requiring data availability statements at submission is a necessary but insufficient step. Verification — confirming that linked repositories are functional and adequately documented before publication — must become standard practice. Several journals have begun piloting such checks, but adoption across the publishing landscape remains limited.

At the institutional level, universities should recognize data stewardship as a professional contribution worthy of formal evaluation in tenure and promotion processes. Embedding data management training into doctoral programs, rather than treating it as an optional supplement, would equip the next generation of researchers with habits the current generation largely was never taught.

Longer-term, the field would benefit from investment in discipline-specific repositories with guaranteed preservation commitments — infrastructure analogous to what the National Archives provides for government records. Science generates a public record. Allowing that record to quietly disappear is not a technical failure. It is an institutional choice, and it is one that can be unmade.

A Recoverable Crisis

The data accessibility problem is serious, but it is not intractable. The technical tools for durable digital preservation exist. The policy frameworks, in embryonic form, are already in place at many institutions. What has been missing is the collective will to treat data stewardship as a scientific priority rather than an administrative afterthought. Until researchers, institutions, journals, and funders align their incentives around preservation, the data graveyard will continue to grow — one retired hard drive, one expired email address, one broken link at a time.

All Articles

Related Articles

Benchmarks Don't Travel: The Reproducibility Gap Threatening Machine Learning's Scientific Credibility

Benchmarks Don't Travel: The Reproducibility Gap Threatening Machine Learning's Scientific Credibility

The Crowd and the Cosmos: How Citizen Science Is Rewriting the Rules of Research Participation

The Crowd and the Cosmos: How Citizen Science Is Rewriting the Rules of Research Participation

Transparency Without Safeguards: Examining the Unintended Costs of the Open Science Movement

Transparency Without Safeguards: Examining the Unintended Costs of the Open Science Movement