SeqDesk · original analysis
Preprint volume has grown roughly twenty-fold in a decade, yet preprints link a dataset far less often than open-access MEDLINE-indexed journal papers. Whole-index counts from Europe PMC, re-counted weekly. The preprint data-link share is shown for transparency but is sensitive to annotation lag — read the gap and the volume, not single-year wobbles.
Europe PMC indexed 7.6K preprints in 2016 and 187K in 2025 — a roughly 25-fold rise that reshaped how biology is disseminated. Preprints are now a major channel for first reporting of results; the question is whether the data underneath travels with them.
In every year, a preprint is far less likely than the published version of record to carry a resolvable data link — about 82% for open-access journal papers versus 6% for preprints in 2025, a roughly 13-fold gap. Two honest caveats about the preprint line: the 2020 bump is COVID-era data-rich preprints, and the apparent recent decline is largely an artefact — Europe PMC’s text-mining lags on very recent preprints, and preprint servers have broadened well beyond data-rich life sciences, diluting the denominator. The durable finding is the size of the gap, not the shape of the preprint curve.
SRC:PPR) per publication yearHAS_DATA:y vs. open-access published papers with HAS_DATA:y, same yearscripts/check-preprints.mjs| Year | Preprints | Link data | Published do |
|---|---|---|---|
| 2026 * | 150K | 4.1% | 75.5% |
| 2025 | 187K | 6.4% | 81.6% |
| 2024 | 182K | 6.5% | 79.7% |
| 2023 | 176K | 10.4% | 76.7% |
| 2022 | 148K | 12.4% | 75.7% |
| 2021 | 154K | 12.6% | 73.5% |
| 2020 | 128K | 15.6% | 69.9% |
| 2019 | 49K | 1.7% | 68.2% |
| 2018 | 31K | 1.1% | 64.1% |
| 2017 | 16K | 1.3% | 60.9% |
| 2016 | 7.6K | 1.4% | 54.3% |
The dissemination counterpart to the open-data card: preprints now carry a huge share of new results, but the data underneath them is linked far less often than for the version of record — re-counted weekly straight from the EMBL-EBI / Europe PMC. Only aggregate percentages and counts are published. · back to all research data
Use & cite this data
These are SeqDesk’s own aggregate figures, refreshed weekly. Download them, drop the live chart into a page, or pull the latest numbers from a small public API — then cite the snapshot you used.
date,value table — opens straight in Excel, R, or pandas.Download CSVAggregate facts compiled by SeqDesk from public archives (EMBL-EBI / Europe PMC) and re-served under each source's terms. Only whole-archive percentages and counts are published — no record-level identifiers. Provided as-is, no warranty. Full method & sources: https://seqdesk.org/data/preprints