SeqDesk · original analysis
Two whole-index desk-checks of open-data practice, re-counted weekly: the share of open-access papers that link at least one dataset (Europe PMC), and how much of DataCite's 130-million-DOI corpus carries a machine-readable licence. Aggregate counts only.
Each point is the share of that year’s open-access papers that Europe PMC can link to at least one dataset — a text-mined accession or a data-availability statement pointing at a repository. The share has climbed from 35% in 2010 to 51% in 2025: data a detectable data link is now present for the majority of open papers, though still far from universal. The denominator is deliberately the open-access corpus — the papers whose full text can actually be inspected for data links; counting paywalled abstracts as “no data” would conflate access with a measurement limit. The most recent year is still in progress (marked * in the table below).
A licence is only half the story. Across DataCite’s 130M DOIs, 70.6% carry some rights statement — but only 30.1% attach a standard SPDX identifier a machine can resolve and act on. The 41-point gap is free-text rights that a human can read but a reuse pipeline cannot, the open-data analogue of the “present but not usable” metadata gap.
Two licences dominate the machine-readable set: CC0 (an explicit public-domain dedication) and CC BY. The long tail of NonCommercial and NoDerivatives variants adds reuse conditions that restrict large-scale text-mining and redistribution. Shares are of the 39M DOIs that carry an SPDX identifier.
HAS_DATA:y), per publication yearscripts/check-open-data.mjs| Year | OA papers | With a data link | Share |
|---|---|---|---|
| 2026 * | 400K | 207K | 51.8% |
| 2025 | 914K | 466K | 51.0% |
| 2024 | 761K | 373K | 49.0% |
| 2023 | 767K | 357K | 46.5% |
| 2022 | 852K | 390K | 45.8% |
| 2021 | 755K | 344K | 45.6% |
| 2020 | 601K | 270K | 44.8% |
| 2019 | 415K | 192K | 46.4% |
| 2018 | 356K | 159K | 44.7% |
| 2017 | 315K | 134K | 42.6% |
| 2016 | 276K | 111K | 40.3% |
| 2015 | 252K | 95K | 37.6% |
| 2014 | 224K | 81K | 36.3% |
| 2013 | 187K | 66K | 35.3% |
| 2012 | 151K | 52K | 34.2% |
| 2011 | 111K | 38K | 34.6% |
| 2010 | 82K | 29K | 35.4% |
The open-science counterpart to the archive counts: not how much data exists, but whether published work actually shares it and licenses it for reuse — re-counted weekly straight from the EMBL-EBI / Europe PMC and DataCite REST API. Only aggregate percentages and counts are published. · back to all research data
Use & cite this data
These are SeqDesk’s own aggregate figures, refreshed weekly. Download them, drop the live chart into a page, or pull the latest numbers from a small public API — then cite the snapshot you used.
date,value table — opens straight in Excel, R, or pandas.Download CSVAggregate facts compiled by SeqDesk from public archives (EMBL-EBI / Europe PMC; DataCite) and re-served under each source's terms. Only whole-archive percentages and counts are published — no record-level identifiers. Provided as-is, no warranty. Full method & sources: https://seqdesk.org/data/open-data