← All research data

SeqDesk · original analysis

Do published papers actually share their data?

Two whole-index desk-checks of open-data practice, re-counted weekly: the share of open-access papers that link at least one dataset (Europe PMC), and how much of DataCite's 130-million-DOI corpus carries a machine-readable licence. Aggregate counts only.

51%
of open-access papers link data · 2025
130M
DataCite DOIs checked
71%
state a licence at all
30%
carry a machine-readable licence
0%25%50%75%100%201020122014201620182020202220242025publication year (open-access papers)

Each point is the share of that year’s open-access papers that Europe PMC can link to at least one dataset — a text-mined accession or a data-availability statement pointing at a repository. The share has climbed from 35% in 2010 to 51% in 2025: data a detectable data link is now present for the majority of open papers, though still far from universal. The denominator is deliberately the open-access corpus — the papers whose full text can actually be inspected for data links; counting paywalled abstracts as “no data” would conflate access with a measurement limit. The most recent year is still in progress (marked * in the table below).

States a licence any rights text — human-readable
70.6%
Machine-readable licence a resolvable SPDX identifier
30.1%

A licence is only half the story. Across DataCite’s 130M DOIs, 70.6% carry some rights statement — but only 30.1% attach a standard SPDX identifier a machine can resolve and act on. The 41-point gap is free-text rights that a human can read but a reuse pipeline cannot, the open-data analogue of the “present but not usable” metadata gap.

CC0 1.0 (public domain) 16M
39.9%
CC BY 4.0 15M
37.4%
CC BY-NC 4.0 4.6M
11.7%
CC BY-SA 4.0 1.1M
2.9%
CC BY-NC-SA 4.0 764K
2.0%
CC BY-NC-ND 4.0 596K
1.5%
CC BY 3.0 343K
0.9%
MIT 196K
0.5%
Apache 2.0 53K
0.1%

Two licences dominate the machine-readable set: CC0 (an explicit public-domain dedication) and CC BY. The long tail of NonCommercial and NoDerivatives variants adds reuse conditions that restrict large-scale text-mining and redistribution. Shares are of the 39M DOIs that carry an SPDX identifier.

Papers metricShare of open-access papers carrying a Europe PMC data link (HAS_DATA:y), per publication year
Licence metricShare of all DataCite DOIs with any rights statement vs. a resolvable SPDX licence identifier
Update cadenceWeekly automated re-count · latest snapshot 2026-06-25
MeasurementWhole-index hit counts from the EMBL-EBI / Europe PMC and DataCite REST API — exact, not a sample
Denominatoropen-access papers, by publication year — the corpus whose full text can be mined for data links
Methodscripts/check-open-data.mjs
YearOA papersWith a data linkShare
2026 *400K207K51.8%
2025914K466K51.0%
2024761K373K49.0%
2023767K357K46.5%
2022852K390K45.8%
2021755K344K45.6%
2020601K270K44.8%
2019415K192K46.4%
2018356K159K44.7%
2017315K134K42.6%
2016276K111K40.3%
2015252K95K37.6%
2014224K81K36.3%
2013187K66K35.3%
2012151K52K34.2%
2011111K38K34.6%
201082K29K35.4%

The open-science counterpart to the archive counts: not how much data exists, but whether published work actually shares it and licenses it for reuse — re-counted weekly straight from the EMBL-EBI / Europe PMC and DataCite REST API. Only aggregate percentages and counts are published. · back to all research data

Use & cite this data

Take this analysis into your own work

These are SeqDesk’s own aggregate figures, refreshed weekly. Download them, drop the live chart into a page, or pull the latest numbers from a small public API — then cite the snapshot you used.

Full dataset · JSONThe complete snapshot behind this page — every series and breakdown, with the source and license inline.Download JSON
Headline series · CSVThe trend shown in the chart as a tidy date,value table — opens straight in Excel, R, or pandas.Download CSV
Chart · SVGThe figure as a crisp vector — scales to any size for slides, posters, or print.
Chart · PNGA high-resolution raster of the figure for quick drops into docs and decks.
License

Aggregate facts compiled by SeqDesk from public archives (EMBL-EBI / Europe PMC; DataCite) and re-served under each source's terms. Only whole-archive percentages and counts are published — no record-level identifiers. Provided as-is, no warranty. Full method & sources: https://seqdesk.org/data/open-data