← All research data

SeqDesk · original analysis

Do published papers actually share their data?

Two whole-index desk-checks of open-data practice, re-counted weekly: the share of open-access papers that link at least one dataset (Europe PMC), and how much of DataCite's 130-million-DOI corpus carries a machine-readable licence. Aggregate counts only.

79%
of open-access papers link data · 2025
134M
DataCite DOIs checked
71%
state a licence at all
31%
carry a machine-readable licence
0%25%50%75%100%201020122014201620182020202220242025publication year (open-access papers)

Each point is the share of that year’s open-access papers that Europe PMC can link to at least one dataset — a text-mined accession or a data-availability statement pointing at a repository. The share has climbed from 40% in 2010 to 79% in 2025: data a detectable data link is now present for the majority of open papers, though still far from universal. The denominator is deliberately the open-access corpus — the papers whose full text can actually be inspected for data links; counting paywalled abstracts as “no data” would conflate access with a measurement limit. The most recent year is still in progress (marked * in the table below).

States a licence any rights text — human-readable
71.0%
Machine-readable licence a resolvable SPDX identifier
30.5%

A licence is only half the story. Across DataCite’s 134M DOIs, 71.0% carry some rights statement — but only 30.5% attach a standard SPDX identifier a machine can resolve and act on. The 41-point gap is free-text rights that a human can read but a reuse pipeline cannot, the open-data analogue of the “present but not usable” metadata gap.

CC BY 4.0 16M
38.7%
CC0 1.0 (public domain) 16M
38.6%
CC BY-NC 4.0 4.7M
11.6%
CC BY-SA 4.0 1.2M
2.8%
CC BY-NC-SA 4.0 798K
1.9%
CC BY-NC-ND 4.0 635K
1.6%
CC BY 3.0 343K
0.8%
MIT 252K
0.6%
Apache 2.0 65K
0.2%

Two licences dominate the machine-readable set: CC0 (an explicit public-domain dedication) and CC BY. The long tail of NonCommercial and NoDerivatives variants adds reuse conditions that restrict large-scale text-mining and redistribution. Shares are of the 41M DOIs that carry an SPDX identifier.

Papers metricShare of open-access papers carrying a Europe PMC data link (HAS_DATA:y), per publication year
Licence metricShare of all DataCite DOIs with any rights statement vs. a resolvable SPDX licence identifier
Update cadenceWeekly automated re-count · latest snapshot 2026-09-07
MeasurementWhole-index hit counts from the EMBL-EBI / Europe PMC and DataCite REST API — exact, not a sample
Denominatoropen-access papers, by publication year — the corpus whose full text can be mined for data links
Methodscripts/check-open-data.mjs
YearOA papersWith a data linkShare
2026 *539K402K74.6%
2025918K726K79.0%
2024770K596K77.4%
2023770K573K74.4%
2022850K622K73.1%
2021751K534K71.1%
2020598K405K67.6%
2019415K275K66.2%
2018356K224K63.0%
2017315K189K59.9%
2016276K148K53.5%
2015252K120K47.6%
2014224K98K44.0%
2013187K77K41.5%
2012151K59K39.4%
2011111K44K39.3%
201082K32K39.6%

The open-science counterpart to the archive counts: not how much data exists, but whether published work actually shares it and licenses it for reuse — re-counted weekly straight from the EMBL-EBI / Europe PMC and DataCite REST API. Only aggregate percentages and counts are published. · back to all research data

Use & cite this data

Take this analysis into your own work

These are SeqDesk’s own aggregate figures, refreshed weekly. Download them, drop the live chart into a page, or pull the latest numbers from a small public API — then cite the snapshot you used.

Full dataset · JSONThe complete snapshot behind this page — every series and breakdown, with the source and license inline.Download JSON
Headline series · CSVThe trend shown in the chart as a tidy date,value table — opens straight in Excel, R, or pandas.Download CSV
Chart · SVGThe figure as a crisp vector — scales to any size for slides, posters, or print.
Chart · PNGA high-resolution raster of the figure for quick drops into docs and decks.
License

Aggregate facts compiled by SeqDesk from public archives (EMBL-EBI / Europe PMC; DataCite) and re-served under each source's terms. Only whole-archive percentages and counts are published — no record-level identifiers. Provided as-is, no warranty. Full method & sources: https://seqdesk.org/data/open-data