SeqDesk · Insights

Research data, tracked

How fast the world's open research data is growing — genomes, sequencing, omics, datasets & more, re-queried from public archives, in the open.

Latest snapshot: 2026-07-20 · updated weekly · solid = measured, dashed = forecast · open any metric for the fit, confidence band, sources & downloads

Our own analyses of open research data: how much of it actually gets shared, how complete it is, and how that’s changing over time. They show us where the gaps are, and help us decide what to build next.

Volume vs context metadata
SeqDesk · original analysis
weekly
12% carry a MIxS environment
5.0M samples · 2025while volume rose 866× since 2008
Bases scale faster than standardized environment metadata · open for the full chart
Sequencing metadata completeness
SeqDesk · original analysis
weekly
38% of MIxS context filled
5.0M samples · 2025coordinates only 26%
How complete is public sequencing metadata? · open for the full analysis
Reporting-standard adoption
SeqDesk · original analysis
weekly
66% on a real MIxS package
5.0M samples · 202533% pick “Generic”
Are samples held to a real reporting standard? · open for the full analysis
Facilities, by metadata completeness
SeqDesk · original analysis
weekly
83% best — New York Genome Center
16 top centers scoredwhole-archive baseline 80%
Ranked by how well they describe data, not volume · open the league table
Reads vs assemblies
SeqDesk · original analysis
weekly
12 assemblies per 100 read sets
42.7M reads · 5.1M assembliescenters range 0–153.7 per 100
The mix of raw reads and assembled genomes each center deposits · open the analysis
Cross-domain data completeness
SeqDesk · original analysis
weekly
96% of GBIF records georeferenced
3.9B records · 6 signals1950 cohort: 70%
How complete is the metadata in major open research archives? · open for the full analysis
Research data availability
SeqDesk · original analysis
weekly
79% still online
148 databases trackeddown from 100% since 2009
How long does research data stay online? · open for the full analysis
Open-data practice
SeqDesk · original analysis
weekly
51% of OA papers link data
up from 35% in 2010only 30% of DataCite DOIs carry a machine-readable licence
Do papers share data - and is it openly licensed? · open for the full analysis
Open-access share
SeqDesk · original analysis
weekly
52% of papers are open access
up from 9% in 2010874K free papers in 2025
How much of the literature is free to read? · open for the full analysis
Preprints and data sharing
SeqDesk · original analysis
weekly
lower data-link rate than journals
186K preprints/yr · ~25× since 20166% link data vs 53% for journals
Preprints grew sharply - did data-link rates keep pace? · open for the full analysis
FAIR scorecard
SeqDesk · original analysis
interactive
F·A·I·R score your facility
interactive toolglobal ENA baseline 80%
Drag the sliders to see F·A·I·R move against the archive · open the scorecard
Facility profile
SeqDesk · original analysis
interactive
16 institutes to explore
metadata · genomes · platformspick any one center
Choose an institute and see its full FAIR profile · open the explorer

Every sample SeqDesk handles ends up — if it’s shared — in a public read archive, stamped with the center that submitted it. So we turn that stamp around: who submits the most, where they are, and which platforms they run. It’s a map of archive submission activity (ENA-centric, surveillance-shaped), not total capacity — read it as “who feeds the commons.” Each card opens one slice of the analysis. ENA snapshot 2026-06-29.

Most active centers
SeqDesk · original analysis
weekly
3.8M runs · Wellcome Sanger Institute
141 centers ≥100 runs61 profiled
Wellcome Sanger Institute
3.8M
COVID-19 Genomics UK Consortium (COG-UK)
591K
Statens Serum Institut
527K
University of California San Diego
468K
Who submits most to the public read archive? · open
Unique submitters to ENA
SeqDesk · original analysis
weekly
1,411 distinct submitter names
141 with ≥100 runs · 61 profileddistinct center_name strings — alias-split & ENA-scoped, not merged facilities
Ever submitted
1,411
≥100 runs
141
Profiled in full
61
How many distinct centers have ever submitted? · open
Where they are
SeqDesk · original analysis
weekly
5.0M runs · United Kingdom
20 countries17 centers lead
United Kingdom
5.0M
United States
1.2M
Denmark
612K
Switzerland
327K
Which countries feed the archive? · open
Platform adoption
SeqDesk · original analysis
weekly
2.5% Oxford Nanopore · 2025
2010–2025Illumina 91% of tracked
How did named platform shares change over time? · open
Platform mix
SeqDesk · original analysis
weekly
90% Illumina
43M runsOxford Nanopore 2.5%
Illumina 90%Oxford Nanopore 2.5%PacBio 2.2%BGI / MGI 1.5%Ion Torrent 1.3%Other 2.2%
Which platforms dominate the archive? · open
Who runs them
SeqDesk · original analysis
weekly
41% universities
4.9% commercialResearch institutes & public lead runs
Academic 25Research institutes & public 33Commercial 3
Academic, public, or commercial? · open
Data brokers, excluded
SeqDesk · original analysis
weekly
17M submissions · EMBL-EBI
14 names excludedarchive aggregators, not sequencers
EMBL-EBI
17M
European Bioinformatics Institute
287K
EBI
220K
Why archive aggregators are excluded · open

A handful of influential papers put hard numbers on how fast research data would grow — and most are now several years old. We take the chance to rebuild their inputs from public archives, check each forecast against what actually happened, and revise the prediction where reality diverged.

ENA raw read datasets
EMBL-EBI / European Nucleotide Archive
43M runs
Gompertz fit46M by 2027
ENA assembled sequence records
EMBL-EBI / ENA
313M records
Gompertz fit330M by 2028
NCBI SRA records
NCBI / Sequence Read Archive
46M records
doubles every 4.9 yr51M by 2027
GenBank nucleotide records
NCBI / GenBank (nuccore)
735M records
Gompertz fit862M by 2029
NCBI genome assemblies
NCBI Datasets
4.3M assemblies
S-curve fit5.5M by 2028
GenBank WGS sequence
NCBI / GenBank release notes
release
33T bases
S-curve fit58T by 2026
Human genomes sequenced
SeqDesk estimate · Berkeley Genomics, UK Biobank, gnomAD, Stephens 2015
2.0M genomes (WGS)
DataCite dataset DOIs
DataCite
72M DOIs
doubles every 1.6 yr132M by 2028
DataCite total DOIs
DataCite
132M DOIs
doubles every 2.0 yr218M by 2028
Zenodo records
CERN / Zenodo
7.0M records
doubles every 2.7 yr9.3M by 2027
OSF public projects
COS / OSF
627K projects
Gompertz fit746K by 2027
EBI BioSamples total
EMBL-EBI / BioSamples
54M samples
Gompertz fit66M by 2027
NCBI GEO samples (GSM)
NCBI / GEO
8.6M samples
Gompertz fit11M by 2028
BioStudies total studies
EMBL-EBI / BioStudies
3.4M datasets
doubles every 3.0 yr4.0M by 2027
ClinVar records
NCBI / ClinVar
4.5M records
Gompertz fit6.0M by 2027
SARS-CoV-2 sequences
NCBI Virus
9.2M records
Gompertz fit9.3M by 2027
FAIR Principles paper citations
OpenAlex (Wilkinson 2016)
18K citations
S-curve fit18K by 2027
Total scholarly works (OpenAlex)
OpenAlex
321M works
S-curve fit340M by 2027
Open-access articles (Europe PMC)
EMBL-EBI / Europe PMC
8.0M records
Gompertz fit9.9M by 2028
Data sources, licensing & using this as an API

Figures here are compiled aggregate counts — single totals updated weekly from public archives and cached by SeqDesk. They are facts, not reproductions of the underlying records. Each carries its source and retrieval date and is provided “as is” with no warranty. SeqDesk is independent and not endorsed by any listed organization.

How the numbers are read. Each is pulled directly from the host archive’s public API and refreshed weekly. Counts are cumulative archive totals; doubling times are shown only when the best-fit forecast is exponential, and each forecast extends the fit only ~10% past the data span — a trend, not a guarantee. Open any metric to switch between exponential, logistic and Gompertz fits.

One endpoint for all of it: /api/research-data serves these cached counts as JSON (filters: ?flagship=1, ?metric=<id>, ?latest=1) — so you can pull every source from one place instead of juggling a dozen APIs. The full source list lives in src/data/research-data-tracker.json.

ProviderCoversLicense
NCBI / U.S. National Library of MedicineSRA, GenBank, RefSeq, ClinVar, dbSNP, Taxonomy, NCBI Datasets, NCBI Virus, GEOUS-gov public domain · no use/distribution restrictions
EMBL-EBIENA, BioSamples, BioStudies/ArrayExpress, Europe PMC, MGnify, OLS/ENVO, ENA checklistsEMBL-EBI Terms of Use · CC0-aligned
UniProt ConsortiumUniProtKB, Swiss-Prot, TrEMBL, InterPro, PfamCC BY 4.0 (attribution required)
RCSB PDB / wwPDBreleased structuresCC0 1.0
AlphaFold DB (Google DeepMind / EMBL-EBI)predicted structuresCC BY 4.0 (attribution required)
OpenAlex (OurResearch)works, dataset works, FAIR-paper citationsCC0
DataCite & Crossrefdataset DOIs, total DOIsCC0 (metadata)
Zenodo · OSF · Dryad · re3datarecords, projects, datasets, repositoriesopen terms · counts are facts
bioRxiv/medRxiv · OBO Foundry · GSC MIxSpreprints, ontologies, MIxS termsopen / CC
GBIF · OBIS · iNaturalistbiodiversity occurrences, datasets, observationsCC0 / CC BY per record · counts are facts
NCBI PubChem · ClinicalTrials.gov (NLM)compounds, substances, bioassays, registered trials & resultsUS-gov public domain
CDS Strasbourg · ESA Gaia · NASA/IPACSIMBAD, VizieR, Gaia DR3, Exoplanet Archive, EOSDIS CMRCC BY 4.0 / Gaia licence / US-gov open
CERN (Open Data · INSPIRE-HEP)physics records & open datasetsCC0 / open
Materials Project · OQMD · NOMADcomputational materials (OPTIMADE)CC BY 4.0
PANGAEA · ESGF (WCRP CMIP6)Earth & climate datasetsCC BY (per dataset)
EMBL-EBI ChEMBL · ChEBI · GWAS Catalog · ENCODE · NeuroMorpho.Orgchemistry, ontologies, associations, experiments, neuron morphologiesCC BY / open
Harvard Dataverse · figshare · data.europa.eu · World Bankcross-domain datasets & development indicatorsopen terms · counts are facts
NHGRI (National Human Genome Research Institute)DNA sequencing cost data (cost per genome, cost per Mb)U.S. Government work / public domain (cite NHGRI)