SeqDesk · Insights
How fast the world's open research data is growing — genomes, sequencing, omics, datasets & more, re-queried from public archives, in the open.
Latest snapshot: 2026-07-20 · updated weekly · solid = measured, dashed = forecast · open any metric for the fit, confidence band, sources & downloads
Our own analyses of open research data: how much of it actually gets shared, how complete it is, and how that’s changing over time. They show us where the gaps are, and help us decide what to build next.
Every sample SeqDesk handles ends up — if it’s shared — in a public read archive, stamped with the center that submitted it. So we turn that stamp around: who submits the most, where they are, and which platforms they run. It’s a map of archive submission activity (ENA-centric, surveillance-shaped), not total capacity — read it as “who feeds the commons.” Each card opens one slice of the analysis. ENA snapshot 2026-06-29.
SeqDesk integrates sequencing tools so researchers don’t have to stitch them together themselves. Tracking how each tool’s adoption moves across the field tells us which ones matter most — and that is how we prioritize what to integrate next.
Beyond individual tools, most sequencing analysis now runs through whole pipelines. We track which workflow technologies the field is adopting — by the citation curves of their papers — how many registered workflows each carries, and, for the two engines with a GitHub-backed registry or repository population, which individual pipelines the community actually uses. That tells us which end-to-end workflows matter most, and where SeqDesk should plug in next.
A handful of influential papers put hard numbers on how fast research data would grow — and most are now several years old. We take the chance to rebuild their inputs from public archives, check each forecast against what actually happened, and revise the prediction where reality diverged.
Figures here are compiled aggregate counts — single totals updated weekly from public archives and cached by SeqDesk. They are facts, not reproductions of the underlying records. Each carries its source and retrieval date and is provided “as is” with no warranty. SeqDesk is independent and not endorsed by any listed organization.
How the numbers are read. Each is pulled directly from the host archive’s public API and refreshed weekly. Counts are cumulative archive totals; doubling times are shown only when the best-fit forecast is exponential, and each forecast extends the fit only ~10% past the data span — a trend, not a guarantee. Open any metric to switch between exponential, logistic and Gompertz fits.
One endpoint for all of it: /api/research-data serves these cached counts as JSON (filters: ?flagship=1, ?metric=<id>, ?latest=1) — so you can pull every source from one place instead of juggling a dozen APIs. The full source list lives in src/data/research-data-tracker.json.
| Provider | Covers | License |
|---|---|---|
| NCBI / U.S. National Library of Medicine | SRA, GenBank, RefSeq, ClinVar, dbSNP, Taxonomy, NCBI Datasets, NCBI Virus, GEO | US-gov public domain · no use/distribution restrictions |
| EMBL-EBI | ENA, BioSamples, BioStudies/ArrayExpress, Europe PMC, MGnify, OLS/ENVO, ENA checklists | EMBL-EBI Terms of Use · CC0-aligned |
| UniProt Consortium | UniProtKB, Swiss-Prot, TrEMBL, InterPro, Pfam | CC BY 4.0 (attribution required) |
| RCSB PDB / wwPDB | released structures | CC0 1.0 |
| AlphaFold DB (Google DeepMind / EMBL-EBI) | predicted structures | CC BY 4.0 (attribution required) |
| OpenAlex (OurResearch) | works, dataset works, FAIR-paper citations | CC0 |
| DataCite & Crossref | dataset DOIs, total DOIs | CC0 (metadata) |
| Zenodo · OSF · Dryad · re3data | records, projects, datasets, repositories | open terms · counts are facts |
| bioRxiv/medRxiv · OBO Foundry · GSC MIxS | preprints, ontologies, MIxS terms | open / CC |
| GBIF · OBIS · iNaturalist | biodiversity occurrences, datasets, observations | CC0 / CC BY per record · counts are facts |
| NCBI PubChem · ClinicalTrials.gov (NLM) | compounds, substances, bioassays, registered trials & results | US-gov public domain |
| CDS Strasbourg · ESA Gaia · NASA/IPAC | SIMBAD, VizieR, Gaia DR3, Exoplanet Archive, EOSDIS CMR | CC BY 4.0 / Gaia licence / US-gov open |
| CERN (Open Data · INSPIRE-HEP) | physics records & open datasets | CC0 / open |
| Materials Project · OQMD · NOMAD | computational materials (OPTIMADE) | CC BY 4.0 |
| PANGAEA · ESGF (WCRP CMIP6) | Earth & climate datasets | CC BY (per dataset) |
| EMBL-EBI ChEMBL · ChEBI · GWAS Catalog · ENCODE · NeuroMorpho.Org | chemistry, ontologies, associations, experiments, neuron morphologies | CC BY / open |
| Harvard Dataverse · figshare · data.europa.eu · World Bank | cross-domain datasets & development indicators | open terms · counts are facts |
| NHGRI (National Human Genome Research Institute) | DNA sequencing cost data (cost per genome, cost per Mb) | U.S. Government work / public domain (cite NHGRI) |