{"schemaVersion":1,"generatedAt":"2026-07-20T09:39:07.064Z","asOf":"2026-07-20","license":"Compiled aggregate counts (facts) re-served by SeqDesk under each source's terms; CC-BY for UniProt & AlphaFold. See https://seqdesk.org/data for full source list & licensing. Provided as-is, no warranty.","categories":{"sequence-archives":"Sequence archives","structures-proteins":"Structures & proteins","datasets-dois":"Datasets, DOIs & repositories","omics-specialized":"Omics & specialized archives","standards-vocab":"Standards & vocabularies","fair-literature":"FAIR adoption & literature","earth-environment":"Earth & environment","physics-materials":"Physics, space & materials","chemistry-compounds":"Chemistry & compounds","biodiversity":"Biodiversity & specimens","clinical-biomed":"Clinical & biomedical","open-data":"Open data & repositories","metadata-completeness":"Metadata & completeness","sequencing-technology":"Sequencing technology"},"licensing":{"summary":"Figures here are compiled aggregate counts — single totals updated weekly from public archives and cached by SeqDesk. They are facts, not reproductions of the underlying records. Each carries its source and retrieval date and is provided “as is” with no warranty. SeqDesk is independent and not endorsed by any listed organization.","providers":[{"name":"NCBI / U.S. National Library of Medicine","scope":"SRA, GenBank, RefSeq, ClinVar, dbSNP, Taxonomy, NCBI Datasets, NCBI Virus, GEO","license":"US-gov public domain · no use/distribution restrictions","url":"https://www.ncbi.nlm.nih.gov/home/about/policies/"},{"name":"EMBL-EBI","scope":"ENA, BioSamples, BioStudies/ArrayExpress, Europe PMC, MGnify, OLS/ENVO, ENA checklists","license":"EMBL-EBI Terms of Use · CC0-aligned","url":"https://www.ebi.ac.uk/about/terms-of-use/"},{"name":"UniProt Consortium","scope":"UniProtKB, Swiss-Prot, TrEMBL, InterPro, Pfam","license":"CC BY 4.0 (attribution required)","url":"https://creativecommons.org/licenses/by/4.0/"},{"name":"RCSB PDB / wwPDB","scope":"released structures","license":"CC0 1.0","url":"https://creativecommons.org/publicdomain/zero/1.0/"},{"name":"AlphaFold DB (Google DeepMind / EMBL-EBI)","scope":"predicted structures","license":"CC BY 4.0 (attribution required)","url":"https://creativecommons.org/licenses/by/4.0/"},{"name":"OpenAlex (OurResearch)","scope":"works, dataset works, FAIR-paper citations","license":"CC0","url":"https://creativecommons.org/publicdomain/zero/1.0/"},{"name":"DataCite & Crossref","scope":"dataset DOIs, total DOIs","license":"CC0 (metadata)","url":"https://datacite.org/"},{"name":"Zenodo · OSF · Dryad · re3data","scope":"records, projects, datasets, repositories","license":"open terms · counts are facts","url":"https://zenodo.org/"},{"name":"bioRxiv/medRxiv · OBO Foundry · GSC MIxS","scope":"preprints, ontologies, MIxS terms","license":"open / CC","url":"https://www.biorxiv.org/"},{"name":"GBIF · OBIS · iNaturalist","scope":"biodiversity occurrences, datasets, observations","license":"CC0 / CC BY per record · counts are facts","url":"https://www.gbif.org/terms"},{"name":"NCBI PubChem · ClinicalTrials.gov (NLM)","scope":"compounds, substances, bioassays, registered trials & results","license":"US-gov public domain","url":"https://www.ncbi.nlm.nih.gov/home/about/policies/"},{"name":"CDS Strasbourg · ESA Gaia · NASA/IPAC","scope":"SIMBAD, VizieR, Gaia DR3, Exoplanet Archive, EOSDIS CMR","license":"CC BY 4.0 / Gaia licence / US-gov open","url":"https://cds.unistra.fr/"},{"name":"CERN (Open Data · INSPIRE-HEP)","scope":"physics records & open datasets","license":"CC0 / open","url":"https://opendata.cern.ch/"},{"name":"Materials Project · OQMD · NOMAD","scope":"computational materials (OPTIMADE)","license":"CC BY 4.0","url":"https://materialsproject.org/about/terms"},{"name":"PANGAEA · ESGF (WCRP CMIP6)","scope":"Earth & climate datasets","license":"CC BY (per dataset)","url":"https://www.pangaea.de/"},{"name":"EMBL-EBI ChEMBL · ChEBI · GWAS Catalog · ENCODE · NeuroMorpho.Org","scope":"chemistry, ontologies, associations, experiments, neuron morphologies","license":"CC BY / open","url":"https://www.ebi.ac.uk/about/terms-of-use/"},{"name":"Harvard Dataverse · figshare · data.europa.eu · World Bank","scope":"cross-domain datasets & development indicators","license":"open terms · counts are facts","url":"https://dataverse.harvard.edu/"},{"name":"NHGRI (National Human Genome Research Institute)","scope":"DNA sequencing cost data (cost per genome, cost per Mb)","license":"U.S. Government work / public domain (cite NHGRI)","url":"https://www.genome.gov/about-genomics/fact-sheets/DNA-Sequencing-Costs-Data"}]},"count":1,"metrics":[{"id":"genbank-wgs-bases","label":"GenBank WGS sequence","category":"sequence-archives","source":"NCBI / GenBank release notes","unit":"bases","tier":"flagship","flagship":true,"cadence":"release","scale":"log","fetch":{"url":"","method":"GET","parse":{"type":"manual"},"auto":false,"appendPolicy":"manual"},"release":null,"series":[["2002-01-01",6702372564],["2004-01-01",35009256228],["2006-01-01",81611376856],["2008-01-01",141374971004],["2010-01-01",177385297156],["2012-01-01",356002922838],["2014-01-01",848977922022],["2016-01-01",1817189565845],["2018-01-01",3656719423096],["2020-01-01",11830842428018],["2022-01-01",19086596616569],["2024-01-01",32983029087303]],"note":"Whole-genome sequence in bases — the GenBank WGS division only (not the larger set-based WGS/TSA/TLS header total in the release notes). Each YYYY-01-01 point is that calendar year's December release, not literally Jan 1: e.g. 2024-01-01 = 32,983,029,087,303 is GenBank Release 264.0 (19 Dec 2024) and 2022-01-01 = 19,086,596,616,569 is Release 253.0 (15 Dec 2022). Spacing is a uniform 2 years, so the trend is intact; only the labels are year-end. Not exposed as a live API — update manually from https://www.ncbi.nlm.nih.gov/genbank/statistics/.","events":[{"date":"1982-01-01","label":"GenBank founded","detail":"The Los Alamos sequence library won a five-year NIGMS grant in 1982 and was christened GenBank, establishing the public nucleotide database whose WGS holdings this curve tracks."},{"date":"2005-09-01","label":"454 NGS arrives","detail":"Margulies et al. published the picolitre-reactor 454 pyrosequencing system in Nature (437:376-380), launching next-generation sequencing with ~100x the throughput of capillary instruments."},{"date":"2006-01-01","label":"Illumina/Solexa GA","detail":"Solexa launched the Genome Analyzer in 2006, bringing sequencing-by-synthesis short reads (1 Gb per run) that would soon dominate sequence submissions."},{"date":"2008-01-22","label":"1000 Genomes Proj.","detail":"An international consortium announced the 1000 Genomes Project on 22 Jan 2008, a population-scale effort that poured large volumes of human sequence into public databases."},{"date":"2014-01-14","label":"$1000 genome","detail":"Illumina introduced the HiSeq X Ten on 14 Jan 2014, the first platform to break the $1,000-per-genome barrier and slash the cost of large-scale sequencing."},{"date":"2020-01-11","label":"SARS-CoV-2 genome","detail":"The first SARS-CoV-2 genome (Wuhan-Hu-1) was made public via virological.org on 11 Jan 2020, igniting an unprecedented global viral-genome sequencing surge."}]}]}