{"schemaVersion":1,"generatedAt":"2026-07-20T09:39:07.064Z","asOf":"2026-07-20","license":"Compiled aggregate counts (facts) re-served by SeqDesk under each source's terms; CC-BY for UniProt & AlphaFold. See https://seqdesk.org/data for full source list & licensing. Provided as-is, no warranty.","categories":{"sequence-archives":"Sequence archives","structures-proteins":"Structures & proteins","datasets-dois":"Datasets, DOIs & repositories","omics-specialized":"Omics & specialized archives","standards-vocab":"Standards & vocabularies","fair-literature":"FAIR adoption & literature","earth-environment":"Earth & environment","physics-materials":"Physics, space & materials","chemistry-compounds":"Chemistry & compounds","biodiversity":"Biodiversity & specimens","clinical-biomed":"Clinical & biomedical","open-data":"Open data & repositories","metadata-completeness":"Metadata & completeness","sequencing-technology":"Sequencing technology"},"licensing":{"summary":"Figures here are compiled aggregate counts — single totals updated weekly from public archives and cached by SeqDesk. They are facts, not reproductions of the underlying records. Each carries its source and retrieval date and is provided “as is” with no warranty. SeqDesk is independent and not endorsed by any listed organization.","providers":[{"name":"NCBI / U.S. National Library of Medicine","scope":"SRA, GenBank, RefSeq, ClinVar, dbSNP, Taxonomy, NCBI Datasets, NCBI Virus, GEO","license":"US-gov public domain · no use/distribution restrictions","url":"https://www.ncbi.nlm.nih.gov/home/about/policies/"},{"name":"EMBL-EBI","scope":"ENA, BioSamples, BioStudies/ArrayExpress, Europe PMC, MGnify, OLS/ENVO, ENA checklists","license":"EMBL-EBI Terms of Use · CC0-aligned","url":"https://www.ebi.ac.uk/about/terms-of-use/"},{"name":"UniProt Consortium","scope":"UniProtKB, Swiss-Prot, TrEMBL, InterPro, Pfam","license":"CC BY 4.0 (attribution required)","url":"https://creativecommons.org/licenses/by/4.0/"},{"name":"RCSB PDB / wwPDB","scope":"released structures","license":"CC0 1.0","url":"https://creativecommons.org/publicdomain/zero/1.0/"},{"name":"AlphaFold DB (Google DeepMind / EMBL-EBI)","scope":"predicted structures","license":"CC BY 4.0 (attribution required)","url":"https://creativecommons.org/licenses/by/4.0/"},{"name":"OpenAlex (OurResearch)","scope":"works, dataset works, FAIR-paper citations","license":"CC0","url":"https://creativecommons.org/publicdomain/zero/1.0/"},{"name":"DataCite & Crossref","scope":"dataset DOIs, total DOIs","license":"CC0 (metadata)","url":"https://datacite.org/"},{"name":"Zenodo · OSF · Dryad · re3data","scope":"records, projects, datasets, repositories","license":"open terms · counts are facts","url":"https://zenodo.org/"},{"name":"bioRxiv/medRxiv · OBO Foundry · GSC MIxS","scope":"preprints, ontologies, MIxS terms","license":"open / CC","url":"https://www.biorxiv.org/"},{"name":"GBIF · OBIS · iNaturalist","scope":"biodiversity occurrences, datasets, observations","license":"CC0 / CC BY per record · counts are facts","url":"https://www.gbif.org/terms"},{"name":"NCBI PubChem · ClinicalTrials.gov (NLM)","scope":"compounds, substances, bioassays, registered trials & results","license":"US-gov public domain","url":"https://www.ncbi.nlm.nih.gov/home/about/policies/"},{"name":"CDS Strasbourg · ESA Gaia · NASA/IPAC","scope":"SIMBAD, VizieR, Gaia DR3, Exoplanet Archive, EOSDIS CMR","license":"CC BY 4.0 / Gaia licence / US-gov open","url":"https://cds.unistra.fr/"},{"name":"CERN (Open Data · INSPIRE-HEP)","scope":"physics records & open datasets","license":"CC0 / open","url":"https://opendata.cern.ch/"},{"name":"Materials Project · OQMD · NOMAD","scope":"computational materials (OPTIMADE)","license":"CC BY 4.0","url":"https://materialsproject.org/about/terms"},{"name":"PANGAEA · ESGF (WCRP CMIP6)","scope":"Earth & climate datasets","license":"CC BY (per dataset)","url":"https://www.pangaea.de/"},{"name":"EMBL-EBI ChEMBL · ChEBI · GWAS Catalog · ENCODE · NeuroMorpho.Org","scope":"chemistry, ontologies, associations, experiments, neuron morphologies","license":"CC BY / open","url":"https://www.ebi.ac.uk/about/terms-of-use/"},{"name":"Harvard Dataverse · figshare · data.europa.eu · World Bank","scope":"cross-domain datasets & development indicators","license":"open terms · counts are facts","url":"https://dataverse.harvard.edu/"},{"name":"NHGRI (National Human Genome Research Institute)","scope":"DNA sequencing cost data (cost per genome, cost per Mb)","license":"U.S. Government work / public domain (cite NHGRI)","url":"https://www.genome.gov/about-genomics/fact-sheets/DNA-Sequencing-Costs-Data"}]},"count":1,"metrics":[{"id":"max-instrument-output-per-run","label":"Max sequencing output per run","category":"sequencing-technology","source":"SeqDesk — compiled from vendor spec sheets","unit":"Gb/run","tier":"secondary","flagship":false,"cadence":"curated","scale":"log","fetch":{"url":"","method":"GET","parse":{"type":"manual"},"auto":false,"appendPolicy":"manual"},"release":null,"headline":{"value":"16 Tb","unit":"per run · NovaSeq X (2023)"},"forecast":false,"series":[["2007-01-01",1],["2010-01-01",600],["2014-01-01",1800],["2017-01-01",6000],["2023-01-01",16000]],"note":"Maximum data output of the highest-throughput instrument available each year, in gigabases per run (log scale). Illumina Genome Analyzer (2007) headlined ~1 Gb/run; HiSeq 2000 reached ~600 Gb with upgraded flow cells (2010); HiSeq X ~1.8 Tb (2014, the '$1,000 genome' machine); NovaSeq 6000 ~6 Tb (2017); NovaSeq X Plus ~16 Tb (2023) — a ~16,000-fold rise in 16 years. For scale, a capillary Sanger run produced ~0.0001 Gb, so the 2007 Genome Analyzer was already ~10,000x more than Sanger. Best-available headline maxima from vendor spec sheets (upgraded-config rather than launch values in places); not forecast, since output jumps at chemistry/flow-cell launches rather than smoothly.","sources":[{"label":"Illumina sequencing platforms & specifications","url":"https://www.illumina.com/systems/sequencing-platforms.html"},{"label":"Illumina sequencing history","url":"https://www.illumina.com/science/technology/next-generation-sequencing/illumina-sequencing-history.html"}],"events":[{"date":"2005-01-01","label":"First NGS instrument","detail":"454 GS20 (2005), the first commercial next-generation sequencer, produced ~25 Mb per run — already well beyond per-run Sanger output."},{"date":"2007-01-01","label":"1 Gb/run","detail":"Illumina/Solexa Genome Analyzer headlined ~1 gigabase per run, the sequencing-by-synthesis breakthrough."},{"date":"2014-01-14","label":"HiSeq X · $1,000 genome","detail":"HiSeq X Ten (announced 14 Jan 2014) delivered ~1.8 Tb per run and the first population-scale $1,000 genome."},{"date":"2023-01-01","label":"NovaSeq X · 16 Tb","detail":"NovaSeq X Plus (shipping from early 2023) reaches ~16 Tb per dual-flow-cell run, the current Illumina output ceiling."}]}]}