← All metrics

Sequence archives

GenBank WGS sequence

NCBI / GenBank release notes

33T
bases · 2024-01-01
S-curve
best-fit model
58T
projected 2026
6.7B138B2.8T58T20022006201020142018202220262345658T 2026
Logistic (S-curve) fitbest fit (AICc)N(t) = L / (1 + e^{−k(t − t₀)})
carrying capacity L 142Tsteepness k 0.374 / yrmidpoint t₀ 2027 0.9907forecast 58T by 2026

Forecast = 10% of the 22-yr record (2.2 yr ahead). Anchored on the latest value; levels off toward L ≈ 142T.

Timeline — what shaped the curve
  • 11982-01GenBank founded — The Los Alamos sequence library won a five-year NIGMS grant in 1982 and was christened GenBank, establishing the public nucleotide database whose WGS holdings this curve tracks.
  • 22005-09454 NGS arrives — Margulies et al. published the picolitre-reactor 454 pyrosequencing system in Nature (437:376-380), launching next-generation sequencing with ~100x the throughput of capillary instruments.
  • 32006-01Illumina/Solexa GA — Solexa launched the Genome Analyzer in 2006, bringing sequencing-by-synthesis short reads (1 Gb per run) that would soon dominate sequence submissions.
  • 42008-01-221000 Genomes Proj. — An international consortium announced the 1000 Genomes Project on 22 Jan 2008, a population-scale effort that poured large volumes of human sequence into public databases.
  • 52014-01-14$1000 genome — Illumina introduced the HiSeq X Ten on 14 Jan 2014, the first platform to break the $1,000-per-genome barrier and slash the cost of large-scale sequencing.
  • 62020-01-11SARS-CoV-2 genome — The first SARS-CoV-2 genome (Wuhan-Hu-1) was made public via virological.org on 11 Jan 2020, igniting an unprecedented global viral-genome sequencing surge.

Whole-genome sequence in bases — the GenBank WGS division only (not the larger set-based WGS/TSA/TLS header total in the release notes). Each YYYY-01-01 point is that calendar year's December release, not literally Jan 1: e.g. 2024-01-01 = 32,983,029,087,303 is GenBank Release 264.0 (19 Dec 2024) and 2022-01-01 = 19,086,596,616,569 is Release 253.0 (15 Dec 2022). Spacing is a uniform 2 years, so the trend is intact; only the labels are year-end. Not exposed as a live API — update manually from https://www.ncbi.nlm.nih.gov/genbank/statistics/.

SourceNCBI / GenBank release notes
Update cadencerelease · manual
Live endpoint
Parsemanual
Datebases
2024-01-0132,983,029,087,303
2022-01-0119,086,596,616,569
2020-01-0111,830,842,428,018
2018-01-013,656,719,423,096
2016-01-011,817,189,565,845
2014-01-01848,977,922,022
2012-01-01356,002,922,838
2010-01-01177,385,297,156
2008-01-01141,374,971,004
2006-01-0181,611,376,856
2004-01-0135,009,256,228
2002-01-016,702,372,564

Machine-readable: /api/research-data?metric=genbank-wgs-bases · back to all metrics

Report errorpmu15@helmholtz-hzi.de