← All research data

SeqDesk · revised predictions

How many human genomes would be sequenced by 2025?

In 2015, Stephens et al. projected 100 million to 2 billion human genomes by 2025. We rebuilt their figure from public archives and put the number that actually happened on it.

100M2.0B
projected by 2025 (Stephens 2015)
501000×
too high vs. ~2.0M distinct genomes (people)
16322×
too high even vs. ~6.2M genome-equivalents of data (the most generous reading)
~19.9 Pb
of human reads in ENA/SRA by 2025
110010K1.0M100M10B1.0T200520102015202020252015 · forecast made
Stephens 2015 — recorded growth (recreated)Stephens 2015 — projected 100M–2B (head-count)ours — distinct people (like-for-like)ours — genome-equivalents of data (most generous)

Stephens’ Figure 1, rebuilt on a log axis with our measured data laid over his. Following the rest of the tracker, solid = measured, dashed = forecast. The dashed vertical marks 2015, when the forecast was made; the shaded wedge is the range his three doubling rates projected to 2025. His headline was a head-count: 100M–2B human genomes (people). The solid black line is his own recorded growth (recreated from his figure); the two grey solid lines are ours, measured from the ENA/SRA archive. The darker grey line is the like-for-like — distinct people actually whole-genome-sequenced (~2.0M) — which fell 501000× short. The lighter grey line is the most generous reading: every genome-equivalent of human sequence data produced (~6.2M) — more than the number of people, because a 30× genome is ~30 genome-equivalents and much sequencing isn’t whole genomes — and even that is 16322× short.

What held upGenomics did become a genuine exabyte-scale big-data field, exactly as the paper argued — the data deluge is real.
What brokeThe headline genome count was badly over-optimistic. Population-scale sequencing was held back by cost, clinical value, consent and reimbursement — not by the sequencers — so even the slowest of their three doubling rates still lands ~50x above reality.
2015 scenarioProjected 2025vs. ~2.0M genomes (people)vs. ~6.2M data-equiv (generous)
Double every 7 months145B72,358× high23,309× high
Double every 12 months1.0B512× high165× high
Double every 18 months102M51× high16× high
Headline range (paper’s claim)100M2.0B501000× high16322× high
Realized 2025~2.0M genomes~6.2M equiv
From one genome to ~2 million — the realized timeline
  • 12001HGP draft genome — The Human Genome Project published the first working draft of the human genome — one genome, after ~13 years and ~$3 billion.
  • 220081000 Genomes Project — An international consortium launched the first population-scale human sequencing effort, the wave Stephens' curve rides through 2012.
  • 32014$1000 genome — Illumina's HiSeq X Ten broke the $1,000-per-genome barrier — the cost collapse that made Stephens' explosive projection look plausible.
  • 42018100k Genomes UK — Genomics England reached 100,000 whole genomes from NHS patients, the first national-scale clinical genomics programme.
  • 52023UK Biobank 500k — UK Biobank released whole-genome sequences for all ~500,000 participants — then the world's largest single set of human genomes.
  • 62025Realized ~2M — Cumulative human whole genomes by 2025 is on the order of 2 million — dominated by a few large biobanks, ~50-1000x below the 2015 projection.

SeqDesk · revised predictions. Stephens' Figure 1 recreated from the published log-scale figure (values approximate); the three projection lines are computed from their 2015 baseline as N(2025) = N(2015) x 2^(120 / doublingMonths). The realized count reuses the curated 'Human genomes sequenced' tracker metric (a deliberately conservative ~2M lower bound; Berkeley Genomics). The paper's headline (100M-2B human genomes by 2025) is a head-count of people, so the like-for-like is our count of distinct individuals actually whole-genome-sequenced (~2M). Our curve runs below Stephens' recorded line because he counted genomes generously — his Figure 1 sits close to total sequencing capacity, not strict distinct individuals — so the gap peaks at ~100x around 2012 and narrows to ~4x by 2015. Early reconstruction anchors: HGP draft (2001), first individual genomes (Venter 2007; Watson and Wang/YH 2008), 1000 Genomes pilot (179 individuals, 2010) and Phase 1 (1,092, 2012). The discrepancy is reported against the paper's stated 100M-2B range; treat any straight-line extrapolation on this page with the same caution. As the MOST GENEROUS possible reading we also count every genome-equivalent of human sequence data produced (total bases / one genome) from the ENA/SRA read archive — this exceeds the number of distinct people (a 30x genome is ~30 genome-equivalents, and much sequencing isn't whole genomes), yet still reaches only roughly 6 million by 2025, ~16-322x under the projection. That data-volume measure corresponds to Figure 1's other axis (sequencing capacity); it counts only data submitted to the public archive, so it too is a lower bound.

What it measuresA 10-year forecast (a head-count of human genomes) checked against reality two ways — distinct people, and a generous upper bound counting every genome-equivalent of data
The paperStephens ZD, et al. (2015) — PLoS Biology 13(7): e1002195 · forecast for 2025
Projection methodN(2025) = N(2015) × 2^(120 / doublingMonths) from their 2015 baseline of 1.0M
Realized count (people)Reuses the curated Human genomes sequenced metric (~2.0M lower bound)
Most generous readingENA Portal API (read_run, tax_eq(9606)) — mirrors NCBI SRA: Σ base_count over human runs ÷ 3.2 Gb/genome = ~6.2M genome-equivalents by 2025
CadenceFigure curated · capacity re-pulled from ENA · validated 2026-06-25
Methodcheck-genome-forecast.mjs · check-human-sequencing-capacity.mjs
  • Three quantities share one axis. The dark line is Stephens' own (generous) count of human genomes; copper is our strict count of distinct PEOPLE whole-genome-sequenced; teal is genome-equivalents of sequencing DATA (total bases / one genome). The 100M-2B headline is a head-count, so copper is the like-for-like. Teal sits above copper because a 30x genome is ~30 genome-equivalents and much human sequencing is exome, transcriptome, amplicon or single-cell rather than whole genomes — so reading data-equivalents as a count of people is a category error.
  • The capacity curve is a lower bound on data produced. It sums only human reads submitted to the open ENA/SRA read archive. A large and growing share of clinical and biobank sequencing sits in controlled-access repositories (EGA, dbGaP) or never leaves a hospital/company, so it is not counted. Stephens measured data PRODUCED; we can only measure data SUBMITTED, which is smaller. first_public also dates archive release, not when the sample was sequenced.
  • Genome-equivalents are a volume proxy, not a genome count. We take total human bases (ENA tax_eq(9606), all read types) divided by one 3.2 Gb genome. That is the right unit for Stephens' data-volume comparison, but it is not a number of sequenced individuals: redundant resequencing, deep coverage and non-WGS assays all inflate bases per person. Choosing a different divisor (e.g. a 30x genome) would shift the curve by that factor.
  • Stephens' Figure 1 is recreated, not the original data. His historical points and the three projection lines are read off the published log-scale figure (values approximate) and computed from a single 2015 baseline. His 100M-2B headline is a stated range, not a point estimate, and the paper's 'genomes' unit is itself ambiguous between sequencing capacity and a count of people.
  • The ~2M genome count is conservative and ultimately unknowable. The distinct-genome curve reuses a deliberately conservative ~2M lower bound (Berkeley Genomics). Component cohorts (UK Biobank, All of Us, gnomAD, Genomics England) must NOT be summed — severe double-counting — the true figure is likely 2-5M, and because most clinical genomes never become public it can't be counted exactly. Early-year anchors are read off milestones, so the curve is sparse before 2015.
  • How we read it. The 2015 forecast overshot either way — ~16-322x in data terms, ~50-1000x in distinct people — so the headline number was too optimistic on both measures. But the paper's core thesis held: genomics did become an exabyte-scale field. The miss was the timeline and the human-genome count, held back by cost, clinical value, consent and reimbursement rather than sequencer throughput. Treat every extrapolation on this page as a trend, not a destiny — that caution is the whole point of revisiting it.
YearCumulative bases (ENA/SRA)Genome-equivalents
2026 (partial)21.5 Pb6,712,727
202519.9 Pb6,208,471
202415.8 Pb4,926,727
202312.4 Pb3,870,250
20229.4 Pb2,939,914
20216.9 Pb2,151,534
20206.1 Pb1,914,824
20194.3 Pb1,329,606
20182.8 Pb875,831
20172.1 Pb658,938
20161.5 Pb476,887
20151.0 Pb327,977
2014680 Tb212,569
2013450 Tb140,671
2012199 Tb62,242
201192 Tb28,781
201040 Tb12,360

The first of our revised predictions: rebuild an influential, ageing forecast from public archives and check it against what actually happened. Only aggregate counts are published. · back to all research data