How many human genomes would be sequenced by 2025?
In 2015, Stephens et al. projected 100 million to 2 billion human genomes by 2025. We rebuilt their figure from public archives and put the number that actually happened on it.
100M–2.0B
projected by 2025 (Stephens 2015)
50–1000×
too high vs. ~2.0M distinct genomes (people)
16–322×
too high even vs. ~6.2M genome-equivalents of data (the most generous reading)
~19.9 Pb
of human reads in ENA/SRA by 2025
Stephens 2015 — recorded growth (recreated)Stephens 2015 — projected 100M–2B (head-count)ours — distinct people (like-for-like)ours — genome-equivalents of data (most generous)
Stephens’ Figure 1, rebuilt on a log axis with our measured data laid over his. Following the rest of the tracker, solid = measured, dashed = forecast. The dashed vertical marks 2015, when the forecast was made; the shaded wedge is the range his three doubling rates projected to 2025. His headline was a head-count: 100M–2B human genomes (people). The solid black line is his own recorded growth (recreated from his figure); the two grey solid lines are ours, measured from the ENA/SRA archive. The darker grey line is the like-for-like — distinct people actually whole-genome-sequenced (~2.0M) — which fell 50–1000× short. The lighter grey line is the most generous reading: every genome-equivalent of human sequence data produced (~6.2M) — more than the number of people, because a 30× genome is ~30 genome-equivalents and much sequencing isn’t whole genomes — and even that is 16–322× short.
Right about the growth, off on the head-count
What held upGenomics did become a genuine exabyte-scale big-data field, exactly as the paper argued — the data deluge is real.
What brokeThe headline genome count was badly over-optimistic. Population-scale sequencing was held back by cost, clinical value, consent and reimbursement — not by the sequencers — so even the slowest of their three doubling rates still lands ~50x above reality.
Each 2015 scenario vs. what happened
2015 scenario
Projected 2025
vs. ~2.0M genomes (people)
vs. ~6.2M data-equiv (generous)
Double every 7 months
145B
72,358× high
23,309× high
Double every 12 months
1.0B
512× high
165× high
Double every 18 months
102M
51× high
16× high
Headline range (paper’s claim)
100M–2.0B
50–1000× high
16–322× high
Realized 2025
—
~2.0M genomes
~6.2M equiv
From one genome to ~2 million — the realized timeline
12001HGP draft genome — The Human Genome Project published the first working draft of the human genome — one genome, after ~13 years and ~$3 billion.
220081000 Genomes Project — An international consortium launched the first population-scale human sequencing effort, the wave Stephens' curve rides through 2012.
32014$1000 genome — Illumina's HiSeq X Ten broke the $1,000-per-genome barrier — the cost collapse that made Stephens' explosive projection look plausible.
42018100k Genomes UK — Genomics England reached 100,000 whole genomes from NHS patients, the first national-scale clinical genomics programme.
52023UK Biobank 500k — UK Biobank released whole-genome sequences for all ~500,000 participants — then the world's largest single set of human genomes.
62025Realized ~2M — Cumulative human whole genomes by 2025 is on the order of 2 million — dominated by a few large biobanks, ~50-1000x below the 2015 projection.
SeqDesk · revised predictions. Stephens' Figure 1 recreated from the published log-scale figure (values approximate); the three projection lines are computed from their 2015 baseline as N(2025) = N(2015) x 2^(120 / doublingMonths). The realized count reuses the curated 'Human genomes sequenced' tracker metric (a deliberately conservative ~2M lower bound; Berkeley Genomics). The paper's headline (100M-2B human genomes by 2025) is a head-count of people, so the like-for-like is our count of distinct individuals actually whole-genome-sequenced (~2M). Our curve runs below Stephens' recorded line because he counted genomes generously — his Figure 1 sits close to total sequencing capacity, not strict distinct individuals — so the gap peaks at ~100x around 2012 and narrows to ~4x by 2015. Early reconstruction anchors: HGP draft (2001), first individual genomes (Venter 2007; Watson and Wang/YH 2008), 1000 Genomes pilot (179 individuals, 2010) and Phase 1 (1,092, 2012). The discrepancy is reported against the paper's stated 100M-2B range; treat any straight-line extrapolation on this page with the same caution. As the MOST GENEROUS possible reading we also count every genome-equivalent of human sequence data produced (total bases / one genome) from the ENA/SRA read archive — this exceeds the number of distinct people (a 30x genome is ~30 genome-equivalents, and much sequencing isn't whole genomes), yet still reaches only roughly 6 million by 2025, ~16-322x under the projection. That data-volume measure corresponds to Figure 1's other axis (sequencing capacity); it counts only data submitted to the public archive, so it too is a lower bound.
How this is tracked
What it measuresA 10-year forecast (a head-count of human genomes) checked against reality two ways — distinct people, and a generous upper bound counting every genome-equivalent of data
The paperStephens ZD, et al. (2015) — PLoS Biology 13(7): e1002195 · forecast for 2025
Projection methodN(2025) = N(2015) × 2^(120 / doublingMonths) from their 2015 baseline of 1.0M
Realized count (people)Reuses the curated Human genomes sequenced metric (~2.0M lower bound)
Most generous readingENA Portal API (read_run, tax_eq(9606)) — mirrors NCBI SRA: Σ base_count over human runs ÷ 3.2 Gb/genome = ~6.2M genome-equivalents by 2025
CadenceFigure curated · capacity re-pulled from ENA · validated 2026-06-25
Three quantities share one axis. The dark line is Stephens' own (generous) count of human genomes; copper is our strict count of distinct PEOPLE whole-genome-sequenced; teal is genome-equivalents of sequencing DATA (total bases / one genome). The 100M-2B headline is a head-count, so copper is the like-for-like. Teal sits above copper because a 30x genome is ~30 genome-equivalents and much human sequencing is exome, transcriptome, amplicon or single-cell rather than whole genomes — so reading data-equivalents as a count of people is a category error.
The capacity curve is a lower bound on data produced. It sums only human reads submitted to the open ENA/SRA read archive. A large and growing share of clinical and biobank sequencing sits in controlled-access repositories (EGA, dbGaP) or never leaves a hospital/company, so it is not counted. Stephens measured data PRODUCED; we can only measure data SUBMITTED, which is smaller. first_public also dates archive release, not when the sample was sequenced.
Genome-equivalents are a volume proxy, not a genome count. We take total human bases (ENA tax_eq(9606), all read types) divided by one 3.2 Gb genome. That is the right unit for Stephens' data-volume comparison, but it is not a number of sequenced individuals: redundant resequencing, deep coverage and non-WGS assays all inflate bases per person. Choosing a different divisor (e.g. a 30x genome) would shift the curve by that factor.
Stephens' Figure 1 is recreated, not the original data. His historical points and the three projection lines are read off the published log-scale figure (values approximate) and computed from a single 2015 baseline. His 100M-2B headline is a stated range, not a point estimate, and the paper's 'genomes' unit is itself ambiguous between sequencing capacity and a count of people.
The ~2M genome count is conservative and ultimately unknowable. The distinct-genome curve reuses a deliberately conservative ~2M lower bound (Berkeley Genomics). Component cohorts (UK Biobank, All of Us, gnomAD, Genomics England) must NOT be summed — severe double-counting — the true figure is likely 2-5M, and because most clinical genomes never become public it can't be counted exactly. Early-year anchors are read off milestones, so the curve is sparse before 2015.
How we read it. The 2015 forecast overshot either way — ~16-322x in data terms, ~50-1000x in distinct people — so the headline number was too optimistic on both measures. But the paper's core thesis held: genomics did become an exabyte-scale field. The miss was the timeline and the human-genome count, held back by cost, clinical value, consent and reimbursement rather than sequencer throughput. Treat every extrapolation on this page as a trend, not a destiny — that caution is the whole point of revisiting it.
Most-generous reading — human sequence data in the archive (genome-equivalents)
Year
Cumulative bases (ENA/SRA)
Genome-equivalents
2026 (partial)
21.5 Pb
6,712,727
2025
19.9 Pb
6,208,471
2024
15.8 Pb
4,926,727
2023
12.4 Pb
3,870,250
2022
9.4 Pb
2,939,914
2021
6.9 Pb
2,151,534
2020
6.1 Pb
1,914,824
2019
4.3 Pb
1,329,606
2018
2.8 Pb
875,831
2017
2.1 Pb
658,938
2016
1.5 Pb
476,887
2015
1.0 Pb
327,977
2014
680 Tb
212,569
2013
450 Tb
140,671
2012
199 Tb
62,242
2011
92 Tb
28,781
2010
40 Tb
12,360
The first of our revised predictions: rebuild an influential, ageing forecast from public archives and check it against what actually happened. Only aggregate counts are published. · back to all research data