← All research data

SeqDesk · revised predictions

Did the ENA really keep doubling every 32 months?

In 2016 the European Nucleotide Archive reported its raw sequence data was doubling every 32 months. We rebuilt ENA's actual base-volume history and drew that 2016 forecast forward — and the archive grew even faster than it predicted.

32 mo
the 2016 forecast — ENA raw bases double every 32 months
3 → 67 Pbp
ENA public raw bases, 2016 → 2025 (108 Pbp incl. embargoed)
~2×
above the 32-month forecast by 2025 — reality ran ahead
~45 mo
but ENA's own doubling has since lengthened (decelerating)
10T100T1.0P10P100P20102015202020252016 — forecast made
ENA raw sequence data (bases)32-month-doubling forecast (from 2016 baseline)

ENA's raw sequence data in bases, on a log axis (copper), against the 32-month-doubling forecast drawn from the 2016 baseline where the prediction was made (dashed, in the same copper, starting at the marked year). Reality tracks the forecast closely through 2018, then pulls above it: by 2025 ENA holds ~67 petabases of public bases (108 including embargoed) versus the ~31 petabases a literal 32-month line predicted. The forecast named the right exponential but, in bases, under-counted the deluge. (The live dataset count grows slower — a ~56-month doubling — and ENA's headline storage figure is a different quantity.)

What held upThe deluge was real, and then some. ENA's raw sequence data kept doubling exponentially across the whole 2016 → 2025 horizon — from 3 petabases to ~67 petabases of public bases (108 Pbp including embargoed). The 32-month forecast named exactly the right phenomenon, and the archive's own operators framing it as an exponential proved sound.
What brokeIf anything it under-shot. By 2025 the archive held about 2× more bases than a literal 32-month line predicted (~31 Pbp). The slowdown such a curve implies did eventually arrive — ENA's self-reported doubling time lengthened from ~10 months (2013) to 32 (2016) to ~45 (2025) — but later and milder than assumed. Mind the units: ENA's 2024 'doubles in just over 3 years' refers to 62.8 PB of storage bytes, which grow slower than bases because of compression, and the live dataset COUNT doubles only ~56 months.
ReportedDoubling timeWhat it refers to
2013~10 monthsraw NGS data (bases)
2015~20 monthsENA data (bases)
2016 — the forecast32 monthsraw sequence data (bases)
2024~3 years62.8 PB storage (bytes, not bases)
2025~45 monthsnucleotides (Nature 2025)
From petabytes to petabases — ENA's growth
  • 12009The petabyte era begins — ENA reports the first 3 months of next-generation submissions brought 10 TB of data — an eighth of everything accumulated in the prior 28 years.
  • 22013~10-month doubling — ENA holds 310 trillion bases and reports raw NGS data doubling roughly every 10 months — growth at its steepest.
  • 32016The 32-month forecast — The ENA 2016 update holds 3 petabases and states raw sequence data is 'doubling currently at 32 months' — the forecast this card revisits.
  • 420188 petabases — ENA holds 8×10¹⁵ base pairs of read data across 1.5 million taxa — still tracking just above the 32-month line.
  • 52021~26 petabases — INSDC raw reads reach ~25.6 petabases of public bases (SRA mirror; ENA tracks within a few percent), already pulling clear of the forecast.
  • 62025~67 petabases — ENA holds ~67 petabases of public raw bases (108 Pbp including embargoed) — roughly 2× what the 2016 32-month line predicted — even as its self-reported doubling lengthens to ~45 months.

SeqDesk · revised predictions. The real line is ENA's total raw-sequence bases (nucleotides), read from successive 'European Nucleotide Archive in YYYY' NAR updates: 50 Tbp (2010), 310 Tbp (2012), 570 Tbp (2013), 1.2 Pbp (2014), 3 Pbp (2016), 8 Pbp (2018), and ~67 Pbp public / 108 Pbp total in Jan 2025 (Karasikov et al., Nature 2025). The 2021 point (~25.6 Pbp public) is the INSDC/SRA mirror, used because ENA's 2021 paper base figure was a likely typo; ENA and SRA track within a few percent. The forecast line is the 2016 baseline projected at 32-month doubling: N(t) = 3 Pbp × 2^(months / 32), reaching ~31 Pbp by 2025. CRITICAL UNIT NOTE: everything plotted is BASES (petabases, 10¹⁵ bp), not storage bytes — ENA's widely-quoted 62.8 PB and its '~3 years' doubling (2024) are storage bytes, a different and slower-growing quantity. The live ena-read-run count is a separate dataset-count cross-check (doubling ~56 months) and is not the byte/base volume the forecast described.

What it measuresA doubling-rate forecast checked in its own units (bases) — ENA raw-sequence growth, 2016 → 2025
The forecastENA 2016 (Toribio et al., NAR 45:D32) — raw sequence data doubling every 32 months, baseline 3 Pbp
Realized (bases)ENA NAR updates + Karasikov et al. (Nature 2025) — ~67 Pbp public / 108 Pbp total by Jan 2025
Live cross-checkena-read-run = 42.6M raw-read datasets (dataset count; doubles ~56 months, slower than bases)
CadenceCurated — literature + live archive count · validated 2026-06-25
Methodscripts/check-revised-prediction.mjs
YearENA raw bases
2025~67 Pbp public (108 Pbp total)
2021~25.6 Pbp (INSDC/SRA mirror)
20188 Pbp
20163 Pbp (forecast baseline)
20141.2 Pbp
2013570 Tbp
2012310 Tbp
201050 Tbp

A revised prediction: an archive's own doubling forecast, checked in its own units — and quietly exceeded. Only aggregate counts are published. · back to all research data