← All metrics

Sequence archives

ENA raw read datasets

EMBL-EBI / European Nucleotide Archive

43M
runs · 2026-07-20
Gompertz
best-fit model
46M
projected 2027
22M30M38M46M20222023202420252026202746M 2027
Gompertz (S-curve) fitbest fit (AICc)N(t) = L·e^{−e^{−k(t − t₀)}}
carrying capacity L 79Msteepness k 0.158 / yrinflection t₀ 2023 0.9969forecast 46M by 2027

Forecast = 10% of the 5-yr record (0.8 yr ahead). Anchored on the latest value; levels off toward L ≈ 79M.

Timeline — what shaped the curve
  • 12007-01SRA launched — NCBI established the Sequence Read Archive in 2007 as the first public repository for raw next-generation sequencing reads, the foundation later mirrored by ENA/EBI.
  • 22007-01Illumina NGS arrives — Solexa's Genome Analyzer (acquired by Illumina in 2007) brought gigabase-scale short-read sequencing to market, triggering the explosive growth in submitted read data.
  • 32008-011000 Genomes Project — Launched in January 2008, this international effort to catalogue human variation drove a large early wave of high-coverage read submissions to public archives.
  • 42015-05Nanopore MinION — Oxford Nanopore's portable MinION became commercially available in May 2015, adding real-time long-read sequencing to the data deposited in ENA/SRA.
  • 52020-01-10SARS-CoV-2 genome — The first SARS-CoV-2 genome (Wuhan-Hu-1) was shared publicly on 10-11 January 2020, kicking off the largest pathogen-sequencing surge ever recorded in the archives.

Each record is one submitted sequencing run. Cumulative; only goes up.

SourceEMBL-EBI / European Nucleotide Archive
Update cadenceweekly
Live endpointhttps://www.ebi.ac.uk/ena/portal/api/count?result=read_run&format=json
Parsejson · count
Dateruns
2026-07-2043,021,658
2026-07-1342,924,583
2026-07-0642,834,140
2026-06-2442,613,406
2025-01-0137,000,000
2024-01-0132,000,000
2023-01-0128,000,000
2022-01-0122,000,000

Machine-readable: /api/research-data?metric=ena-read-run · back to all metrics

Report errorpmu15@helmholtz-hzi.de