← All research data

SeqDesk · revised predictions

Did we sequence 2.5 million plant and animal genomes by 2025?

The same 2015 paper behind our genome-count card also predicted at least 2.5 million plant and animal genomes by 2025. Count every genome NCBI holds and you clear it; count only the plants and animals it was actually about and you fall ~64× short.

2.5M
plant + animal genomes projected by 2025 (Stephens 2015)
~39k
plant + animal genome assemblies actually built
~64×
short of the prediction as written
3.35M
all genome assemblies by end-2025 (every organism)
1001.0K10K100K1.0M10M2010201520202025
All genome assemblies NCBI holdsStephens 2.5M target (2025)Plant + animal assemblies only

On a log axis: the copper line is every genome assembly NCBI holds — it blew past the dashed 2.5 million target during 2024 and stands at 4.18M today. But that line is overwhelmingly bacteria, viruses and archaea from metagenomic and pathogen-surveillance pipelines. The red line is the plants and animals the prediction was actually about — climbing from a few hundred genomes in 2012 to about 39,000 by 2026, but staying ~64× below the target the whole way, and concentrated in a few thousand re-sequenced species rather than the ~1.2-million-species coverage the method assumed.

What held upThe order of magnitude and the direction were broadly right: total public genome assemblies did blow past 2.5 million by 2025 (3.35M at end-2025), so as a raw 'millions of genomes' headline the forecast lands — a striking contrast with the same paper's human-genome leg (100M–2B projected, ~2M realized, 50–1000× over).
What brokeRead literally, the prediction was about PLANT AND ANIMAL genomes covering most of ~1.2 million described species. Reality there is ~39,000 plant+animal assemblies (~64× short), concentrated in a few thousand re-sequenced species rather than broad species coverage. The 2.5M that did materialize is overwhelmingly bacterial, viral and archaeal — driven by metagenomic and pathogen-surveillance pipelines the 2015 framing did not foreground. It is 'the leg they got right' only if you quietly swap the subject of the sentence.
What you countBy 2025vs the 2.5M target
Plant + animal assemblies (as written)~39,000~64× short
Every genome assembly NCBI holds3.35M (end-2025)~1.3× over — met
Distinct plant/animal species covereda few thousandfar below the ~1.2M-species premise
How a plant-and-animal forecast became a microbial one
  • 12015The forecast — Stephens et al. estimate 'at least 2.5 million plant and animal genome sequences by 2025', anticipating coverage of most of the ~1.2 million described plant and animal species plus thousands of individuals of high-value species.
  • 22019The metagenome surge — NCBI genome assemblies jump from ~298k (end-2018) to ~639k (end-2019) as metagenome-assembled genomes (MAGs) and large-scale microbial pipelines come online — the growth that would carry the all-organism count, not plants and animals.
  • 320211,000,000th assembly — NCBI passes one million genome assemblies — overwhelmingly bacterial and viral, accelerated by SARS-CoV-2 genomic surveillance.
  • 42024All-organism count clears 2.5M — Total genome assemblies cross the 2.5 million mark during 2024 (2.81M at end-2024) — the forecast's number, reached by changing what is counted.
  • 52026~39k plants & animals vs 4.18M total — NCBI holds 4.18 million assemblies overall but only ~38,984 plant+animal assemblies (29,705 animals + 9,279 green plants) — the prediction's actual subject sits ~64× below its target.

SeqDesk · revised predictions. The 2.5 million figure is a stated number in the paper text, so the forecast side needs no rebuild. The realized all-organism count reuses the curated ncbi-assemblies tracker metric (NCBI Datasets, taxon root). The like-for-like plant+animal line is reconstructed from NCBI Datasets: for each year-end we compute cumulative = total − count(first_release_date ≥ next Jan 1), summed over Metazoa (taxon 33208) and Viridiplantae (taxon 33090); the current totals are 29,705 + 9,279 = 38,984 (June 2026). It is a current-snapshot reconstruction (assemblies removed since release are not counted) and is NOT yet a standing tracker metric, so it is frozen and will drift. These are assembly counts, which double-count species (many assemblies per high-value species) and so overstate the species coverage the method assumed. We report both lines and label the all-organism figure as a generous over-count that changes the subject of the sentence.

What it measuresA non-human leg of Stephens 2015 — projected plant+animal genomes vs reality, the companion to our genome-count card
The paperStephens et al. 2015, PLOS Biology — '≥2.5 million plant and animal genome sequences by 2025'
All-organism realizedReuses the ncbi-assemblies metric — 3.35M (end-2025), 4.18M live
Plant+animal realizedNCBI Datasets taxon queries: Metazoa (33208) + Viridiplantae (33090) = 38,984 as of June 2026
CadenceCurated — literature + live archive count · validated 2026-06-25
Methodscripts/check-revised-prediction.mjs
YearGenome assemblies
2026 (live)4,183,542
20253,351,364
20242,814,679
20232,204,898
20221,658,238
20201,013,201
2018298,005
2015114,446
201236,218
200914,421

A revised prediction: same paper as our genome-count card, a different axis — and a reminder that a forecast can be 'right' only by changing what you count. Only aggregate counts are published. · back to all research data