← All research data

SeqDesk · revised predictions

How long until we had a structure for every protein? The forecasts said decades. AlphaFold did it in a week.

For two decades, structural biology was organized around a widening gap: protein sequences piled up exponentially while experimentally solved structures crawled forward a few thousand a year. The experimental forecast was basically right — the PDB grew on trend to ~256k. But in July 2022 AlphaFold dumped 214 million predicted structures into a public database, and the question the forecast answered simply stopped mattering.

256k
experimental structures in the PDB (2026)
214M
AlphaFold predicted structures
~840×
predicted structures vs every experimental one ever solved
<0.1% → ~100%
sequence–structure gap, before → after AlphaFold
101001.0K10K100K1.0M10M100M1.0B19801990200020102020
Experimental structures (RCSB PDB)AlphaFold predicted structures

Both lines on a log axis. The teal line is experimental structures in the PDB — a smooth, near-linear climb across 50 years to 256,006 by 2026, almost exactly the pre-2020 trajectory. The purple line is AlphaFold: a near-vertical cliff in 2021–2022 from ~360,000 to 214 million predicted structures — roughly 840× every experimental structure ever solved. The forecast of experimental growth was sound; what broke is the premise that experiment was the only road to a structure.

What held upExperimental structure determination grew almost exactly as a pre-2020 extrapolation predicted: the PDB reached ~256,006 structures by 2026 at ~15,000–17,600 per year, on trend. A forecast of experimental structures — and the multi-decade horizon to characterize protein space by experiment alone — was sound.
What brokeThe forecasts assumed the only route to a structure was experiment, so the sequence–structure gap (<0.1% of sequences covered) would persist for decades. AlphaFold made that premise obsolete: 214 million predicted structures (~840× all experimental structures ever solved) collapsed the gap to near-complete UniProt coverage in roughly a week, July 2022. The forecast wasn't beaten by 'more of the same, faster' — a method change (ML prediction) retired the question.
YearExperimental structures (PDB)Structural coverage of known sequences
2004~28,000~2%
2010~69,000~0.7%
2020~173,000<0.1% — gap widening
2022~200,000 experimental + 214M predicted~100% by prediction
2026256,006 experimentalnear-complete (predicted)
From a $764M initiative to a one-week sweep
  • 12000Protein Structure Initiative begins — NIGMS launches the PSI (2000–2015, ~$764M) to bring high-throughput experiment to bear on protein fold space; output across all phases is only a few thousand unique structures.
  • 22015PSI ends, gap intact — PSI:Biology closes after NIH declines renewal; fewer than ~0.1% of known sequences have an experimental structure — the gap is wider, not narrower, than at the start.
  • 32020AlphaFold2 wins CASP14 — AlphaFold2 posts a median GDT_TS of 92.4/100 and best predictions on 88 of 97 targets; organizers call the 50-year protein-folding problem essentially solved.
  • 42021AlphaFold DB launches — DeepMind + EMBL-EBI release ~360,000 predicted structures (human proteome + 20 model organisms); ~1 million by December 2021.
  • 52022214 million predicted structures — The database expands to 214,684,311 structures — 'almost the whole of UniProt', ~200× the prior release and ~1,000× the ~190k experimental structures in the PDB at the time.
  • 62026PDB on trend at ~256k — Experimental structures reach 256,006 by 2026 at ~15–18k/year — exactly the pre-2020 trajectory. The experimental forecast was right; the question changed.

SeqDesk · revised predictions. Predicted structures are not experimental structures: AlphaFold's 214 million models are computational predictions with per-residue confidence (pLDDT) that varies widely, and they cover single chains in isolation — no ligands, cofactors, complexes, alternative conformations or dynamics. So 'near-complete coverage of UniProt' is coverage by prediction, not experimental validation — the AlphaFold and PDB series should never be summed. The sequence denominator is itself slippery: UniProt counts churn with redundancy clustering and reclassification (the SeqDesk uniprot-total series has a known anomalous drop in 2026), so the gap percentages use the stable, curated Swiss-Prot set and are approximate. The small delta between the 214,684,311 July-2022 release and the tracker's 214,683,829 is ordinary inter-release revision.

What it measuresA pre-2020 structure-growth forecast vs a method change — experimental PDB growth and the sequence–structure gap
The forecast / framingProtein Structure Initiative (NIGMS, 2000–2015, ~$764M) and the widening sequence–structure gap (<0.1% covered pre-AlphaFold)
Experimental realizedReuses the pdb-structures metric — 256,006 structures, on the pre-2020 trend
The disruptionReuses alphafold (214,683,829 predicted); swissprot supplies the sequence denominator
CadenceCurated — literature + live archive counts · validated 2026-06-25
Methodscripts/check-revised-prediction.mjs
YearCumulative structures
2026 (live)256,006
2024229,635
2022199,677
2020172,809
2014105,050
201069,486
200013,583
1990507
197613

A revised prediction: sometimes a forecast isn't beaten — its question is retired. The PDB grew exactly as expected; AlphaFold made 'how many structures do we have?' a different question. Only aggregate counts are published. · back to all research data