SeqDesk · original analysis
Sequence data reaches ENA as raw reads, as assembled genomes, or both. This shows the mix each top submitting center deposits — assemblies per 100 read sets — straight from the public ENA archive.
Each bar is how many assembled genomes a center deposits per 100 raw read sets; the grey tick is the whole-archive average (12 per 100). The mix varies widely with the kind of sequencing a center does — pathogen surveillance and clinical isolates are often shared as raw reads, while genome and environmental projects more often deposit assemblies. Both are valid: raw reads are the right deposit for many uses, and a rate can exceed 100 because a genome is built from many reads and assemblies can be submitted without raw reads. A read set and an assembly are not one-to-one — a genome is assembled from many read sets, and assemblies can be submitted without raw reads — so this is a submission-volume ratio, not a per-sample share. Which mix is appropriate depends on the work; raw reads are the right deposit for many uses.
| Center | Raw read sets | Assemblies | Per 100 read sets |
|---|---|---|---|
| ETH Zurich | 124K | 190K | 153.7 |
| Quadram Institute Bioscience | 57K | 34K | 59.5 |
| Washington University (McDonnell Genome Institute) | 69K | 13K | 19 |
| INRAE | 114K | 9.2K | 8.1 |
| BGI | 116K | 6.5K | 5.6 |
| Broad Institute of MIT and Harvard | 422K | 14K | 3.4 |
| Wellcome Sanger Institute | 3.8M | 70K | 1.8 |
| CNRS (France) | 86K | 906 | 1 |
| Statens Serum Institut | 527K | 2.9K | 0.6 |
| Swiss Pathogen Surveillance Platform (SPSP) | 149K | 731 | 0.5 |
| Baylor College of Medicine (Human Genome Sequencing Center) | 161K | 624 | 0.4 |
| University of California San Diego | 468K | 455 | 0.1 |
| COVID-19 Genomics UK Consortium (COG-UK) | 591K | 0 | 0 |
| UK Health Security Agency (Colindale) | 158K | 0 | 0 |
| Public Health Wales / Pathogen Genomics Unit | 153K | 0 | 0 |
| New York Genome Center | 50K | 0 | 0 |
result=read_run (raw reads) and result=assembly (genome assemblies). Exact whole-archive totals, aggregate-only.center_name aliases.scripts/check-assembly-rate.mjsRaw reads and assembled genomes are different deposit types, each appropriate for different work; this simply shows the balance each center strikes. Re-counted weekly straight from the ENA Portal API (result=read_run and result=assembly, center_name filters). Only aggregate counts are published. · back to all research data