← All research data

SeqDesk · original analysis

Raw reads and assembled genomes

Sequence data reaches ENA as raw reads, as assembled genomes, or both. This shows the mix each top submitting center deposits — assemblies per 100 read sets — straight from the public ENA archive.

43M
raw read sets in ENA
5.1M
assembled genomes
12
assemblies per 100 read sets
3/16
centers above the archive average
1.ETH Zurich CH · academic
153.7/100
2.Quadram Institute Bioscience GB · research institute
59.5/100
3.Washington University (McDonnell Genome Institute) US · academic
19/100
4.INRAE FR · research institute
8.1/100
5.BGI CN · commercial
5.6/100
6.Broad Institute of MIT and Harvard US · research institute
3.4/100
7.Wellcome Sanger Institute GB · research institute
1.8/100
8.CNRS (France) FR · national genomics
1/100
9.Statens Serum Institut DK · public health
0.6/100
10.Swiss Pathogen Surveillance Platform (SPSP) CH · consortium
0.5/100
11.Baylor College of Medicine (Human Genome Sequencing Center) US · academic
0.4/100
12.University of California San Diego US · academic
0.1/100
13.COVID-19 Genomics UK Consortium (COG-UK) GB · consortium
0/100
14.UK Health Security Agency (Colindale) GB · public health
0/100
15.Public Health Wales / Pathogen Genomics Unit GB · public health
0/100
16.New York Genome Center US · research institute
0/100

Each bar is how many assembled genomes a center deposits per 100 raw read sets; the grey tick is the whole-archive average (12 per 100). The mix varies widely with the kind of sequencing a center does — pathogen surveillance and clinical isolates are often shared as raw reads, while genome and environmental projects more often deposit assemblies. Both are valid: raw reads are the right deposit for many uses, and a rate can exceed 100 because a genome is built from many reads and assemblies can be submitted without raw reads. A read set and an assembly are not one-to-one — a genome is assembled from many read sets, and assemblies can be submitted without raw reads — so this is a submission-volume ratio, not a per-sample share. Which mix is appropriate depends on the work; raw reads are the right deposit for many uses.

CenterRaw read setsAssembliesPer 100 read sets
ETH Zurich124K190K153.7
Quadram Institute Bioscience57K34K59.5
Washington University (McDonnell Genome Institute)69K13K19
INRAE114K9.2K8.1
BGI116K6.5K5.6
Broad Institute of MIT and Harvard422K14K3.4
Wellcome Sanger Institute3.8M70K1.8
CNRS (France)86K9061
Statens Serum Institut527K2.9K0.6
Swiss Pathogen Surveillance Platform (SPSP)149K7310.5
Baylor College of Medicine (Human Genome Sequencing Center)161K6240.4
University of California San Diego468K4550.1
COVID-19 Genomics UK Consortium (COG-UK)591K00
UK Health Security Agency (Colindale)158K00
Public Health Wales / Pathogen Genomics Unit153K00
New York Genome Center50K00
Data sourceLive counts from the ENA Portal APIresult=read_run (raw reads) and result=assembly (genome assemblies). Exact whole-archive totals, aggregate-only.
Finishing rateassemblies ÷ read sets × 100, per center, matched across self-reported center_name aliases.
Important caveatA read set and an assembly are not 1:1 — a genome is assembled from many reads, and assemblies can be deposited without raw reads. This is a submission-volume ratio, not a per-sample share (so a rate can exceed 100).
Update cadenceRe-counted weekly · latest snapshot 2026-06-29 · code: scripts/check-assembly-rate.mjs
CompanionsPairs with the facility metadata league table and the volume leaderboard.

Raw reads and assembled genomes are different deposit types, each appropriate for different work; this simply shows the balance each center strikes. Re-counted weekly straight from the ENA Portal API (result=read_run and result=assembly, center_name filters). Only aggregate counts are published. · back to all research data

Report errorpmu15@helmholtz-hzi.de