← All research data

SeqDesk · original analysis

How completely do top centers describe their samples?

The world's highest-volume submitting centers, compared not by how much they deposit but by how completely their public ENA samples are described — geography, collection date, and a real MIxS reporting standard. A description of practice, not a verdict.

16
top centers scored
83%
most complete — New York Genome Center
80%
whole-archive baseline
4/16
above the baseline
1.New York Genome Center US · research institute
83%
2.Broad Institute of MIT and Harvard US · research institute
83%
3.BGI CN · commercial
82%
4.INRAE FR · research institute
82%
5.Technical University of Denmark (DTU) DK · academic
77%
6.Quadram Institute Bioscience GB · research institute
77%
7.Comenius University Science Park SK · research institute
76%
8.Washington University (McDonnell Genome Institute) US · academic
75%
9.UK Health Security Agency (Colindale) GB · public health
75%
10.Public Health Wales / Pathogen Genomics Unit GB · public health
75%
11.Statens Serum Institut DK · public health
74%
12.CNRS (France) FR · national genomics
72%
13.ETH Zurich CH · academic
70%
14.Wellcome Sanger Institute GB · research institute
66%
15.University of California San Diego US · academic
65%
16.Baylor College of Medicine (Human Genome Sequencing Center) US · academic
64%

Each bar is a center’s completeness score — the mean of a persistent accession (near-universal once deposited), a real MIxS package rather than the permissive “Generic” default, and the universal context that helps make a sample reusable (geographic location + collection date). The grey tick is the whole-archive average (80%). Rank is deliberately blind to volume: the biggest depositor is not always the one whose records are most complete. What a record needs to carry depends on the sample type, so read this as a description of practice, not a verdict on a center. Mean of: a persistent accession (≈universal), a real MIxS package (not Generic), and universal context (geographic location + collection date). MIxS environment & coordinates are reported but kept out of the score because they mostly apply to environmental samples — and what a record needs depends on the sample type.

CenterSamplesGeographic locationCollection dateReal MIxS packageMIxS environmentScore
New York Genome Center16K67.7%67.9%67.7%0.0%83%
Broad Institute of MIT and Harvard487K59.8%57.3%74.7%1.2%83%
BGI81K59.8%38.4%81.9%12.8%82%
INRAE157K83.1%76.1%51.4%36.4%82%
Technical University of Denmark (DTU)111K91.7%78.8%24.0%48.0%77%
Quadram Institute Bioscience91K97.0%96.1%11.9%15.6%77%
Comenius University Science Park26K100.0%100.0%4.5%0.0%76%
Washington University (McDonnell Genome Institute)166K53.0%50.3%51.3%21.0%75%
UK Health Security Agency (Colindale)489K100.0%100.0%0.0%0.0%75%
Public Health Wales / Pathogen Genomics Unit246K100.0%100.0%0.0%0.0%75%
Statens Serum Institut590K97.1%95.5%0.9%1.6%74%
CNRS (France)55K48.2%45.6%44.1%13.5%72%
ETH Zurich412K75.7%77.7%4.3%21.9%70%
Wellcome Sanger Institute3.5M66.4%65.4%0.1%0.0%66%
University of California San Diego78K25.1%31.6%35.2%11.2%65%
Baylor College of Medicine (Human Genome Sequencing Center)72K22.9%23.3%35.8%13.4%64%
Data sourceLive counts from the ENA Portal API (result=sample) — exact whole-archive totals, not a sample, survey, or estimate. Only aggregate percentages are published; no sample identifiers leave the archive.
Which centersThe highest-volume curated submitting centers from the facilities leaderboard, each matched across its self-reported center_name aliases (OR-merged).
How each % is computedFor each center: samples carrying a field ÷ that center’s total samples. e.g. center_name=… AND country="*" for geography, a valid collection_date range, or ncbi_reporting_standard other than “Generic”.
ScoreThe mean of three fractions: a persistent accession (≈universal once deposited), a real MIxS package (not “Generic”), and universal context (geographic location + collection date).
Why exclude environment?A MIxS environment only applies to environmental samples, so scoring a human/clinical center on it would be unfair — it is shown in the table, not baked into the rank.
Update cadenceRe-counted weekly by an automated job · latest snapshot 2026-06-29 · code: scripts/check-facility-metadata.mjs
CompanionPairs with the FAIR scorecard — compare your own facility against these real centers.

The volume leaderboard asks who deposits the most; this asks how completely samples are described — and the two are rarely the same center. Re-counted weekly straight from the ENA Portal API (result=sample, center_name filters). Only aggregate percentages and counts are published. · back to all research data

Report errorpmu15@helmholtz-hzi.de