← All research data

SeqDesk · original analysis

How completely do top centers describe their samples?

The world's highest-volume submitting centers, compared not by how much they deposit but by how completely their public ENA samples are described — geography, collection date, and a real MIxS reporting standard. A description of practice, not a verdict.

16
top centers scored
91%
most complete — Chinese Academy of Sciences
80%
whole-archive baseline
5/16
above the baseline
1.Chinese Academy of Sciences CN · national genomics
91%
2.Broad Institute of MIT and Harvard US · research institute
83%
3.BGI CN · commercial
82%
4.INRAE FR · research institute
82%
5.New York Genome Center US · research institute
82%
6.Quadram Institute Bioscience GB · research institute
77%
7.Technical University of Denmark (DTU) DK · academic
77%
8.Washington University (McDonnell Genome Institute) US · academic
75%
9.UK Health Security Agency (Colindale) GB · public health
75%
10.Public Health Wales / Pathogen Genomics Unit GB · public health
75%
11.Statens Serum Institut DK · public health
74%
12.CNRS (France) FR · national genomics
72%
13.ETH Zurich CH · academic
70%
14.Wellcome Sanger Institute GB · research institute
66%
15.University of California San Diego US · academic
65%
16.Baylor College of Medicine (Human Genome Sequencing Center) US · academic
64%

Each bar is a center’s completeness score — the mean of a persistent accession (near-universal once deposited), a real MIxS package rather than the permissive “Generic” default, and the universal context that helps make a sample reusable (geographic location + collection date). The grey tick is the whole-archive average (80%). Rank is deliberately blind to volume: the biggest depositor is not always the one whose records are most complete. What a record needs to carry depends on the sample type, so read this as a description of practice, not a verdict on a center. Mean of: a persistent accession (≈universal), a real MIxS package (not Generic), and universal context (geographic location + collection date). MIxS environment & coordinates are reported but kept out of the score because they mostly apply to environmental samples — and what a record needs depends on the sample type.

CenterSamplesGeographic locationCollection dateReal MIxS packageMIxS environmentScore
Chinese Academy of Sciences394K74.4%76.9%89.4%24.4%91%
Broad Institute of MIT and Harvard490K59.2%57.1%74.6%1.2%83%
BGI82K60.3%39.5%80.8%14.0%82%
INRAE162K83.3%76.5%50.4%36.1%82%
New York Genome Center17K65.2%65.4%65.2%0.0%82%
Quadram Institute Bioscience91K97.0%96.1%12.1%16.0%77%
Technical University of Denmark (DTU)121K92.4%80.5%22.1%52.1%77%
Washington University (McDonnell Genome Institute)167K53.2%50.5%51.5%20.9%75%
UK Health Security Agency (Colindale)489K100.0%100.0%0.0%0.0%75%
Public Health Wales / Pathogen Genomics Unit246K100.0%100.0%0.0%0.0%75%
Statens Serum Institut592K97.1%95.5%0.9%1.6%74%
CNRS (France)57K47.5%45.0%42.7%13.0%72%
ETH Zurich413K75.7%77.7%4.3%21.9%70%
Wellcome Sanger Institute3.6M66.6%65.6%0.1%0.0%66%
University of California San Diego78K25.1%31.6%35.3%11.2%65%
Baylor College of Medicine (Human Genome Sequencing Center)73K23.4%23.2%36.2%13.3%64%
Data sourceLive counts from the ENA Portal API (result=sample) — exact whole-archive totals, not a sample, survey, or estimate. Only aggregate percentages are published; no sample identifiers leave the archive.
Which centersThe highest-volume curated submitting centers from the facilities leaderboard, each matched across its self-reported center_name aliases (OR-merged).
How each % is computedFor each center: samples carrying a field ÷ that center’s total samples. e.g. center_name=… AND country="*" for geography, a valid collection_date range, or ncbi_reporting_standard other than “Generic”.
ScoreThe mean of three fractions: a persistent accession (≈universal once deposited), a real MIxS package (not “Generic”), and universal context (geographic location + collection date).
Why exclude environment?A MIxS environment only applies to environmental samples, so scoring a human/clinical center on it would be unfair — it is shown in the table, not baked into the rank.
Update cadenceRe-counted weekly by an automated job · latest snapshot 2026-09-07 · code: scripts/check-facility-metadata.mjs
CompanionPairs with the FAIR scorecard — compare your own facility against these real centers.

The volume leaderboard asks who deposits the most; this asks how completely samples are described — and the two are rarely the same center. Re-counted weekly straight from the ENA Portal API (result=sample, center_name filters). Only aggregate percentages and counts are published. · back to all research data

Report errorpmu15@helmholtz-hzi.de