Pick any of the top submitting centers and see its aggregated profile in one place — how completely it describes its samples, how often it turns reads into assembled genomes, its volume and platform mix — all measured from public ENA data, against the whole-archive baseline. No single center is featured; it is a lens on whichever one you choose.
Choose an institute
Wellcome Sanger Institute
Hinxton · United Kingdom · research institute · est. 1992 · website ↗
3.5M
samples in ENA
#14 / 16
metadata league · 66%
#7 / 16
finishing genomes · 1.8/100
How completely are samples described?
Each row is one metadata field that supports FAIR reuse; which fields are applicable can depend on sample type and disclosure constraints. The dark bar is the share of Wellcome’s samples that have it; the grey tick marks the whole-archive average, so anything past the tick is above the field as a whole.
Geographic location
66%baseline 59%+7 pts
Share of samples that record the country of origin — the most basic “where did this come from”. Often the only geography a record carries.
Collection date
65%baseline 56%+9 pts
Share with a real sampling date (a valid day between 1900 and today), not left blank or filled with a placeholder like “not provided”.
Lat/lon coordinates
0%baseline 26%-26 pts
Share carrying precise latitude/longitude, not just a country — what is needed to place a sample precisely on a map. Often the least frequently populated field.
MIxS environment
0%baseline 12%-12 pts
Share tagged with a standardised habitat/biome term from the MIxS vocabulary. Mostly applies to environmental samples, so it is reported but kept out of the rank.
Real MIxS package
0%baseline 66%-66 pts
Share submitted under an environment-specific MIxS checklist (e.g. host-associated, soil, water) rather than the general-purpose “Generic” checklist, which requests fewer mandatory fields. Generic is a valid choice when no package fits.
How it’s measuredFor each center we count its public ENA samples, then for every field count how many actually carry a value — subtracting known placeholders (“not provided”, “not collected”, “missing”, …) — and divide by the total. Centers are matched on ENA’s center_name via a hand-curated list of name aliases.
ScopeThe top submitting centers by volume (those with at least a few hundred public samples). The grey tick on each bar is the whole-archive average for that field.
RefreshedRegenerated automatically every Monday from the live archive and committed as a snapshot. This view reflects ENA as of 2026-06-29; all computation runs in your browser — nothing is fetched or sent when you pick a center.
How accurate is this — and where it can be wrong
Treat these figures as a directional estimate, not an audited statistic. A few things can skew them:
Name matching. ENA’s center_name is free text and inconsistent — the same institute appears under many spellings, and one label can cover several labs. The curated aliases can miss a center’s submissions or sweep in ones that aren’t theirs, which moves the totals up or down.
Present ≠ correct. We can only check whether a field is filled and not a known placeholder. A populated field may still hold a wrong, vague, or imprecise value — so real completeness is likely a little lower than shown.
ENA-only & point-in-time. Only public ENA submissions count — embargoed data, and anything deposited solely to NCBI or DDBJ, is invisible here. The archive also changes daily, so a weekly snapshot lags reality.
Submission behaviour, not capacity. The mix is surveillance-heavy and a single large project can dominate a center’s profile. This reflects what a center chose to deposit and how — not its true sequencing output or capability.
Spotted your institute looking off? It is almost always a name-alias gap — let us know and we can fix the mapping.
A per-center lens on the facility analyses — built for picking your own institute out of the crowd. · back to all research data