← All research data

SeqDesk · interactive

Explore a sequencing center

Pick any of the top submitting centers and see its aggregated profile in one place — how completely it describes its samples, how often it turns reads into assembled genomes, its volume and platform mix — all measured from public ENA data, against the whole-archive baseline. No single center is featured; it is a lens on whichever one you choose.

Choose an institute
Wellcome Sanger Institute
Hinxton · United Kingdom · research institute · est. 1992 · website ↗
3.5M
samples in ENA
#14 / 16
metadata league · 66%
#7 / 16
finishing genomes · 1.8/100
How completely are samples described?

Each row is one metadata field that supports FAIR reuse; which fields are applicable can depend on sample type and disclosure constraints. The dark bar is the share of Wellcome’s samples that have it; the grey tick marks the whole-archive average, so anything past the tick is above the field as a whole.

Geographic location
66%baseline 59%+7 pts

Share of samples that record the country of origin — the most basic “where did this come from”. Often the only geography a record carries.

Collection date
65%baseline 56%+9 pts

Share with a real sampling date (a valid day between 1900 and today), not left blank or filled with a placeholder like “not provided”.

Lat/lon coordinates
0%baseline 26%-26 pts

Share carrying precise latitude/longitude, not just a country — what is needed to place a sample precisely on a map. Often the least frequently populated field.

MIxS environment
0%baseline 12%-12 pts

Share tagged with a standardised habitat/biome term from the MIxS vocabulary. Mostly applies to environmental samples, so it is reported but kept out of the rank.

Real MIxS package
0%baseline 66%-66 pts

Share submitted under an environment-specific MIxS checklist (e.g. host-associated, soil, water) rather than the general-purpose “Generic” checklist, which requests fewer mandatory fields. Generic is a valid choice when no package fits.

Platform mix
Illumina 95%Nanopore 2%PacBio 0%Ion Torrent 0%Other 3%

Every percentage is the share of Wellcome’s public ENA samples; the grey tick is the whole-archive baseline. See the full metadata league table and finishing-rate ranking.

SourcePublic records in the ENA Portal API (sample-level count queries) — the same committed snapshots behind the metadata league, finishing-rate and facilities analyses.
How it’s measuredFor each center we count its public ENA samples, then for every field count how many actually carry a value — subtracting known placeholders (“not provided”, “not collected”, “missing”, …) — and divide by the total. Centers are matched on ENA’s center_name via a hand-curated list of name aliases.
ScopeThe top submitting centers by volume (those with at least a few hundred public samples). The grey tick on each bar is the whole-archive average for that field.
RefreshedRegenerated automatically every Monday from the live archive and committed as a snapshot. This view reflects ENA as of 2026-06-29; all computation runs in your browser — nothing is fetched or sent when you pick a center.
How accurate is this — and where it can be wrong

Treat these figures as a directional estimate, not an audited statistic. A few things can skew them:

  • Name matching. ENA’s center_name is free text and inconsistent — the same institute appears under many spellings, and one label can cover several labs. The curated aliases can miss a center’s submissions or sweep in ones that aren’t theirs, which moves the totals up or down.
  • Present ≠ correct. We can only check whether a field is filled and not a known placeholder. A populated field may still hold a wrong, vague, or imprecise value — so real completeness is likely a little lower than shown.
  • ENA-only & point-in-time. Only public ENA submissions count — embargoed data, and anything deposited solely to NCBI or DDBJ, is invisible here. The archive also changes daily, so a weekly snapshot lags reality.
  • Submission behaviour, not capacity. The mix is surveillance-heavy and a single large project can dominate a center’s profile. This reflects what a center chose to deposit and how — not its true sequencing output or capability.

Spotted your institute looking off? It is almost always a name-alias gap — let us know and we can fix the mapping.

A per-center lens on the facility analyses — built for picking your own institute out of the crowd. · back to all research data