The MIxS iceberg
MIxS is a widely used standard for describing a sequenced sample — one shared vocabulary across 29 ENA/GSC checklists. Scroll down to see how much is shared across all checklists and how much is specific to individual ones.
What do all 29 checklists actually agree on?
Start at the waterline
Picture the 29 checklists as one body of water. The waterline marks what almost every checklist agrees on. Above it is consensus; below it, the specialised vocabulary each community brought of its own.
One field is truly universal
Exactly one field is on all 29 checklists: collection_date. Not the organism, not the location — just when the sample was taken. It sits right at the tip, and it’s the only field every checklist also makes mandatory.
Six fields are near-universal
Widen the bar to “shared by ≥25 of 29” and the tip grows to 6 fields — collection_date, geographic location (×4) and project_name. So both numbers are true: 1 is universal, 6 are near-universal. That tiny tip is the whole shared core.
Below the surface: the 60%
Now draw the rest. 527 fields (60%) sit below the waterline — each used by a single checklist. The shared core is small; the bulk of MIxS fields are specific to individual checklists.
The deep is sorted by domain
The deep isn’t noise. Every one-off field belongs to one community — soil, the human gut, a viral pathogen — so it sorts into domain families (the colours). Lower a sounding line through the whole standard just below.
One field on all of them, almost none shared by most
Across the 29 checklists there are 2,798 field-slots but only 880 distinct fields. Exactly one — collection_date — is on every checklist; just 6 are shared by 25 or more; and 527 (60%) appear in a single checklist. The deep isn’t random, though: each one-off field belongs to a community — soil, the human gut, a viral pathogen — so the tail sorts itself into domain families, which is what the colours in the grid below show.
Lower a sounding line through the standard
A field × checklist heatmap. Each row is a field; each column a checklist. Click any cell to see exactly where that field lives.
The 24 rows above are a domain-balanced core sample; 856 more fields sit in the long tail. Only 75 of 880 fields are ever required in any checklist — the other 805 are never mandatory anywhere.
Hover the grid to read a cell; click a field name or a cell to pin it; click a checklist header to see what it requires.
That shape is why SeqDesk maps each order to a specific checklist instead of one generic form, and why we read metadata completeness against the checklist a sample was actually submitted under.
How completely is a sample actually described?
MIxS stands for Minimum Information about any Sequence — but the enforced minimum is tiny. Of all 2,798 field-slots, only 271 (9.7%) are required; the rest are optional. A checklist can offer 273 fields and still mandate as few as 2. Drag the sort below to see who asks for the most — and who actually requires it.
Sorted by size. The big environmental checklists offer the most fields — yet require almost none.
The real backbone is ~8 fields
75 fields are required somewhere, but only these are enforced across many checklists. The pip shows how often each is required where it appears.
And what they ask for is mostly unstructured: 91% of slots are free text, not controlled vocabularies. Even required slots are only 20% controlled (optional ones 8%) — so a filled-in value often still isn’t standardised.
Data & method. Computed at build time from the 29 MIxS sample checklists SeqDesk mirrors from the ENA checklist registry (snapshot 2026-06-03). A field is a distinct field name; the “shared core” is the 6 fields used by ≥25 checklists. Iceberg areas are illustrative; the heatmap shows a domain-balanced core sample of 24 fields, and the sediment strip summarises all 880. Every printed number is exact.