← All research data

SeqDesk · original analysis

How full is the world's sequencing metadata?

Every public sample in the European Nucleotide Archive, scored for whether its contextual (MIxS) metadata is actually filled in — by submission year, re-counted weekly. Missing-value sentinels like “not collected” count as empty.

5.0M
samples · 2025
38%
mean MIxS context filled
26%
have lat/lon coordinates
12%
have a MIxS environment
0%25%50%75%100%2008201020122014201620182020202220242025sample submission year (first public)12345

Hover the chart to read every shown field at that year · hover or click a legend item to isolate one line · use “Add more fields” to chart the extended metadata catalog (host, environment, sampling, taxonomy & standards).

Each line is the share of that year’s public samples whose field is validly filled — present and not an INSDC “missing value” placeholder. Geographic location and collection date have climbed steeply, but precise coordinates and a standardized MIxS environment stay rare, and host-associated context (isolation source, host organism, tissue) sits in between. Each year is the whole cohort that went public that year, so the curve reflects who submitted (a few very large projects can swing a recent year) — read the long arc, not single-year wobbles.

Timeline — the standards behind the curve
  • 12008-05MIGS standard — The Genomic Standards Consortium's “Minimum Information about a Genome Sequence” (Field et al., Nat Biotechnol) — the first community checklist for genome & metagenome context.
  • 22011-05MIxS specification — MIMARKS & MIxS (Yilmaz et al., Nat Biotechnol) define the minimum contextual fields — geography, environment, collection — that this analysis scores against.
  • 32016-03-15FAIR Principles — Wilkinson et al. (Sci Data) coin the FAIR principles, putting rich, machine-readable metadata at the centre of data stewardship.
  • 42020-03-11COVID-19 sequencing surge — The WHO declared a pandemic and SARS-CoV-2 submissions exploded; these viral-genome batches now dominate recent cohorts, so post-2020 swings reflect what was submitted as much as facility practice.
  • 52023-03-03INSDC spatiotemporal standards — INSDC tightened collection-date & lat/lon reporting and standardised the missing-value vocabulary — the very placeholders this tracker counts as empty.
Location field present any value at all
72% filled
Location actually usable excludes “missing”, “not collected”, …
59% filled

In 2025, a geographic location is recorded for 72% of samples — but 13.3% of all samples carry one of the controlled missing-value placeholders (missing, not collected, not provided, not applicable, restricted access) in that field, so the genuinely usable share is only 59%. Counting placeholders as “filled” would hide this gap — every field here subtracts them.

Geographic location
59% filled
Collection date
56% filled
Lat / lon coordinates
26% filled
MIxS environment
12% filled
Isolation source
28% filled
Host organism
27% filled
Tissue type
26% filled
CohortEvery public sample in the ENA, grouped by the year it was first made public
Update cadenceWeekly automated re-count · latest snapshot 2026-06-29
“Filled” meansA real value — not blank and not an INSDC missing-value placeholder (missing, not collected, not provided, not applicable, restricted access)
MeasurementWhole-archive counts from the EMBL-EBI / European Nucleotide Archive (Portal API /count, whole-archive) — exact, not a sample
Coordinates / datesLat/lon counted via the location field; collection dates via a valid date-range filter
Methodscripts/check-metadata-completeness.mjs
YearSamplesGeographic locationCollection dateLat / lon coordinatesMIxS environmentIsolation sourceHost organismTissue type
2026 *2.2M35%34%18%8%18%16%18%
20255.0M59%56%26%12%28%27%26%
20245.0M52%51%24%12%29%27%21%
20236.6M60%58%20%13%26%30%15%
20227.1M76%74%16%9%39%48%11%
20216.8M75%71%13%8%34%41%10%
20203.2M45%39%23%12%21%20%19%
20192.6M39%31%26%14%23%17%19%
20181.9M38%33%20%12%19%20%20%
20171.6M32%28%17%11%16%16%15%
20161.2M31%25%18%13%14%15%14%
2015931K23%20%12%10%11%12%12%
2014544K24%22%12%12%11%14%10%
2013434K12%10%4%4%9%8%3%
2012467K5%4%1%1%4%4%1%
2011238K4%4%1%1%28%3%1%
2010154K2%1%0%0%1%2%1%
200932K0%0%2%0%0%1%4%
20085.7K0%1%1%1%1%1%20%

The quality counterpart to the archive counts: instead of how many samples exist, this measures how much of each sample’s contextual (MIxS) metadata is actually filled in — re-counted weekly straight from the EMBL-EBI / European Nucleotide Archive (Portal API /count, whole-archive). Only aggregate percentages and counts are published. · back to all research data

Report errorpmu15@helmholtz-hzi.de

Use & cite this data

Take this analysis into your own work

These are SeqDesk’s own aggregate figures, refreshed weekly. Download them, drop the live chart into a page, or pull the latest numbers from a small public API — then cite the snapshot you used.

Full dataset · JSONThe complete snapshot behind this page — every series and breakdown, with the source and license inline.Download JSON
Headline series · CSVThe trend shown in the chart as a tidy date,value table — opens straight in Excel, R, or pandas.Download CSV
Chart · SVGThe figure as a crisp vector — scales to any size for slides, posters, or print.
Chart · PNGA high-resolution raster of the figure for quick drops into docs and decks.
License

Aggregate facts compiled by SeqDesk from public archives (EMBL-EBI / European Nucleotide Archive) and re-served under each source's terms. Only whole-archive percentages and counts are published — no record-level identifiers. Provided as-is, no warranty. Full method & sources: https://seqdesk.org/data/metadata-completeness