← All research data

SeqDesk · original analysis

Surging data volume vs lagging contextual metadata

One chart for the FAIR argument: how many samples the ENA takes in each year, against the share of them that carry a valid MIxS environment field (which applies mainly to environmental samples).

5.0M
samples · 2025
866×
more samples than 2008
12%
carry a MIxS environment
0.9%
did in 2008
1.0K10K100K1.0M10M0%25%50%75%100%2008201020122014201620182020202220242025sample submission year (first public)12345
Samples deposited / year (log)Share with valid MIxS environment context (right axis)

Left axis is logarithmic — the volume line climbs across four orders of magnitude while the context line never leaves the floor · hover to read both.

The copper line is how many samples the ENA took in each year — from 5.7K in 2008 to 5.0M in 2025, a 866× rise (note the logarithmic axis). The orange line is the share of those samples that carry a valid MIxS environment, and it stays low, ending at just 12%. Sequence production rose by orders of magnitude over this period, while the share carrying this standardized contextual field — one signal among many that support findability and reuse — rose only modestly. A MIxS environment applies mainly to environmental samples, so read this as one field’s trend, not a verdict on every record. Each year is the whole cohort that went public that year — read the long arc, not single-year wobbles.

Timeline — the standards behind the curve
  • 12008-05MIGS standard — The Genomic Standards Consortium's “Minimum Information about a Genome Sequence” (Field et al., Nat Biotechnol) — the first community checklist for genome & metagenome context.
  • 22011-05MIxS specification — MIMARKS & MIxS (Yilmaz et al., Nat Biotechnol) define the minimum contextual fields — geography, environment, collection — this analysis scores against.
  • 32016-03-15FAIR Principles — Wilkinson et al. (Sci Data) coin the FAIR principles, putting rich, machine-readable metadata at the centre of data stewardship.
  • 42020-03-11COVID-19 sequencing surge — The WHO declared a pandemic and SARS-CoV-2 submissions exploded; these viral-genome batches dominate recent cohorts, so post-2020 swings reflect what was submitted as much as practice.
  • 52023-03-03INSDC spatiotemporal standards — INSDC tightened collection-date & lat/lon reporting and standardised the missing-value vocabulary — the very placeholders the metadata tracker counts as empty.
CohortEvery public sample in the ENA, grouped by the year it was first made public (partial current year excluded)
VolumeWhole-archive sample counts per first-public year — exact, not a sample
“MIxS environment”A valid, non-placeholder environment context (broad-scale / local / medium), the rarest contextual field
Update cadenceWeekly automated re-count · latest snapshot 2026-06-29
CompanionPairs with sequencing metadata completeness (the same data, field by field) and reporting-standard adoption
Methodscripts/check-metadata-completeness.mjs
YearSamplesWith a MIxS environment
20255.0M11.9%
20245.0M11.7%
20236.6M12.8%
20227.1M9.2%
20216.8M8.2%
20203.2M11.7%
20192.6M14.1%
20181.9M11.9%
20171.6M10.6%
20161.2M13.1%
2015931K9.6%
2014544K11.8%
2013434K4.2%
2012467K1.4%
2011238K0.6%
2010154K0.3%
200932K0.3%
20085.7K0.9%

The headline of the metadata story: produced volume against described context, re-counted weekly straight from the EMBL-EBI / European Nucleotide Archive (Portal API /count, whole-archive). Only aggregate percentages and counts are published. · back to all research data

Report errorpmu15@helmholtz-hzi.de