← All research data

SeqDesk · original analysis

Are sequencing samples held to a real reporting standard?

Every public sample in the European Nucleotide Archive, by submission year, scored for whether it declares a GSC/NCBI reporting standard at all — and whether that standard is a real MIxS package or the permissive “Generic” default that mandates no contextual fields. Re-counted weekly.

5.0M
samples · 2025
66%
on a real MIxS package
99%
declare any standard
33%
of those pick “Generic”

Worth a closer look. Nearly every sample now declares a standard — but 1.6M of them (33%) declare “Generic”, the package that requires no contextual fields. Generic is a valid choice when no MIxS package fits, but a high “declares a standard” rate doesn’t by itself mean rich metadata. Where a real package does fit, SeqDesk defaults studies to it rather than to Generic.

0%25%50%75%100%2008201020122014201620182020202220242025sample submission year (first public)1234
On a real MIxS package“Generic” defaultNo standard declared

Each year’s bands stack to 100% of that year’s public samples · hover to read the split.

Every public sample falls into exactly one band, so each year stacks to 100%: on a real GSC MIxS package (the dark band), on the permissive “Generic” default (mid-grey), or declaring no standard at all (the lightest band). The dark band is the share whose checklist asked for geography, environment, and host fields; the mid-grey band is on Generic, which asks for none — a valid choice where no package fits, but not a guarantee of contextual metadata. Each year is the whole cohort that went public that year, so a few very large projects can swing a recent year — read the long arc.

Timeline — the standards behind the curve
  • 12008-05MIGS standard — The Genomic Standards Consortium's “Minimum Information about a Genome Sequence” (Field et al., Nat Biotechnol) — the first community checklist a submission could declare.
  • 22011-05MIxS specification — MIMARKS & MIxS (Yilmaz et al., Nat Biotechnol) define the package family — survey, specimen, the MIMS environmental package — a sample can be reported against.
  • 32016-03-15FAIR Principles — Wilkinson et al. (Sci Data) put rich, standardized, machine-readable metadata at the centre of data stewardship.
  • 42020-03-11COVID-19 sequencing surge — SARS-CoV-2 submissions exploded; pathogen/clinical packages now dominate recent cohorts, so post-2020 swings reflect what was submitted as much as policy.
Generic mandates no contextual fields
33%
Human
7%
Plant
6%
MIMARKS.survey
5%
Pathogen.cl
4%
MIMS.me
3%
Invertebrate
3%
Pathogen.env
1%
MIGS.eu
1%

Share of the 2025 cohort declaring each named package (bars scaled to the largest). These are the submissions held to a standard that actually asks for contextual metadata — the rest fall back to “Generic” or declare nothing at all.

CohortEvery public sample in the ENA, grouped by the year it was first made public
Update cadenceWeekly automated re-count · latest snapshot 2026-06-25
“Real package” meansDeclares a ncbi_reporting_standard other than the permissive “Generic” default
MeasurementWhole-archive counts from the EMBL-EBI / European Nucleotide Archive (Portal API /count, whole-archive) — exact, not a sample
CompanionPairs with sequencing metadata completeness — the fields a standard would demand
Methodscripts/check-reporting-standard.mjs
YearSamplesDeclares any standardOn a real MIxS package“Generic” default
2026 *2.2M100%39%61%
20255.0M99%66%33%
20245.0M97%57%40%
20236.6M70%43%27%
20227.1M63%51%12%
20216.8M61%46%15%
20203.2M80%46%34%
20192.6M75%42%34%
20181.9M79%42%37%
20171.6M85%33%52%
20161.2M72%31%42%
2015931K81%25%56%
2014544K78%23%55%
2013434K84%4%80%
2012467K91%1%90%
2011238K93%0%93%
2010154K95%0%95%
200932K100%2%98%
20085.7K100%6%95%

The prior question to metadata completeness: was the sample ever held to a standard that would demand contextual metadata? Re-counted weekly straight from the EMBL-EBI / European Nucleotide Archive (Portal API /count, whole-archive). Only aggregate percentages and counts are published. · back to all research data

Report errorpmu15@helmholtz-hzi.de

Use & cite this data

Take this analysis into your own work

These are SeqDesk’s own aggregate figures, refreshed weekly. Download them, drop the live chart into a page, or pull the latest numbers from a small public API — then cite the snapshot you used.

Full dataset · JSONThe complete snapshot behind this page — every series and breakdown, with the source and license inline.Download JSON
Headline series · CSVThe trend shown in the chart as a tidy date,value table — opens straight in Excel, R, or pandas.Download CSV
Chart · SVGThe figure as a crisp vector — scales to any size for slides, posters, or print.
Chart · PNGA high-resolution raster of the figure for quick drops into docs and decks.
License

Aggregate facts compiled by SeqDesk from public archives (EMBL-EBI / European Nucleotide Archive) and re-served under each source's terms. Only whole-archive percentages and counts are published — no record-level identifiers. Provided as-is, no warranty. Full method & sources: https://seqdesk.org/data/reporting-standard