SeqDesk · original analysis
Every public sample in the European Nucleotide Archive, by submission year, scored for whether it declares a GSC/NCBI reporting standard at all — and whether that standard is a real MIxS package or the permissive “Generic” default that mandates no contextual fields. Re-counted weekly.
Worth a closer look. Nearly every sample now declares a standard — but 1.6M of them (33%) declare “Generic”, the package that requires no contextual fields. Generic is a valid choice when no MIxS package fits, but a high “declares a standard” rate doesn’t by itself mean rich metadata. Where a real package does fit, SeqDesk defaults studies to it rather than to Generic.
Each year’s bands stack to 100% of that year’s public samples · hover to read the split.
Every public sample falls into exactly one band, so each year stacks to 100%: on a real GSC MIxS package (the dark band), on the permissive “Generic” default (mid-grey), or declaring no standard at all (the lightest band). The dark band is the share whose checklist asked for geography, environment, and host fields; the mid-grey band is on Generic, which asks for none — a valid choice where no package fits, but not a guarantee of contextual metadata. Each year is the whole cohort that went public that year, so a few very large projects can swing a recent year — read the long arc.
Share of the 2025 cohort declaring each named package (bars scaled to the largest). These are the submissions held to a standard that actually asks for contextual metadata — the rest fall back to “Generic” or declare nothing at all.
ncbi_reporting_standard other than the permissive “Generic” defaultscripts/check-reporting-standard.mjs| Year | Samples | Declares any standard | On a real MIxS package | “Generic” default |
|---|---|---|---|---|
| 2026 * | 2.2M | 100% | 39% | 61% |
| 2025 | 5.0M | 99% | 66% | 33% |
| 2024 | 5.0M | 97% | 57% | 40% |
| 2023 | 6.6M | 70% | 43% | 27% |
| 2022 | 7.1M | 63% | 51% | 12% |
| 2021 | 6.8M | 61% | 46% | 15% |
| 2020 | 3.2M | 80% | 46% | 34% |
| 2019 | 2.6M | 75% | 42% | 34% |
| 2018 | 1.9M | 79% | 42% | 37% |
| 2017 | 1.6M | 85% | 33% | 52% |
| 2016 | 1.2M | 72% | 31% | 42% |
| 2015 | 931K | 81% | 25% | 56% |
| 2014 | 544K | 78% | 23% | 55% |
| 2013 | 434K | 84% | 4% | 80% |
| 2012 | 467K | 91% | 1% | 90% |
| 2011 | 238K | 93% | 0% | 93% |
| 2010 | 154K | 95% | 0% | 95% |
| 2009 | 32K | 100% | 2% | 98% |
| 2008 | 5.7K | 100% | 6% | 95% |
The prior question to metadata completeness: was the sample ever held to a standard that would demand contextual metadata? Re-counted weekly straight from the EMBL-EBI / European Nucleotide Archive (Portal API /count, whole-archive). Only aggregate percentages and counts are published. · back to all research data
Use & cite this data
These are SeqDesk’s own aggregate figures, refreshed weekly. Download them, drop the live chart into a page, or pull the latest numbers from a small public API — then cite the snapshot you used.
date,value table — opens straight in Excel, R, or pandas.Download CSVAggregate facts compiled by SeqDesk from public archives (EMBL-EBI / European Nucleotide Archive) and re-served under each source's terms. Only whole-archive percentages and counts are published — no record-level identifiers. Provided as-is, no warranty. Full method & sources: https://seqdesk.org/data/reporting-standard