SeqDesk · original analysis
Every public sample in the European Nucleotide Archive, scored for whether its contextual (MIxS) metadata is actually filled in — by submission year, re-counted weekly. Missing-value sentinels like “not collected” count as empty.
Hover the chart to read every shown field at that year · hover or click a legend item to isolate one line · use “Add more fields” to chart the extended metadata catalog (host, environment, sampling, taxonomy & standards).
Each line is the share of that year’s public samples whose field is validly filled — present and not an INSDC “missing value” placeholder. Geographic location and collection date have climbed steeply, but precise coordinates and a standardized MIxS environment stay rare, and host-associated context (isolation source, host organism, tissue) sits in between. Each year is the whole cohort that went public that year, so the curve reflects who submitted (a few very large projects can swing a recent year) — read the long arc, not single-year wobbles.
In 2025, a geographic location is recorded for 72% of samples — but 13.3% of all samples carry one of the controlled missing-value placeholders (missing, not collected, not provided, not applicable, restricted access) in that field, so the genuinely usable share is only 59%. Counting placeholders as “filled” would hide this gap — every field here subtracts them.
location field; collection dates via a valid date-range filterscripts/check-metadata-completeness.mjs| Year | Samples | Geographic location | Collection date | Lat / lon coordinates | MIxS environment | Isolation source | Host organism | Tissue type |
|---|---|---|---|---|---|---|---|---|
| 2026 * | 2.2M | 35% | 34% | 18% | 8% | 18% | 16% | 18% |
| 2025 | 5.0M | 59% | 56% | 26% | 12% | 28% | 27% | 26% |
| 2024 | 5.0M | 52% | 51% | 24% | 12% | 29% | 27% | 21% |
| 2023 | 6.6M | 60% | 58% | 20% | 13% | 26% | 30% | 15% |
| 2022 | 7.1M | 76% | 74% | 16% | 9% | 39% | 48% | 11% |
| 2021 | 6.8M | 75% | 71% | 13% | 8% | 34% | 41% | 10% |
| 2020 | 3.2M | 45% | 39% | 23% | 12% | 21% | 20% | 19% |
| 2019 | 2.6M | 39% | 31% | 26% | 14% | 23% | 17% | 19% |
| 2018 | 1.9M | 38% | 33% | 20% | 12% | 19% | 20% | 20% |
| 2017 | 1.6M | 32% | 28% | 17% | 11% | 16% | 16% | 15% |
| 2016 | 1.2M | 31% | 25% | 18% | 13% | 14% | 15% | 14% |
| 2015 | 931K | 23% | 20% | 12% | 10% | 11% | 12% | 12% |
| 2014 | 544K | 24% | 22% | 12% | 12% | 11% | 14% | 10% |
| 2013 | 434K | 12% | 10% | 4% | 4% | 9% | 8% | 3% |
| 2012 | 467K | 5% | 4% | 1% | 1% | 4% | 4% | 1% |
| 2011 | 238K | 4% | 4% | 1% | 1% | 28% | 3% | 1% |
| 2010 | 154K | 2% | 1% | 0% | 0% | 1% | 2% | 1% |
| 2009 | 32K | 0% | 0% | 2% | 0% | 0% | 1% | 4% |
| 2008 | 5.7K | 0% | 1% | 1% | 1% | 1% | 1% | 20% |
The quality counterpart to the archive counts: instead of how many samples exist, this measures how much of each sample’s contextual (MIxS) metadata is actually filled in — re-counted weekly straight from the EMBL-EBI / European Nucleotide Archive (Portal API /count, whole-archive). Only aggregate percentages and counts are published. · back to all research data
Use & cite this data
These are SeqDesk’s own aggregate figures, refreshed weekly. Download them, drop the live chart into a page, or pull the latest numbers from a small public API — then cite the snapshot you used.
date,value table — opens straight in Excel, R, or pandas.Download CSVAggregate facts compiled by SeqDesk from public archives (EMBL-EBI / European Nucleotide Archive) and re-served under each source's terms. Only whole-archive percentages and counts are published — no record-level identifiers. Provided as-is, no warranty. Full method & sources: https://seqdesk.org/data/metadata-completeness