SeqDesk · original analysis
One chart for the FAIR argument: how many samples the ENA takes in each year, against the share of them that carry a valid MIxS environment field (which applies mainly to environmental samples).
Left axis is logarithmic — the volume line climbs across four orders of magnitude while the context line never leaves the floor · hover to read both.
The copper line is how many samples the ENA took in each year — from 5.7K in 2008 to 5.0M in 2025, a 866× rise (note the logarithmic axis). The orange line is the share of those samples that carry a valid MIxS environment, and it stays low, ending at just 12%. Sequence production rose by orders of magnitude over this period, while the share carrying this standardized contextual field — one signal among many that support findability and reuse — rose only modestly. A MIxS environment applies mainly to environmental samples, so read this as one field’s trend, not a verdict on every record. Each year is the whole cohort that went public that year — read the long arc, not single-year wobbles.
scripts/check-metadata-completeness.mjs| Year | Samples | With a MIxS environment |
|---|---|---|
| 2025 | 5.0M | 11.9% |
| 2024 | 5.0M | 11.7% |
| 2023 | 6.6M | 12.8% |
| 2022 | 7.1M | 9.2% |
| 2021 | 6.8M | 8.2% |
| 2020 | 3.2M | 11.7% |
| 2019 | 2.6M | 14.1% |
| 2018 | 1.9M | 11.9% |
| 2017 | 1.6M | 10.6% |
| 2016 | 1.2M | 13.1% |
| 2015 | 931K | 9.6% |
| 2014 | 544K | 11.8% |
| 2013 | 434K | 4.2% |
| 2012 | 467K | 1.4% |
| 2011 | 238K | 0.6% |
| 2010 | 154K | 0.3% |
| 2009 | 32K | 0.3% |
| 2008 | 5.7K | 0.9% |
The headline of the metadata story: produced volume against described context, re-counted weekly straight from the EMBL-EBI / European Nucleotide Archive (Portal API /count, whole-archive). Only aggregate percentages and counts are published. · back to all research data