Who makes sequencing data FAIR?
Sequencing got cheap and the public archives filled up fast, but a raw read is not FAIR data. This is the story, in five short parts, of the gap between how much sequence the world produces and how much of it anyone can actually reuse, and of the effort to close it.
Sequencing got cheap, and the data piled up
The cost of sequencing a genome dropped from about $95 million to a few hundred dollars, so labs produced far more data than anyone could keep track of. Scroll to see the drop.
2001
NHGRI's $1000-genome curve, revisited · as of 2026-06-25 · curated · /data →
$95 million for one genome
In September 2001, the National Human Genome Research Institute put the cost of a single genome at about $95 million. One sequence was a national-scale project, and the public archives were nearly empty.
Moore's law was the optimistic benchmark
The hopeful case was that sequencing would keep pace with computing, halving in cost every two years, like Moore's law. The dashed line traces that trajectory from the same 2001 starting point.
In 2008 it broke away
With the jump from Sanger to next-generation sequencing, the cost per genome fell from roughly $3 million to about $230,000 in a single year, dropping below the Moore line. NHGRI dates this out-pacing of Moore's law to January 2008.
By 2022, about $525
The decline flattened after 2015, but the result was still large: $524.62 per genome by May 2022, about 141× lower than a Moore's-law trajectory from 2001.
This is why the data piled up
Cost per genome fell about 181,000-fold in twenty-one years. When sequencing gets that cheap, the public archives fill quickly, and much of that data now arrives faster than it can be described. That gap is part of why FAIR matters.
Most samples are missing the basics
Every public sample in the ENA, checked for whether simple facts, like where and when it was collected, are actually filled in. Scroll to see how many are blank.
scroll to reveal each field, or pick one below
0 are missing it
Five million samples
Every tile is a slice of the 4.95M public samples submitted to the European Nucleotide Archive in 2025, one of three INSDC archives that mirror the same data across three continents. Raw sequence is only reusable if you know what each sample is. How much of that context is actually filled in?
Where did it come from?
Only 59% record a geographic location. More than four in ten public samples never say which country, or sea, they were taken from.
When was it taken?
56% record a collection date. The one field that every one of the 29 MIxS checklists agrees you must capture (collection_date) is blank on nearly half of all samples.
Exactly where?
Just 26% carry latitude and longitude, and that only counts whether the field is present, not whether it is correct or lands on the right continent. Precise, mappable location is missing on three out of four.
What environment?
Only 12% describe the environment the sample came from. Coverage falls off steadily: more than half record roughly where, but fewer than one in eight record the habitat. There is a lot of sequence, and little context to go with it.
This is dark data
This is sometimes called dark data: deposited and public, but hard to find or reuse because the context that gives it meaning is missing. There is more sequence than ever, and a smaller share of it is findable, interoperable or reusable. Closing that gap, one sample at a time, is a central goal of FAIR. The next parts look at who is working on it.
How the world is responding
As the data piled up, groups around the world built standards, funding and rules to keep it usable. They appear on the map in the year each one started. Scroll the years forward.
Hover or tap a dot to see who it is.
2008 – 2012
The archives start to fill
Next-generation sequencing arrives and the public archives start to fill up fast. By 2012 the INSDC archives (ENA, NCBI/SRA and DDBJ) are already mirroring the same data across three continents. The problem is no longer storage; it is finding and reusing what is inside.
2013 – 2016
Standards and a shared rulebook
The response is coordination. GA4GH (2013) sets standards for sensitive human data; ELIXIR (2014) links Europe’s archives; and in 2016 the FAIR principles give everyone the same rulebook: Findable (a persistent ID you can search), Accessible (a documented way to retrieve it, which can still be behind controlled access, not the same as “open”), Interoperable (shared vocabularies), and Reusable (enough context to trust and reuse it), for machines, not just people.
2017 – 2020
Money and mandates
Now the funders move. Plan S and EOSC (2018), Australia’s BioCommons (2019), and Germany’s NFDI (2020, up to €90M a year) turn FAIR from a principle into budgeted national infrastructure, while the data keeps doubling.
2021 – 2023
FAIR gets enforced, and still slips
NFDI4Microbiota (2021) targets microbial sequence data; the Federated EGA nodes go live (2022). Then the mandates bite: the US OSTP “Nelson” memo (Aug 2022), the NIH data-sharing policy (Jan 2023), and the ENA making geographic location and collection date mandatory (May 2023). And yet, as the curve above showed, completeness peaked around 2022 and then fell. Mandates apply going forward; the archive is cumulative.
2024 – now
It comes down to the lab bench
The archives now hold millions of human genomes’ worth of data, and FAIR is funded and mandated worldwide, yet the completeness line still has not recovered. A mandate applies to the archive, but the country, the date and the checklist are decided at the bench, months earlier. That last step is where SeqDesk fits, turning a facility’s everyday orders into INSDC-ready, FAIR records.
Is any of it working?
Four simple measures of open and FAIR practice, showing what is improving and what still lags. Scroll through them one at a time.
Literature free to read
52%up from 9% in 2010Open-access papers linking a dataset
51%barely moved in 15 yearsPreprint volume growth in a decade
25×but only 6% link any dataShared databases still reachable
79%31 of 148 already goneIs it actually working?
FAIR is funded and mandated. Billions in grants, journal policies, national archives. But is the published record actually getting more findable and reusable? Four honest gauges.
Access is the real win
About 52% of the MEDLINE-indexed literature is now free to read, up from just 9% in 2010. On the “A” in FAIR, the literature has genuinely opened up.
Papers still rarely share data
Only 51% of even open-access papers link a single dataset, and that share has barely shifted in fifteen years. Free to read is not the same as free to reuse.
Preprints grew, their data did not
Preprint volume grew roughly 25× in a decade, to 186k a year. Yet preprints link a dataset only about 6% of the time, versus 53% for published papers. More records, but less provenance.
Shared data does not always stay online
Of a live cohort of 148 genomic and biomedical databases, only 79% are still reachable; 31 are already gone, with an estimated half-life near 24 years. Findability erodes over time.
Progress on access, not yet on the rest
Real, measurable progress on getting the literature open. Persistence and actual data-sharing still lag well behind. The data is being made FAIR, slowly and unevenly.
It starts at the lab bench
Every rule and archive depends on one thing: the details someone enters when a sample is first logged. Scroll to see where the data is really made.
How FAIR is decided, layer by layer
It comes down to data entry
Every mandate, every archive, every standard funnels down to one moment. FAIR data does not come from a policy document or an archive's submission portal. It comes from the details someone enters at the lab bench.
Global mandates & funders
It starts at the top. Funders and journals require that sequence data be deposited and openly available. The intent is sound, but a mandate only sets the destination, not how the context gets there.
Deposit in INSDC + cite an accession
So the reads land in an INSDC archive and a paper cites an accession. The bytes are public. Yet a public accession with no context is still a dead end, findable in name only.
MIxS-compliant, machine-readable metadata
What makes it reusable is MIxS-compliant, machine-readable metadata: the right checklist, the right fields, the right controlled vocabularies. This is the layer that turns a deposit into something another lab, or a machine, can actually find and trust.
Captured correctly at the sample bench
And all of that depends on one thing the archive can never recover for you: capturing the context correctly when the sample is taken. Yet only about 12% of public samples describe their environment, and only about 26% carry coordinates. The context is lost here, at the bench, before anything reaches an archive.
Captured at the source
This is the gap SeqDesk closes. SeqDesk maps each order to the right MIxS checklist, captures the country, date and coordinates the ENA now requires, and carries the accession through to submission, so FAIR is captured at the source, not back-filled.how SeqDesk submits to the ENASeqDesk
The data and the response
Two things grew over the same years. One is the data: a large and rising volume of sequence in the public archives. The other is the response to it, which has four parts. A funder pays for data infrastructure, an archive stores the sequences, a set of standards makes them interoperable, and a policy sets the expectation that data will be shared, within the limits of consent and controlled access. Germany has NFDI and NFDI4Microbiota; the EU has EOSC and Horizon Europe; the US has the SRA and the NIH and OSTP policies; Australia has BioCommons; Japan has DDBJ. The three large archives are not rivals: they hold the single INSDC dataset, mirrored across three continents.
Across funders, archives and journals, sharing sequencing data is now the default expectation, within the limits of privacy, consent and embargo. The open question is who makes it happen at the moment a sample is logged.
The full data behind this piece (the cost curve, metadata completeness, open access, link rot and more) is tracked live and updated weekly on the data pages.
| NHGRI's $1000-genome curve, revisited | as of 2026-06-25 | Curated — NHGRI cost series + live archive count | |
| Sequencing metadata completeness | as of 2026-06-29 | Weekly automated re-count | CSV · API |
| Open-access share | as of 2026-06-25 | Weekly automated re-count | CSV · API |
| Open-data practice | as of 2026-06-25 | Weekly automated re-count | CSV · API |
| Preprints and data sharing | as of 2026-06-29 | Weekly automated re-count | CSV · API |
| Research data availability | as of 2026-06-24 | Weekly automated check | CSV · API |
| Human-sequencing capacity (estimate) | as of 2026-06-25 | snapshot |
Sources & notes. The dark-data wall scores every public sample in the ENA (4,953,259 in 2025) for whether each MIxS field is actually filled in (missing-value sentinels such as “not collected” count as empty), from the ENA Portal API, re-counted weekly. Other figures are drawn from primary sources and time-stamped where they are snapshots: the FAIR principles (Wilkinson et al., Scientific Data, 2016; 22,000+ Google Scholar citations); ENA 62.8 PB and the May 2023 mandatory spatiotemporal fields (EMBL-EBI, 2024); NCBI SRA 25.6 petabase-pairs (NAR, 2021); INSDC >9 PB and ~10×/4-year growth (NAR, 2020); NFDI €90M/year and NFDI4Microbiota ~€3M/year over 2021–26 (DFG, HZI); EOSC ≥€1bn and Horizon Europe €95.5bn (European Commission); NIH Data Management & Sharing policy (Jan 2023) and the OSTP “Nelson” memo (Aug 2022); Australian BioCommons ~A$74M since 2019 (biocommons.org.au). The climbing metric is cumulative human-sequence data in the public read archive, in genome-equivalents (3.2 Gb each), counted per submission year from the ENA Portal API (which mirrors much of NCBI SRA); 2026 is a part-year. Country outlines: Natural Earth (110m), simplified. The NFDI4Microbiota annual figure is reported by partners rather than published by the DFG.