Infrastructure · interactive

Who makes sequencing data FAIR?

Sequencing got cheap and the public archives filled up fast, but a raw read is not FAIR data. This is the story, in five short parts, of the gap between how much sequence the world produces and how much of it anyone can actually reuse, and of the effort to close it.

Part 1 · the cost of sequencing

Sequencing got cheap, and the data piled up

The cost of sequencing a genome dropped from about $95 million to a few hundred dollars, so labs produced far more data than anyone could keep track of. Scroll to see the drop.

cost per genome (USD)log scale
$100$1K$10K$100K$1M$10M$100M200220062010201420182022$95.3M
$95.3Mcost per genome
2001

NHGRI's $1000-genome curve, revisited · as of 2026-06-25 · curated · /data →

$95 million for one genome

In September 2001, the National Human Genome Research Institute put the cost of a single genome at about $95 million. One sequence was a national-scale project, and the public archives were nearly empty.

Moore's law was the optimistic benchmark

The hopeful case was that sequencing would keep pace with computing, halving in cost every two years, like Moore's law. The dashed line traces that trajectory from the same 2001 starting point.

In 2008 it broke away

With the jump from Sanger to next-generation sequencing, the cost per genome fell from roughly $3 million to about $230,000 in a single year, dropping below the Moore line. NHGRI dates this out-pacing of Moore's law to January 2008.

By 2022, about $525

The decline flattened after 2015, but the result was still large: $524.62 per genome by May 2022, about 141× lower than a Moore's-law trajectory from 2001.

This is why the data piled up

Cost per genome fell about 181,000-fold in twenty-one years. When sequencing gets that cheap, the public archives fill quickly, and much of that data now arrives faster than it can be described. That gap is part of why FAIR matters.

And it is not all human. Most of the sequence in the public archives comes from microbes, viruses, plants, animals and whole environments, such as a scoop of soil or a litre of seawater. That is why the rest of this story applies to all of it, not only to people. A human sample carries a lot of implied context; an environmental sample is hard to reuse unless someone records where and what it came from. The “human genomes’ worth” figure used later is just a familiar unit for the volume.
microbesvirusesplants & animalsenvironmentalhuman
Part 2 · missing metadata

Most samples are missing the basics

Every public sample in the ENA, checked for whether simple facts, like where and when it was collected, are actually filled in. Scroll to see how many are blank.

share of ENA samples with this filled inall samples

scroll to reveal each field, or pick one below

filled in missingeach tile ≈ 12.9K samples
100%4.95M of 4.95M samples
0 are missing it
Core MIxS context
Host & sample
Environment (MIxS)
Geography
Sampling provenance
Taxonomy detail
Standards

Five million samples

Every tile is a slice of the 4.95M public samples submitted to the European Nucleotide Archive in 2025, one of three INSDC archives that mirror the same data across three continents. Raw sequence is only reusable if you know what each sample is. How much of that context is actually filled in?

Where did it come from?

Only 59% record a geographic location. More than four in ten public samples never say which country, or sea, they were taken from.

When was it taken?

56% record a collection date. The one field that every one of the 29 MIxS checklists agrees you must capture (collection_date) is blank on nearly half of all samples.

Exactly where?

Just 26% carry latitude and longitude, and that only counts whether the field is present, not whether it is correct or lands on the right continent. Precise, mappable location is missing on three out of four.

What environment?

Only 12% describe the environment the sample came from. Coverage falls off steadily: more than half record roughly where, but fewer than one in eight record the habitat. There is a lot of sequence, and little context to go with it.

This is dark data

This is sometimes called dark data: deposited and public, but hard to find or reuse because the context that gives it meaning is missing. There is more sequence than ever, and a smaller share of it is findable, interoperable or reusable. Closing that gap, one sample at a time, is a central goal of FAIR. The next parts look at who is working on it.

Part 3 · what’s being done

How the world is responding

As the data piled up, groups around the world built standards, funding and rules to keep it usable. They appear on the map in the year each one started. Scroll the years forward.

scroll to advance the year · hover a dot2008
ENA · EMBL-EBINCBI · SRADDBJ
INSDC archive standard · funder · mandate

Hover or tap a dot to see who it is.

0
human genomes’ worth of sequence in the public read archive (cumulative, INSDC; most under controlled access)
02008201420202026

2008 – 2012

The archives start to fill

Next-generation sequencing arrives and the public archives start to fill up fast. By 2012 the INSDC archives (ENA, NCBI/SRA and DDBJ) are already mirroring the same data across three continents. The problem is no longer storage; it is finding and reusing what is inside.

2013 – 2016

Standards and a shared rulebook

The response is coordination. GA4GH (2013) sets standards for sensitive human data; ELIXIR (2014) links Europe’s archives; and in 2016 the FAIR principles give everyone the same rulebook: Findable (a persistent ID you can search), Accessible (a documented way to retrieve it, which can still be behind controlled access, not the same as “open”), Interoperable (shared vocabularies), and Reusable (enough context to trust and reuse it), for machines, not just people.

2017 – 2020

Money and mandates

Now the funders move. Plan S and EOSC (2018), Australia’s BioCommons (2019), and Germany’s NFDI (2020, up to €90M a year) turn FAIR from a principle into budgeted national infrastructure, while the data keeps doubling.

2021 – 2023

FAIR gets enforced, and still slips

NFDI4Microbiota (2021) targets microbial sequence data; the Federated EGA nodes go live (2022). Then the mandates bite: the US OSTP “Nelson” memo (Aug 2022), the NIH data-sharing policy (Jan 2023), and the ENA making geographic location and collection date mandatory (May 2023). And yet, as the curve above showed, completeness peaked around 2022 and then fell. Mandates apply going forward; the archive is cumulative.

2024 – now

It comes down to the lab bench

The archives now hold millions of human genomes’ worth of data, and FAIR is funded and mandated worldwide, yet the completeness line still has not recovered. A mandate applies to the archive, but the country, the date and the checklist are decided at the bench, months earlier. That last step is where SeqDesk fits, turning a facility’s everyday orders into INSDC-ready, FAIR records.

Part 4 · is it working?

Is any of it working?

Four simple measures of open and FAIR practice, showing what is improving and what still lags. Scroll through them one at a time.

four honest gauges of FAIR progressscorecard

Literature free to read

52%up from 9% in 2010

Open-access papers linking a dataset

51%barely moved in 15 years

Preprint volume growth in a decade

25×but only 6% link any data

Shared databases still reachable

79%31 of 148 already gone
Four FAIR gauges1 access, 3 lagging

Is it actually working?

FAIR is funded and mandated. Billions in grants, journal policies, national archives. But is the published record actually getting more findable and reusable? Four honest gauges.

Access is the real win

About 52% of the MEDLINE-indexed literature is now free to read, up from just 9% in 2010. On the “A” in FAIR, the literature has genuinely opened up.

Papers still rarely share data

Only 51% of even open-access papers link a single dataset, and that share has barely shifted in fifteen years. Free to read is not the same as free to reuse.

Preprints grew, their data did not

Preprint volume grew roughly 25× in a decade, to 186k a year. Yet preprints link a dataset only about 6% of the time, versus 53% for published papers. More records, but less provenance.

Shared data does not always stay online

Of a live cohort of 148 genomic and biomedical databases, only 79% are still reachable; 31 are already gone, with an estimated half-life near 24 years. Findability erodes over time.

Progress on access, not yet on the rest

Real, measurable progress on getting the literature open. Persistence and actual data-sharing still lag well behind. The data is being made FAIR, slowly and unevenly.

Part 5 · at the lab bench

It starts at the lab bench

Every rule and archive depends on one thing: the details someone enters when a sample is first logged. Scroll to see where the data is really made.

How FAIR is decided, layer by layer

It comes down to data entry

Every mandate, every archive, every standard funnels down to one moment. FAIR data does not come from a policy document or an archive's submission portal. It comes from the details someone enters at the lab bench.

Global mandates & funders

It starts at the top. Funders and journals require that sequence data be deposited and openly available. The intent is sound, but a mandate only sets the destination, not how the context gets there.

Deposit in INSDC + cite an accession

So the reads land in an INSDC archive and a paper cites an accession. The bytes are public. Yet a public accession with no context is still a dead end, findable in name only.

MIxS-compliant, machine-readable metadata

What makes it reusable is MIxS-compliant, machine-readable metadata: the right checklist, the right fields, the right controlled vocabularies. This is the layer that turns a deposit into something another lab, or a machine, can actually find and trust.

Captured correctly at the sample bench

And all of that depends on one thing the archive can never recover for you: capturing the context correctly when the sample is taken. Yet only about 12% of public samples describe their environment, and only about 26% carry coordinates. The context is lost here, at the bench, before anything reaches an archive.

Captured at the source

This is the gap SeqDesk closes. SeqDesk maps each order to the right MIxS checklist, captures the country, date and coordinates the ENA now requires, and carries the accession through to submission, so FAIR is captured at the source, not back-filled.how SeqDesk submits to the ENASeqDesk

The data and the response

Two things grew over the same years. One is the data: a large and rising volume of sequence in the public archives. The other is the response to it, which has four parts. A funder pays for data infrastructure, an archive stores the sequences, a set of standards makes them interoperable, and a policy sets the expectation that data will be shared, within the limits of consent and controlled access. Germany has NFDI and NFDI4Microbiota; the EU has EOSC and Horizon Europe; the US has the SRA and the NIH and OSTP policies; Australia has BioCommons; Japan has DDBJ. The three large archives are not rivals: they hold the single INSDC dataset, mirrored across three continents.

Across funders, archives and journals, sharing sequencing data is now the default expectation, within the limits of privacy, consent and embargo. The open question is who makes it happen at the moment a sample is logged.

The full data behind this piece (the cost curve, metadata completeness, open access, link rot and more) is tracked live and updated weekly on the data pages.

The data behind this story
NHGRI's $1000-genome curve, revisitedas of 2026-06-25Curated — NHGRI cost series + live archive count
Sequencing metadata completenessas of 2026-06-29Weekly automated re-count
Open-access shareas of 2026-06-25Weekly automated re-count
Open-data practiceas of 2026-06-25Weekly automated re-count
Preprints and data sharingas of 2026-06-29Weekly automated re-count
Research data availabilityas of 2026-06-24Weekly automated check
Human-sequencing capacity (estimate)as of 2026-06-25snapshot

Sources & notes. The dark-data wall scores every public sample in the ENA (4,953,259 in 2025) for whether each MIxS field is actually filled in (missing-value sentinels such as “not collected” count as empty), from the ENA Portal API, re-counted weekly. Other figures are drawn from primary sources and time-stamped where they are snapshots: the FAIR principles (Wilkinson et al., Scientific Data, 2016; 22,000+ Google Scholar citations); ENA 62.8 PB and the May 2023 mandatory spatiotemporal fields (EMBL-EBI, 2024); NCBI SRA 25.6 petabase-pairs (NAR, 2021); INSDC >9 PB and ~10×/4-year growth (NAR, 2020); NFDI €90M/year and NFDI4Microbiota ~€3M/year over 2021–26 (DFG, HZI); EOSC ≥€1bn and Horizon Europe €95.5bn (European Commission); NIH Data Management & Sharing policy (Jan 2023) and the OSTP “Nelson” memo (Aug 2022); Australian BioCommons ~A$74M since 2019 (biocommons.org.au). The climbing metric is cumulative human-sequence data in the public read archive, in genome-equivalents (3.2 Gb each), counted per submission year from the ENA Portal API (which mirrors much of NCBI SRA); 2026 is a part-year. Country outlines: Natural Earth (110m), simplified. The NFDI4Microbiota annual figure is reported by partners rather than published by the DFG.