SeqDesk · original analysis
A live cohort of genomic & biomedical databases from the literature, re-checked every week — survival by how the data is hosted.
A summary view: the share of the cohort still reachable over time. Reconstructed history — for every database now offline, the last date the Internet Archive holds a working snapshot estimates when it went dark; databases still live are counted alive throughout. The final point (2026-06-24) is the live measurement; earlier points are an archive-based estimate, so read the long-run slope as indicative rather than exact.
Estimated share still online that many years after publication — from each database’s birth (paper year), death (its last web-archive snapshot) and right-censoring (still live). Median half-life ≈ 24 years.
Share of databases still reachable, plotted against how old the data is — the cohort’s direct answer to “how long does research data stay online?”. Each point is a publication-era bucket from the latest check; older data is harder to keep online.
Each discipline ranked by the share of its tracked databases still online — the dark part of each bar is still reachable, the grey part is gone, so the split reads straight off the bar. The line beneath gives the database count and the typical (median) age of the data. Microbial & pathogen genomics has the highest reachable share in this cohort (94% online); antimicrobial resistance, taxonomy and sequence/structure resources the lowest (~67%) — differences to read with caution given the small per-field counts.
Not every dead resource is a 404 — the largest share simply stop responding (timeouts, DNS failures), and several still return a working 200 from a parked or placeholder page, which a naïve uptime check would miss.
scripts/check-link-rot.mjs · /api/link-rot| Breakdown | Group | n | Reachable |
|---|---|---|---|
| Hosting | University / lab server | 79 | 73% |
| Institutional archive | 36 | 89% | |
| DOI repository | 15 | 100% | |
| GitHub / GitLab | 9 | 100% | |
| Project site | 5 | 40% | |
| Other | 4 | 25% | |
| Data age | before 2008 | 17 | 65% |
| 2008–2012 | 18 | 61% | |
| 2013–2016 | 43 | 79% | |
| 2017–2020 | 46 | 85% | |
| 2021 onward | 24 | 92% | |
| Field | Microbial & pathogen genomics | 18 | 94% |
| Other / general-purpose | 15 | 87% | |
| Functional annotation & pathways | 13 | 85% | |
| Human & medical genomics | 53 | 79% | |
| Microbiome & metagenomics | 13 | 77% | |
| Taxonomy, phylogenetics & evolution | 18 | 67% | |
| Antimicrobial resistance | 9 | 67% | |
| Sequences, proteins & structures | 9 | 67% | |
| Resource type | Software + data | 29 | 93% |
| Reference dataset | 29 | 83% | |
| Reference database | 86 | 73% |
A fixed cohort of genomic & biomedical databases drawn from the literature, re-checked weekly. Survival is the fraction of the cohort still reachable, broken down by hosting class and by how old the data is. The tracked list of resources is kept private; only these aggregate figures are published. The weekly check runs via GitHub Actions. · back to all research data
Use & cite this data
These are SeqDesk’s own aggregate figures, refreshed weekly. Download them, drop the live chart into a page, or pull the latest numbers from a small public API — then cite the snapshot you used.
date,value table — opens straight in Excel, R, or pandas.Download CSVAggregate facts compiled by SeqDesk from public archives (SeqDesk; Internet Archive) and re-served under each source's terms. Only whole-archive percentages and counts are published — no record-level identifiers. Provided as-is, no warranty. Full method & sources: https://seqdesk.org/data/availability