← All research data

SeqDesk · original analysis

How long does research data stay online?

A live cohort of genomic & biomedical databases from the literature, re-checked every week — survival by how the data is hosted.

148
databases tracked
117
reachable now
31
offline now
79%
still online · 2026-06-24
79%86%93%100%200920122015201820212024

A summary view: the share of the cohort still reachable over time. Reconstructed history — for every database now offline, the last date the Internet Archive holds a working snapshot estimates when it went dark; databases still live are counted alive throughout. The final point (2026-06-24) is the live measurement; earlier points are an archive-based estimate, so read the long-run slope as indicative rather than exact.

Survival odds Kaplan-Meier · n=147
95%reachable at 5y
84%reachable at 10y
74%reachable at 15y
62%reachable at 20y

Estimated share still online that many years after publication — from each database’s birth (paper year), death (its last web-archive snapshot) and right-censoring (still live). Median half-life ≈ 24 years.

61%71%82%92%3y6y9y12y15y18y

Share of databases still reachable, plotted against how old the data is — the cohort’s direct answer to “how long does research data stay online?”. Each point is a publication-era bucket from the latest check; older data is harder to keep online.

still online offlineranked by share still reachable
Microbial & pathogen genomics94% online
18 databases · median age ~10 yr · 1 offline
Other / general-purpose87% online
15 databases · median age ~12 yr · 2 offline
Functional annotation & pathways85% online
13 databases · median age ~8 yr · 2 offline
Human & medical genomics79% online
53 databases · median age ~8 yr · 11 offline
Microbiome & metagenomics77% online
13 databases · median age ~9 yr · 3 offline
Taxonomy, phylogenetics & evolution67% online
18 databases · median age ~11.5 yr · 6 offline
Antimicrobial resistance67% online
9 databases · median age ~12 yr · 3 offline
Sequences, proteins & structures67% online
9 databases · median age ~12 yr · 3 offline

Each discipline ranked by the share of its tracked databases still online — the dark part of each bar is still reachable, the grey part is gone, so the split reads straight off the bar. The line beneath gives the database count and the typical (median) age of the data. Microbial & pathogen genomics has the highest reachable share in this cohort (94% online); antimicrobial resistance, taxonomy and sequence/structure resources the lowest (~67%) — differences to read with caution given the small per-field counts.

University / lab server n=79
73% live · 21 down
Institutional archive n=36
89% live · 4 down
DOI repository n=15
100% live
GitHub / GitLab n=9
100% live
Project site n=5
40% live · 3 down
Other n=4
25% live · 3 down
before 2008 n=17
65% live · 6 down
2008–2012 n=18
61% live · 7 down
2013–2016 n=43
79% live · 9 down
2017–2020 n=46
85% live · 7 down
2021 onward n=24
92% live · 2 down
Software + data n=29
93% live · 2 down
Reference dataset n=29
83% live · 5 down
Reference database n=86
73% live · 23 down
Parked / placeholder (200)7
Timed out6
Host unreachable (DNS / refused)5
Not found (404 / 410)4
Login or forbidden (401 / 403)3
Server error (5xx)2
TLS / certificate error2
Redirect to nowhere (3xx)2

Not every dead resource is a 404 — the largest share simply stop responding (timeouts, DNS failures), and several still return a working 200 from a parked or placeholder page, which a naïve uptime check would miss.

Cohort148 genomic & biomedical databases drawn from the literature
Update cadenceWeekly automated check · tracking since 2026-06-24
“Reachable” meansA live 2xx response serving the resource — not a parked domain, error page, or login wall
HistoryPre-tracking points reconstructed from Internet Archive (Wayback Machine) snapshots
Web-archive backup95% have a Wayback snapshot · 7 live but unarchived (no safety net)
PrivacyResource names, URLs & source papers are kept private; only the aggregate figures here are published
Method & APIscripts/check-link-rot.mjs · /api/link-rot
BreakdownGroupnReachable
HostingUniversity / lab server7973%
Institutional archive3689%
DOI repository15100%
GitHub / GitLab9100%
Project site540%
Other425%
Data agebefore 20081765%
2008–20121861%
2013–20164379%
2017–20204685%
2021 onward2492%
FieldMicrobial & pathogen genomics1894%
Other / general-purpose1587%
Functional annotation & pathways1385%
Human & medical genomics5379%
Microbiome & metagenomics1377%
Taxonomy, phylogenetics & evolution1867%
Antimicrobial resistance967%
Sequences, proteins & structures967%
Resource typeSoftware + data2993%
Reference dataset2983%
Reference database8673%

A fixed cohort of genomic & biomedical databases drawn from the literature, re-checked weekly. Survival is the fraction of the cohort still reachable, broken down by hosting class and by how old the data is. The tracked list of resources is kept private; only these aggregate figures are published. The weekly check runs via GitHub Actions. · back to all research data

Report errorpmu15@helmholtz-hzi.de

Use & cite this data

Take this analysis into your own work

These are SeqDesk’s own aggregate figures, refreshed weekly. Download them, drop the live chart into a page, or pull the latest numbers from a small public API — then cite the snapshot you used.

Full dataset · JSONThe complete snapshot behind this page — every series and breakdown, with the source and license inline.Download JSON
Headline series · CSVThe trend shown in the chart as a tidy date,value table — opens straight in Excel, R, or pandas.Download CSV
Chart · SVGThe figure as a crisp vector — scales to any size for slides, posters, or print.
Chart · PNGA high-resolution raster of the figure for quick drops into docs and decks.
License

Aggregate facts compiled by SeqDesk from public archives (SeqDesk; Internet Archive) and re-served under each source's terms. Only whole-archive percentages and counts are published — no record-level identifiers. Provided as-is, no warranty. Full method & sources: https://seqdesk.org/data/availability