SeqDesk · original analysis
Beyond raw record counts: how much of the world's open research data actually carries rich, reusable metadata — coordinates, verified identifications, images, posted results — across biodiversity, marine and clinical archives.
Share of GBIF records that carry coordinates, by collection year — each point is the share of records collected in that year that carry usable coordinates (a position, no geospatial issues). Specimens digitised from older field notes are sparsely georeferenced; modern, GPS- and smartphone-era records usually are. Read the long arc, not single-year wobbles — one very large dataset can nudge a recent year.
Six metadata-completeness signals beyond a bare record count: do biodiversity occurrences carry coordinates and images, are observations verified to research grade, are marine records resolved to species, and have registered clinical trials actually posted their results? Each is the share of that source’s records meeting the bar — high for coordinates, far lower for posted trial results.
year facetscripts/update-cross-domain-completeness.mjs| Collection year | Georeferenced |
|---|---|
| 2026 | 99.8% |
| 2025 | 99.7% |
| 2024 | 99.8% |
| 2023 | 99.8% |
| 2022 | 99.8% |
| 2021 | 99.4% |
| 2020 | 99.8% |
| 2019 | 99.6% |
| 2018 | 99.5% |
| 2017 | 99.5% |
| 2016 | 99.4% |
| 2015 | 99.0% |
| 2014 | 98.9% |
| 2013 | 98.6% |
| 2012 | 97.2% |
| 2011 | 98.5% |
| 2010 | 98.3% |
| 2009 | 97.9% |
| 2008 | 97.7% |
| 2007 | 97.3% |
| 2006 | 97.2% |
| 2005 | 97.0% |
| 2004 | 96.8% |
| 2003 | 96.4% |
| 2002 | 96.4% |
| 2001 | 96.1% |
| 2000 | 95.8% |
| 1999 | 95.4% |
| 1998 | 94.7% |
| 1997 | 94.2% |
| 1996 | 94.0% |
| 1995 | 93.6% |
| 1994 | 92.9% |
| 1993 | 93.0% |
| 1992 | 92.6% |
| 1991 | 92.5% |
| 1990 | 92.1% |
| 1989 | 91.8% |
| 1988 | 91.1% |
| 1987 | 90.8% |
| 1986 | 89.3% |
| 1985 | 88.8% |
| 1984 | 87.7% |
| 1983 | 87.9% |
| 1982 | 87.4% |
| 1981 | 86.5% |
| 1980 | 86.8% |
| 1979 | 84.6% |
| 1978 | 84.8% |
| 1977 | 83.6% |
| 1976 | 81.3% |
| 1975 | 81.2% |
| 1974 | 79.5% |
| 1973 | 77.7% |
| 1972 | 77.2% |
| 1971 | 74.7% |
| 1970 | 74.5% |
| 1969 | 70.3% |
| 1968 | 68.4% |
| 1967 | 66.5% |
| 1966 | 65.8% |
| 1965 | 66.3% |
| 1964 | 68.0% |
| 1963 | 69.0% |
| 1962 | 67.4% |
| 1961 | 66.2% |
| 1960 | 69.1% |
| 1959 | 73.2% |
| 1958 | 72.9% |
| 1957 | 73.0% |
| 1956 | 70.1% |
| 1955 | 71.0% |
| 1954 | 70.4% |
| 1953 | 67.7% |
| 1952 | 67.7% |
| 1951 | 67.6% |
| 1950 | 70.0% |
The quality counterpart to the archive counts: instead of how many records exist, this measures how much of each record’s context is actually filled in — re-counted weekly from public APIs. Only aggregate percentages and counts are published. · back to all research data
Use & cite this data
These are SeqDesk’s own aggregate figures, refreshed weekly. Download them, drop the live chart into a page, or pull the latest numbers from a small public API — then cite the snapshot you used.
date,value table — opens straight in Excel, R, or pandas.Download CSVAggregate facts compiled by SeqDesk from public archives (GBIF; iNaturalist, OBIS, ClinicalTrials.gov) and re-served under each source's terms. Only whole-archive percentages and counts are published — no record-level identifiers. Provided as-is, no warranty. Full method & sources: https://seqdesk.org/data/completeness