There are 214 occurrence records for a particular ground beetle in one midwestern county in the aggregated database most people use, and 189 of them share the same latitude and longitude to six decimal places. That coordinate is the county courthouse.
Nobody collected 189 beetles at a courthouse. What happened is that the original specimen labels said the name of the county and nothing else — which is what a label said in 1931 — and at some point in the digitisation chain a georeferencing step converted the county name into the county's administrative centroid, or in this case a point somebody picked to stand for the county, and wrote it into a coordinate field with six decimal places.
Six decimal places is a precision of about eleven centimetres.
I want to walk through what happens to that record downstream, because the interesting part is not the georeferencing, which is a known and documented practice that the aggregators are entirely open about. The interesting part is what the eleven centimetres does to everybody who touches the record afterwards.
The aggregator publishes a coordinate uncertainty field alongside the point. For our beetle it is populated, and it says something like 35000 metres, which is honest and correct and tells you exactly what you need to know. So far nothing is wrong.
Then somebody builds a species distribution model. A distribution model takes occurrence points, extracts the environmental conditions at each point — temperature, precipitation, land cover, elevation — and learns which conditions the species is associated with. And the extraction step reads a raster at a coordinate. It does not read an uncertainty field, because a raster lookup is a lookup, and the function signature takes a point.
So the model learns that this beetle is associated with whatever the land cover is at the courthouse, which is developed, impervious surface, in the middle of a town. One hundred and eighty-nine times.
I checked how often the uncertainty field is used. I looked at published distribution models in one subfield over four years — a sample I assembled by hand, not a systematic review, and I will not pretend it was one — and of the papers I could evaluate, the methods section mentioned filtering on coordinate uncertainty in a minority of them. Several mentioned removing duplicate coordinates, which would have caught our beetle by accident, and is why I do not think the sky is falling.
The rest is the part I have not been able to stop thinking about. A model trained on centroid-inflated records does not produce obviously wrong output. It produces slightly wrong output, biased in a specific and consistent direction: toward whatever conditions prevail at administrative centres, which are disproportionately low-elevation, disproportionately developed, and disproportionately near roads, because that is where people put county seats.
Here is where the objection lands, and it is a good one that I have had put to me by people who build these databases and know far more about them than I do.
Georeferencing legacy records is not a corner-cut. It is what makes a century of museum specimens usable at all. The alternative to a centroid with a 35-kilometre uncertainty is not a precise coordinate; it is discarding the record, and discarding it means discarding the only evidence that this beetle was in that county in 1931, which is exactly the evidence that makes it possible to say anything about change over time. The aggregators publish the uncertainty. They document the method. They have done their part of this properly and it would be unfair to leave that unsaid.
Which is why my complaint is not about the database. It is about the join, again, which is where I always seem to end up.
The uncertainty field exists on one side. The raster extraction exists on the other. Between them is a modeller writing a script, and there is no point in the pipeline where anything forces the two to meet. The default behaviour of every tool I have used is to accept a point and return a value. To use the uncertainty you have to know it is there, remember it, and add a filter, and the reward for doing so is a smaller dataset and a weaker-looking model.
The fix is not a rule about data quality. It is that the extraction function should refuse a record with no uncertainty value and warn on one above a threshold the modeller has to set explicitly. Make the default noisy. Nobody would have to be persuaded of anything.
I drove past that courthouse in June. It is a limestone building with a lawn and eleven parking spaces, and there is no habitat within four hundred metres of it that would support the animal in question, and somewhere in a model that has been cited a hundred and forty times it is prime.
