I obtained my own record from one of the large commercial voter files, which you can do, and it is four hundred and some fields long. Most of it is accurate and dull: address history, the elections I voted in, which are public record in my state, my age band, a household composition estimate that is right.
Then there is a field with a two-digit number in it representing my probability of supporting one of the two parties. My state does not register voters by party. I have never declared one to anybody, ever, including at a primary, because my state does not have that either.
The number is 71.
It is not a lie and I am not alleging one. It is a model output, and the vendor documents that it is a model output, and anybody reading the technical appendix knows exactly what they are holding. That is not where the problem is.
The problem is what happens on the next hop. A campaign buys the file and builds its own targeting model. My 71 goes in as a feature. The campaign's model produces a score of its own, which is used to decide whether I get contacted and with what, and that contact history — did he open it, did he give — is written back to the file as behavioural data, where it becomes an input to the vendor's next refresh of the very score it descended from.
I have not added information to that system at any point. I have been modelled, targeted on the basis of the model, and had my response to the targeting fed back as evidence for the model.
The technical term for this is contamination and it is well understood by the people who build these things, several of whom I have talked to and none of whom were remotely defensive about it. It is understood, documented, and structurally impossible to avoid, because the only large behavioural signal available about people in non-registration states is how they responded to something a model told somebody to send them.
Here is why I am writing about it rather than shrugging. These files are now the sampling frame for a substantial amount of survey research. When a pollster stratifies by party in a non-registration state, the party variable frequently comes from a file like this one. So a modelled quantity enters a survey as though it were a demographic fact, the survey's finding is reported as evidence about the electorate, and the finding is then used to tune — you can see where this goes.
I want to be careful about how much I am claiming, because I have watched this argument get stretched into a claim that all polling is circular, which is not what I think and is not what the evidence supports. Most stratification variables are genuine — age, geography, education, actual turnout history, which is a public record and is the strongest predictor in the file. The modelled party score is one variable among many and the good survey shops know precisely which of their variables are modelled and say so in the methodology.
But the number is in there, in a field that looks exactly like the fields around it, and the fields around it are facts.
I do not know what my 71 is built from and neither does anybody outside the vendor. My best guess, from the documentation, is that it is mostly geography and consumer data. Which means a two-digit number describing my politics is substantially a description of my street, and the model has been told my street's opinion often enough that it now reports it as mine.
