Tell me again, tell me over and over. Latent-Regressor Estimation with Repeated Noisy Measurements
-
SeriesResearch Master Defense
-
Speaker
-
LocationVU HG-01-A44
Amsterdam -
Date and time
June 22, 2026
15:00 - 17:00
Machine-learning tools, including large language models, increasingly generate variables from unstructured data. The resulting classifications can be wrong and can differ across repetitions of the same task, which may bias the analyses that use them. This thesis develops a maximum likelihood estimator that, under a conditional-independence assumption, uses repeated classifications to estimate regression coefficients and measurement quality jointly. The estimator can also incorporate validation data and self-reported certainty where these are available.
In an application to constitutional provisions and civil liberties, the estimated coefficients vary across repeated LLM classifications. When the model answers without the constitutional text, it frequently classifies absent provisions as present. The estimator recovers this false-positive pattern and shows that the LLM's self-reported certainty can be highest where classifications are most likely to be wrong.
Estimating measurement quality alongside the coefficients improves the estimates and makes the reliability of the classifications explicit, even when expert validation is limited.