Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents
In this audit, every LLM failed as a true stand-in for human survey data.
The paper tests 37 models against a Lithuanian organizational-psychology dataset and scores whether they preserve human survey structure, reliability, demographic effects, and downstream validity. LLMs often matched the rough direction of human relationships, but a Gaussian-copula statistical baseline beat every model on key similarity measures. The authors also found a strong acquiescence bias, weak predictive validity on held-out humans, and fabricated mediation effects on placebo paths. HF Daily Papers' note
The paper tests 37 models against a Lithuanian organizational-psychology dataset and scores whether they preserve human survey structure, reliability, demographic effects, and downstream validity. LLMs often matched the rough direction of human relationships, but a Gaussian-copula statistical baseline beat every model on key similarity measures. The authors also found a strong acquiescence bias, weak predictive validity on held-out humans, and fabricated mediation effects on placebo paths. HF Daily Papers' note
score 4