Before You Poll with LLMs: A Deliberative Diagnostic Framework
The paper says LLM personas can match static opinions while failing to mimic how people change their minds after new information.
Wali and Tayyab propose a deliberative polling diagnostic that compares human and model belief shifts after the same informational interventions. Tested on five frontier models using America in One Room data, every model failed in a distinct way. GPT-5.1 showed reversal on outgroup questions, while Gemini 2.0 Flash, Claude Sonnet 4.5, and Llama 3.3 70B overshot human movement; DeepSeek V3 barely moved. The authors call the pattern self-sycophancy: models conforming to internal persona stereotypes instead of reasoning from the supplied material. ArXiv · AI/CL/LG's note
Wali and Tayyab propose a deliberative polling diagnostic that compares human and model belief shifts after the same informational interventions. Tested on five frontier models using America in One Room data, every model failed in a distinct way. GPT-5.1 showed reversal on outgroup questions, while Gemini 2.0 Flash, Claude Sonnet 4.5, and Llama 3.3 70B overshot human movement; DeepSeek V3 barely moved. The authors call the pattern self-sycophancy: models conforming to internal persona stereotypes instead of reasoning from the supplied material. ArXiv · AI/CL/LG's note
score 4