Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models
LLM estimates often fail basic probability checks when subpopulation answers are aggregated back up.
The paper tests whether model outputs obey the law of total probability across recursive population partitions. Using binary trees, the authors prompt frontier models on increasingly fine-grained subgroups, then compare aggregated estimates with direct population-level estimates. They report widespread self-consistency violations across domains and models. In persona prompting, finer-grained subgroup estimates often match human reference data better than direct aggregate answers, a pattern they call the “macro fallacy.” ArXiv · AI/CL/LG's note
The paper tests whether model outputs obey the law of total probability across recursive population partitions. Using binary trees, the authors prompt frontier models on increasingly fine-grained subgroups, then compare aggregated estimates with direct population-level estimates. They report widespread self-consistency violations across domains and models. In persona prompting, finer-grained subgroup estimates often match human reference data better than direct aggregate answers, a pattern they call the “macro fallacy.” ArXiv · AI/CL/LG's note
score 4