When Text and Numbers Disagree: Evidence Arbitration in Large Language Models
LLMs showed patterned biases when forced to choose between conflicting text, numbers, and tool outputs.
The paper introduces a synthetic benchmark where only one evidence source matches the ground truth. In tests on open-weight instruction-tuned models, choices were not random: models showed distinct preferences between textual and numerical evidence. They followed newer evidence more reliably than explicit reliability labels, and sometimes leaned too heavily on external forecasts despite conflicting context. Source: ArXiv · AI/CL/LG's note
The paper introduces a synthetic benchmark where only one evidence source matches the ground truth. In tests on open-weight instruction-tuned models, choices were not random: models showed distinct preferences between textual and numerical evidence. They followed newer evidence more reliably than explicit reliability labels, and sometimes leaned too heavily on external forecasts despite conflicting context. Source: ArXiv · AI/CL/LG's note
score 5