Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
The paper says LLMs often know the commonsense answer, but salient distractors can suppress it.
The authors introduce SaliTrap, a benchmark built around four kinds of commonsense “trap” prompts. Across 12 state-of-the-art models, they report that explicit but useless details, including numbers, can pull models into unnecessary computation and wrong compliance. Stripping away the misleading task framing recovered more than 90% of the relevant knowledge in one failure category, pointing to elicitation rather than absence of knowledge. They also report that lightweight inference-time prompting narrowed the gap without retraining. ArXiv · AI/CL/LG's note
The authors introduce SaliTrap, a benchmark built around four kinds of commonsense “trap” prompts. Across 12 state-of-the-art models, they report that explicit but useless details, including numbers, can pull models into unnecessary computation and wrong compliance. Stripping away the misleading task framing recovered more than 90% of the relevant knowledge in one failure category, pointing to elicitation rather than absence of knowledge. They also report that lightweight inference-time prompting narrowed the gap without retraining. ArXiv · AI/CL/LG's note
score 5