Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
The paper says LLMs often know the commonsense answer, but salient distractors can suppress it.
The authors introduce SaliTrap, a benchmark testing how explicit but irrelevant details can hijack commonsense reasoning. Across 12 state-of-the-art models, performance worsened as distractors became denser. A context-free probe recovered over 90% of the failures tied to sycophantic compliance, suggesting the knowledge was present but crowded out. Lightweight inference-time prompting narrowed the gap without retraining. HF Daily Papers' note
The authors introduce SaliTrap, a benchmark testing how explicit but irrelevant details can hijack commonsense reasoning. Across 12 state-of-the-art models, performance worsened as distractors became denser. A context-free probe recovered over 90% of the failures tied to sycophantic compliance, suggesting the knowledge was present but crowded out. Lightweight inference-time prompting narrowed the gap without retraining. HF Daily Papers' note
score 4