Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
The paper tests whether LLMs’ stated reasons actually match what changes their decisions.
The authors treat cited factors as either necessary, where changing one changes the output, or sufficient, where keeping one preserves it. Across Claude, GPT, and Gemini models, cited factor rankings only moderately tracked measured influence in two synthetic tasks. In advisor recommendations, uncited factors often outranked cited ones under both tests; prompt monitoring performed better but still showed gaps. ArXiv · AI/CL/LG's note
The authors treat cited factors as either necessary, where changing one changes the output, or sufficient, where keeping one preserves it. Across Claude, GPT, and Gemini models, cited factor rankings only moderately tracked measured influence in two synthetic tasks. In advisor recommendations, uncited factors often outranked cited ones under both tests; prompt monitoring performed better but still showed gaps. ArXiv · AI/CL/LG's note
score 4