Attention-Path Fragility as an Uncertainty Signal in Large Language Models
The paper tests uncertainty by seeing whether confident answers break when attention heads are masked.
The authors introduce ASMI, a training-free estimator that compares subnetworks created by masking attention heads. In grounded QA, they say it adds error-prediction signal beyond confidence and entropy, especially for confident-but-fragile predictions. The method is strongest when answers depend on provided context, and falls back to weak or baseline behavior on parametric QA as expected. Sem-ASMI can read the signal from one greedy response, while stronger variants reuse samples already drawn for comparison baselines. ArXiv · AI/CL/LG's note
The authors introduce ASMI, a training-free estimator that compares subnetworks created by masking attention heads. In grounded QA, they say it adds error-prediction signal beyond confidence and entropy, especially for confident-but-fragile predictions. The method is strongest when answers depend on provided context, and falls back to weak or baseline behavior on parametric QA as expected. Sem-ASMI can read the signal from one greedy response, while stronger variants reuse samples already drawn for comparison baselines. ArXiv · AI/CL/LG's note
score 5