Megadose AI progress, ranked and analyzed.

When Should LLMs Abstain? Chain-of-Self-Questioning for Selective Risk Control

· ArXiv · AI/CL/LG ·
CoSQ cuts wrong commitments by making models justify whether they have enough support to answer.

The paper tests a prompt-only abstention method on TruthfulQA across 11 model families. In its balanced-option setting, Grounded-CoSQ at `τ=0.90` lowers mean wrong commitments from 13.1% with chain-of-thought prompting to 8.9%, while answering 87.6% of questions. Answered accuracy rises from 86.9% to 89.7%, and the reported gains hold across all evaluated models and thresholds. A Natural Questions short-answer test is cited as supporting open-form evidence. ArXiv · AI/CL/LG's note

score 5

Categories: Research