Megadose AI progress, ranked and analyzed.

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

· HF Daily Papers ·
Training on refusal rationales alone cut false refusals while preserving similar safety performance.

Kim and Kim split safety-tuning responses into boilerplate refusal statements and the explanations that justify them. Their experiments found that the stock refusal language pushes models toward shallow risk cues, making benign prompts with risky wording more likely to be rejected. Removing that boilerplate and training only on rationales improved discrimination between harmful and harmless inputs, with the effect also appearing in in-context learning tests. HF Daily Papers' note

score 4

Categories: Research