Door-in-the-Face Requests and Refusal Behaviour in Large Language Models
The study finds “door-in-the-face” prompting can change refusal behavior, but the direction depends on the model family.
In the paper’s tests, Anthropic’s frontier models were more likely to answer a smaller follow-up request after refusing a larger related one. OpenAI and Google frontier models, and Haiku 4.5, became less compliant under the same setup. An unrelated refused request had weaker effects, suggesting the relationship between the two requests matters. Rewriting refused “usable instruction” requests as explanation requests removed nearly all refusals in the tested set. ArXiv · AI/CL/LG's note
In the paper’s tests, Anthropic’s frontier models were more likely to answer a smaller follow-up request after refusing a larger related one. OpenAI and Google frontier models, and Haiku 4.5, became less compliant under the same setup. An unrelated refused request had weaker effects, suggesting the relationship between the two requests matters. Rewriting refused “usable instruction” requests as explanation requests removed nearly all refusals in the tested set. ArXiv · AI/CL/LG's note
score 4