Megadose Built for builders and researchers.

Personalized Privacy Control in LLMs via Attention Head Intervention

· ArXiv · AI/CL/LG ·
Prompted privacy rules were ignored often enough that the paper proposes changing model attention at inference time.

The authors introduce “personalized privacy” for cases where disclosure limits differ by user even in the same context. They also present P3Bench, a benchmark built from contextual privacy policies plus user-specific disclosure policies. In experiments, prompt-based controls failed frequently, with reported policy ignorance ratios of 51.25% for Qwen2.5-7B and 74.28% for Gemma3-4B. Their proposed method, Repair, intervenes on attention heads during inference to push responses toward the user’s stated privacy policy. ArXiv · AI/CL/LG's note

score 5

Categories: Research