When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
Persona skills can leak not just stored facts, but the style and traits that let an agent imitate a user.
The paper introduces AntiSkillBench, a benchmark for testing privacy leakage and impersonation risk in persona-skill pipelines. It uses 7,500 persona-grounded dialogue traces from 50 detailed profiles, then evaluates three skill-distillation strategies across three frontier agents. The authors report that risks persist across model backbones and distillation methods, including explicit attribute disclosure and behavioral impersonation. The defenses tested were limited and did not generalize reliably across risks or distillation setups. HF Daily Papers' note
The paper introduces AntiSkillBench, a benchmark for testing privacy leakage and impersonation risk in persona-skill pipelines. It uses 7,500 persona-grounded dialogue traces from 50 detailed profiles, then evaluates three skill-distillation strategies across three frontier agents. The authors report that risks persist across model backbones and distillation methods, including explicit attribute disclosure and behavioral impersonation. The defenses tested were limited and did not generalize reliably across risks or distillation setups. HF Daily Papers' note
score 4