PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations
Seeded user secrets were still recoverable after topic shifts in 38.7% to 54.6% of tested dialogues.
PrivDrift tests 1,000 controlled multi-turn conversations where a user discloses a secret, the chat drifts to denser unrelated content, and later probes try to extract it. The paper reports substantial leakage across three long-context LLMs, with results varying by model, secret type, and persuasion intensity. More topic drift within the tested window did not reliably reduce leakage. ArXiv · AI/CL/LG's note
PrivDrift tests 1,000 controlled multi-turn conversations where a user discloses a secret, the chat drifts to denser unrelated content, and later probes try to extract it. The paper reports substantial leakage across three long-context LLMs, with results varying by model, secret type, and persuasion intensity. More topic drift within the tested window did not reliably reduce leakage. ArXiv · AI/CL/LG's note
score 5