PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
The paper tests privacy leakage at the moment an agent gathers data, before it says or does anything with it.
PrivacyPeek covers 1,182 cases across seven acquisition behaviors and 16 domains. It inspects tool-call paths to see when agents collect sensitive information beyond the task, then probes whether that undisclosed data can be elicited later. Tests on 10 agents across four model families found unnecessary sensitive-data acquisition to be widespread. Prompt-level defenses reduced only a small share of the leakage. HF Daily Papers' note
PrivacyPeek covers 1,182 cases across seven acquisition behaviors and 16 domains. It inspects tool-call paths to see when agents collect sensitive information beyond the task, then probes whether that undisclosed data can be elicited later. Tests on 10 agents across four model families found unnecessary sensitive-data acquisition to be widespread. Prompt-level defenses reduced only a small share of the leakage. HF Daily Papers' note
score 5