The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads
In MirageBench, all 12 tested LLMs fabricated unsupported user attributes at high rates.
The paper defines the failure as over-inference: personalization claims that go beyond the evidence in a user profile. Across 143,616 judged claims, every model over-inferred 35% to 49% of the time, with a cross-model mean of 41.6%. The authors also report a “Self-Monitoring Inversion,” where models that rated themselves as less prone to over-inference tended to be judged as worse on it. Their conclusion is that external checking is a firmer basis for trustworthy personalization than a model’s own self-audit. HF Daily Papers' note
The paper defines the failure as over-inference: personalization claims that go beyond the evidence in a user profile. Across 143,616 judged claims, every model over-inferred 35% to 49% of the time, with a cross-model mean of 41.6%. The authors also report a “Self-Monitoring Inversion,” where models that rated themselves as less prone to over-inference tended to be judged as worse on it. Their conclusion is that external checking is a firmer basis for trustworthy personalization than a model’s own self-audit. HF Daily Papers' note
score 5