The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning
AdaPop changes unlearning pressure based on how entrenched a fact is.
The paper argues that popular facts are harder for LLMs to forget because they are memorized more deeply during pretraining. Its AdaPop method uses token confidence plus a popularity signal, such as Wikidata sitelinks or LLM-as-judge, to scale forgetting per fact. The authors report lower leakage than competing methods: about 5x less under paraphrased queries and about 1.6x less under adversarial reformulations. HF Daily Papers' note
The paper argues that popular facts are harder for LLMs to forget because they are memorized more deeply during pretraining. Its AdaPop method uses token confidence plus a popularity signal, such as Wikidata sitelinks or LLM-as-judge, to scale forgetting per fact. The authors report lower leakage than competing methods: about 5x less under paraphrased queries and about 1.6x less under adversarial reformulations. HF Daily Papers' note
score 5