Megadose AI progress, ranked and analyzed.

Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning

· ArXiv · AI/CL/LG ·
Forgetting a fact in one language can still leave it reachable through another.

The paper tests that failure mode across 174 language-script pairs and 25 paraphrase types. It frames the fix as “language budgeted multilingual unlearning,” where only selected languages get forget supervision. The authors propose COVER, a selection method that uses benign calibration data and the frozen model to choose better source languages. Across three model families and two forget sets, COVER cut held-out residual access by 7.8% to 27.3% versus uniform source selection. ArXiv · AI/CL/LG's note

score 6

Categories: Research