ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models
The benchmark tests whether unlearning can block harmful uses of a concept without damaging its benign uses.
Kale and Harris introduce ConceptGuard, built around “dual-use concepts” that appear in both unsafe and legitimate contexts. The paper argues that existing forget/retain benchmarks mostly test isolated factual recall, missing whether a model can separate intent. In their experiments, current unlearning methods show weak contextual separation and strong trade-offs between forgetting and utility. The dataset is publicly available. ArXiv · AI/CL/LG's note
Kale and Harris introduce ConceptGuard, built around “dual-use concepts” that appear in both unsafe and legitimate contexts. The paper argues that existing forget/retain benchmarks mostly test isolated factual recall, missing whether a model can separate intent. In their experiments, current unlearning methods show weak contextual separation and strong trade-offs between forgetting and utility. The dataset is publicly available. ArXiv · AI/CL/LG's note
score 4