The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs
The paper says current trigger-tag misuse detectors fail once an attacker can alter outputs or model weights.
The authors formalize token-level and weight-level trigger-tags, then test attacks against both using phishing as the case study. Their Untag framework maps the attack surfaces across the two mechanism types. In their evaluation, the attacks make the representative trigger-tag systems “entirely ineffective.” The authors conclude these mechanisms may help in controlled settings, but should not be treated as robust detectors for open-weight models. ArXiv · AI/CL/LG's note
The authors formalize token-level and weight-level trigger-tags, then test attacks against both using phishing as the case study. Their Untag framework maps the attack surfaces across the two mechanism types. In their evaluation, the attacks make the representative trigger-tag systems “entirely ineffective.” The authors conclude these mechanisms may help in controlled settings, but should not be treated as robust detectors for open-weight models. ArXiv · AI/CL/LG's note
score 5