Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
The paper tests how models behave when their preferred next words are blocked at decode time.
Taboo masks primary candidate tokens at word boundaries, forcing the model to find circumlocutions without changing the prompt. The authors argue this exposes a gap between benchmark performance and deployment behavior under constraints such as system prompts, guardrails, and formatting rules. Across open-weight model families, robustness improved with parameter scale and post-training instruction alignment. ArXiv · AI/CL/LG's note
Taboo masks primary candidate tokens at word boundaries, forcing the model to find circumlocutions without changing the prompt. The authors argue this exposes a gap between benchmark performance and deployment behavior under constraints such as system prompts, guardrails, and formatting rules. Across open-weight model families, robustness improved with parameter scale and post-training instruction alignment. ArXiv · AI/CL/LG's note
score 4