Megadose AI progress, ranked and analyzed.

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

· ArXiv · AI/CL/LG ·
The paper tests how models behave when their preferred next words are blocked at decode time.

Taboo masks primary candidate tokens at word boundaries, forcing the model to find circumlocutions without changing the prompt. The authors argue this exposes a gap between benchmark performance and deployment behavior under constraints such as system prompts, guardrails, and formatting rules. Across open-weight model families, robustness improved with parameter scale and post-training instruction alignment. ArXiv · AI/CL/LG's note

score 4

Categories: Research