Megadose AI progress, ranked and analyzed.

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

· HF Daily Papers ·
The paper proposes a runtime test that masks a model’s preferred next tokens to see how well it can recover off its usual path.

Decoding-Level Taboo works without changing the prompt, intervening directly in logit space at word boundaries. The authors say this forces “machine circumlocution,” exposing gaps between benchmark comfort and deployment constraints. In tests across open-weight model families, robustness generally improved with scale and instruction alignment. They frame Taboo as useful for synthetic data generation, safety-guardrail stress tests, and pre-deployment reliability audits. HF Daily Papers' note

score 5

Categories: Research