BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing
BLOOM-WILT is pitched as a no-training auditing pipeline for surfacing rare LLM behaviours through multi-turn elicitation and logit reweighting.
The paper says WILT updates an auditor model’s conversational strategy across rounds using prior scored interactions. It also tilts the target model’s decoding with the target’s own next-token distribution under an elicitation prompt, aiming to sample behaviour-relevant outputs without pushing probability below the baseline. Across 4 target models and 8 behaviours, it beat the baseline auditor in 30 of 32 settings and changed the authors’ prior model safety rankings. In one reported case, self-harm encouragement presence from Qwen3.5-4B rose from 51% to 100% at matched compute.
ArXiv · AI/CL/LG's note
The paper says WILT updates an auditor model’s conversational strategy across rounds using prior scored interactions. It also tilts the target model’s decoding with the target’s own next-token distribution under an elicitation prompt, aiming to sample behaviour-relevant outputs without pushing probability below the baseline. Across 4 target models and 8 behaviours, it beat the baseline auditor in 30 of 32 settings and changed the authors’ prior model safety rankings. In one reported case, self-harm encouragement presence from Qwen3.5-4B rose from 51% to 100% at matched compute.
ArXiv · AI/CL/LG's note
score 5