Selecting The Most Informative Tokens in Natural Language Autoencoders
Auditors may not need to explain every token to catch prompt-injection or concealment behavior.
The paper tests 4.7 million explanations from natural language autoencoders and asks which token positions are worth inspecting. A ranker based only on chat structure usually picked more relevant positions than signals from model computation. In three of four datasets, explaining just 5% of token positions kept nearly all of the success rate of explaining every position. The authors also report that pretrained verbalizers could recover words a model had learned to conceal through fine-tuning. HF Daily Papers' note
The paper tests 4.7 million explanations from natural language autoencoders and asks which token positions are worth inspecting. A ranker based only on chat structure usually picked more relevant positions than signals from model computation. In three of four datasets, explaining just 5% of token positions kept nearly all of the success rate of explaining every position. The authors also report that pretrained verbalizers could recover words a model had learned to conceal through fine-tuning. HF Daily Papers' note
score 4