Megadose AI progress, ranked and analyzed.

Selecting The Most Informative Tokens in Natural Language Autoencoders

· HF Daily Papers ·
Auditors may not need to explain every token to catch prompt-injection or concealment behavior.

The paper tests 4.7 million explanations from natural language autoencoders and asks which token positions are worth inspecting. A ranker based only on chat structure usually picked more relevant positions than signals from model computation. In three of four datasets, explaining just 5% of token positions kept nearly all of the success rate of explaining every position. The authors also report that pretrained verbalizers could recover words a model had learned to conceal through fine-tuning. HF Daily Papers' note

score 4

Categories: Research