Megadose AI progress, ranked and analyzed.

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

· ArXiv · AI/CL/LG ·
The paper argues that reconstruction scores can bless activation explanations while missing false individual claims.

Dingeto reports that a Qwen-2.5-7B verbalizer reconstructed activations above chance even though only about 2% of specific claims were reconstruction-dependent. In synthetic tests, the standard setup learned private codes in all five runs. The proposed RECAP method trains auxiliary linear heads so designated internal content stays independently decodable, with probes still flagging lies under adversarial edits. ArXiv · AI/CL/LG's note

score 5

Categories: Research