StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions
The paper tests whether answer shifts between “support” and “elimination” prompts show up as distinct internal model states.
The authors use minimally varied multiple-choice prompts and add an untrained `[STATE]` token as a place to read and intervene on residual-stream activations. In the models they test, support- and elimination-framed prompts produce separable activations, mainly in intermediate layers. Swapping those activations between paired prompts changes predictions and improves agreement across framings. ArXiv · AI/CL/LG's note
The authors use minimally varied multiple-choice prompts and add an untrained `[STATE]` token as a place to read and intervene on residual-stream activations. In the models they test, support- and elimination-framed prompts produce separable activations, mainly in intermediate layers. Swapping those activations between paired prompts changes predictions and improves agreement across framings. ArXiv · AI/CL/LG's note
score 4