How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
The paper argues that alignment tuning installs distinct, steerable bias directions inside LLM representations.
Across five model families and seven bias types, the authors find that pretrained base models show little of this cue-driven susceptibility. In aligned models, each bias can be decoded from hidden states and causally shifted to recover unbiased answers. The directions remain mostly separate, even when the behaviors look similar. ArXiv · AI/CL/LG's note
Across five model families and seven bias types, the authors find that pretrained base models show little of this cue-driven susceptibility. In aligned models, each bias can be decoded from hidden states and causally shifted to recover unbiased answers. The directions remain mostly separate, even when the behaviors look similar. ArXiv · AI/CL/LG's note
score 5