Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation
Hidden traits can survive “clean” distillation data by drifting along measurable preference directions during fine-tuning.
The paper argues that biased teacher outputs can leave small preference gaps even in neutral-looking numeric data. During SFT, those gaps push the student along a trait-aligned direction until behavior transfers. The authors propose probe-space corridor regularization to constrain that drift, reporting malicious-response transfer cut from 29.55% to 6.45% with low main-task accuracy cost. ArXiv · AI/CL/LG's note
The paper argues that biased teacher outputs can leave small preference gaps even in neutral-looking numeric data. During SFT, those gaps push the student along a trait-aligned direction until behavior transfers. The authors propose probe-space corridor regularization to constrain that drift, reporting malicious-response transfer cut from 29.55% to 6.45% with low main-task accuracy cost. ArXiv · AI/CL/LG's note
score 6