The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation
RIDE pushes the student past the RL teacher by extrapolating hidden-state changes, not output probabilities.
The paper argues that output-space extrapolation is noisy and unstable because the model head weakens parts of the representation shift before they reach logits. Its method, RIDE, measures the residual between an RL-trained teacher and its pre-RL checkpoint at each layer and token position, then trains the student toward points beyond the teacher along that residual. The authors frame this as following a directional reward while keeping the student near the teacher. Across four teacher/base pairs, RIDE approaches or exceeds the RL teacher in every case and beats output-space extrapolation, especially when the teacher remains close to its base. HF Daily Papers' note
The paper argues that output-space extrapolation is noisy and unstable because the model head weakens parts of the representation shift before they reach logits. Its method, RIDE, measures the residual between an RL-trained teacher and its pre-RL checkpoint at each layer and token position, then trains the student toward points beyond the teacher along that residual. The authors frame this as following a directional reward while keeping the student near the teacher. Across four teacher/base pairs, RIDE approaches or exceeds the RL teacher in every case and beats output-space extrapolation, especially when the teacher remains close to its base. HF Daily Papers' note
score 4