Learning Length-Extrapolatable Recurrent Models
The paper argues that recurrent models fail at longer lengths because future losses lose usable “state credit,” not simply because gradients decay.
Hanwen Jiang proposes Credit Stabilization through Time, a backward-pass rescaling method that leaves the forward computation unchanged. CST stabilizes the norm of the state-credit signal without rotating the corrected component. The paper reports better performance beyond the training horizon on both synthetic tasks and real data, with gains seen up to 128x the training length. ArXiv · AI/CL/LG's note
Hanwen Jiang proposes Credit Stabilization through Time, a backward-pass rescaling method that leaves the forward computation unchanged. CST stabilizes the norm of the state-credit signal without rotating the corrected component. The paper reports better performance beyond the training horizon on both synthetic tasks and real data, with gains seen up to 128x the training length. ArXiv · AI/CL/LG's note
score 5