Distance generalization in transformers: why bother with positional encoding?
The paper tests whether transformers can handle unseen token delays without changing context length.
Nevermann and Gros define “distance generalization” as performance when source-to-recall gaps shift between training and inference. They use two synthetic delay-copy tasks, with full and selective copying, to compare RoPE, ALiBi, and no positional encoding. The study asks how positional schemes, training-distance diversity, and transfer across distances affect results, but the abstract stops at saying the mechanisms need closer study. ArXiv · AI/CL/LG's note
Nevermann and Gros define “distance generalization” as performance when source-to-recall gaps shift between training and inference. They use two synthetic delay-copy tasks, with full and selective copying, to compare RoPE, ALiBi, and no positional encoding. The study asks how positional schemes, training-distance diversity, and transfer across distances affect results, but the abstract stops at saying the mechanisms need closer study. ArXiv · AI/CL/LG's note
score 4