Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
Longer training context helped only up to a point, then made models lean less on stored knowledge.
The paper argues that abundant relevant context can lower the pressure for a model to internalize facts parametrically. In pretraining, larger context windows improved several measures until an intermediate optimum, after which performance declined. In fine-tuning, task-relevant context helped when support was present at test time but hurt robustness when context was missing or misleading. The authors trace the effect to training pressure shifting from feed-forward networks toward attention modules. ArXiv · AI/CL/LG's note
The paper argues that abundant relevant context can lower the pressure for a model to internalize facts parametrically. In pretraining, larger context windows improved several measures until an intermediate optimum, after which performance declined. In fine-tuning, task-relevant context helped when support was present at test time but hurt robustness when context was missing or misleading. The authors trace the effect to training pressure shifting from feed-forward networks toward attention modules. ArXiv · AI/CL/LG's note
score 6