Megadose AI progress, ranked and analyzed.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

· ArXiv · AI/CL/LG ·
Long-context models fail when their reasoning traces copy the prompt instead of anchoring on the evidence that matters.

The paper identifies “repetitive copying” as a common failure mode across frontier long-context LLMs, worsening as context grows. The authors trace the issue to weak grounding: models copy from both key evidence and distractors, and poor focus on key evidence correlates with wrong answers. They propose GEAR, a reward-shaping method that rewards overlap with relevant evidence and penalizes overlap with irrelevant context. Across model scales and benchmarks, it improves over accuracy-only reinforcement learning by up to 4.6 average points while shortening reasoning traces and reducing copying. ArXiv · AI/CL/LG's note

score 5

Categories: Research