TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
The paper proposes denser, turn-level supervision for tool-using LLMs by learning from execution-conditioned hindsight.
TurnSight targets long tool-interaction chains where trajectory-level reinforcement learning gives weak credit assignment. It builds multiple hindsight views with different lookahead horizons, then keeps signals that agree directionally across horizons. Those signals are normalized across sibling rollouts and used to modulate RL advantages without flipping the original optimization direction. The authors report gains across three benchmarks. ArXiv · AI/CL/LG's note
TurnSight targets long tool-interaction chains where trajectory-level reinforcement learning gives weak credit assignment. It builds multiple hindsight views with different lookahead horizons, then keeps signals that agree directionally across horizons. Those signals are normalized across sibling rollouts and used to modulate RL advantages without flipping the original optimization direction. The authors report gains across three benchmarks. ArXiv · AI/CL/LG's note
score 5