Megadose Built for builders and researchers.

Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL

· ArXiv · AI/CL/LG ·
A simple gradient-size control rivaled targeted attribution, making much of the rollout “cause” signal look confounded.

The paper introduces BehaviorTrace, an evaluation harness for testing training-data attribution during online RL fine-tuning with GRPO. In experiments on Qwen2.5-1.5B, ranking steps by gradient magnitude alone reached 4.2 to 4.5 times chance and matched or beat the best targeted estimator on two of three seeds. At saturated checkpoints, fluency predicted the behavior label at least as well as the gradient methods tested. The author says only one signal held across all seeds: trigger-token gradients aligned with a target built where the behavior actually appeared. ArXiv · AI/CL/LG's note

score 5

Categories: OSS & Tools, Research