Megadose Built for builders and researchers.

Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

· ArXiv · AI/CL/LG ·
DepGPO trains terminal agents by tracing which commands actually fed the verifier’s final judgment.

The paper says existing credit assignment can reward or penalize terminal steps that did not matter to the outcome. Its method builds a dependency graph from execution traces, then works backward from the resources checked by the task verifier. Credit is assigned to relevant writes and the reads that supported them, redistributing trajectory advantages across steps. The authors report better task performance and more stable training on complex terminal tasks. ArXiv · AI/CL/LG's note

score 5

Categories: Research