OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport
OnTrack tries to catch failing agent runs while they are still unfolding, not after the damage is done.
The paper proposes a streaming monitor that compares an agent’s steps and dependencies with successful prior trajectories, adding about a millisecond per step. Its usefulness depends on how much reference data is available, from detecting plan violations with full access to spotting loops, stalls, and repeated tool calls with only live logs. On SWE-bench trajectories, the authors say the first eight steps were enough to rank failing runs below successful ones better than content-similarity baselines. An abort policy saved about 18% of compute on failing runs, with 5 of 6 aborts judged correct. ArXiv · AI/CL/LG's note
The paper proposes a streaming monitor that compares an agent’s steps and dependencies with successful prior trajectories, adding about a millisecond per step. Its usefulness depends on how much reference data is available, from detecting plan violations with full access to spotting loops, stalls, and repeated tool calls with only live logs. On SWE-bench trajectories, the authors say the first eight steps were enough to rank failing runs below successful ones better than content-similarity baselines. An abort policy saved about 18% of compute on failing runs, with 5 of 6 aborts judged correct. ArXiv · AI/CL/LG's note
score 5