ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
ASCENT trains an agent’s LoRA weights during deployment from its own verified task attempts, rather than relying on retrieved memories or a separate teacher.
The paper frames long-horizon agent runs as experience that can be reused across a stream of related tasks. Instead of imitating the agent’s raw single attempt, ASCENT uses a frozen initial copy of the model to self-distill the verified trajectory with hindsight, then applies the signal to persistent fast weights. The method also drops invalid-action turns before distillation to improve execution efficiency. In tests on ALFWorld, WebShop, and AppWorld, the authors report better task success and interaction efficiency as experience accumulates. HF Daily Papers' note
The paper frames long-horizon agent runs as experience that can be reused across a stream of related tasks. Instead of imitating the agent’s raw single attempt, ASCENT uses a frozen initial copy of the model to self-distill the verified trajectory with hindsight, then applies the signal to persistent fast weights. The method also drops invalid-action turns before distillation to improve execution efficiency. In tests on ALFWorld, WebShop, and AppWorld, the authors report better task success and interaction efficiency as experience accumulates. HF Daily Papers' note
score 5