Megadose AI progress, ranked and analyzed.

CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding

· HF Daily Papers ·
The paper claims a VLM can improve ultra-long video grounding by evolving both its tool policy and the tools themselves, without retraining the model.

CoEvoWhen distills experience from the model’s reasoning traces into a reusable skill for choosing between long-range image observations and fine-grained video observations. Its updater can also modify existing media tools or create new ones for evidence collection. Across five benchmarks and three VLMs, the authors report better temporal grounding accuracy and lower visual token cost at inference. They also say the evolved skill transfers to general long-video QA without task-specific evolution. HF Daily Papers' note

score 4

Categories: Research