VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction
VepAgent predicts future video events by reasoning through missing causal steps, not just summarizing what the model has already seen.
The paper proposes an agent framework for Video Event Prediction that tracks the transition from the last observed state to likely future events. It adds a FutureBench-4K chain-of-thought dataset for supervised fine-tuning, then uses tools for state tracking, frame retrieval, and region magnification during inference. Its reward design targets prediction accuracy, causal coherence, and visual grounding. The authors report state-of-the-art results on FutureBench and NEPBench, including gains over larger multimodal LLMs. HF Daily Papers' note
The paper proposes an agent framework for Video Event Prediction that tracks the transition from the last observed state to likely future events. It adds a FutureBench-4K chain-of-thought dataset for supervised fine-tuning, then uses tools for state tracking, frame retrieval, and region magnification during inference. Its reward design targets prediction accuracy, causal coherence, and visual grounding. The authors report state-of-the-art results on FutureBench and NEPBench, including gains over larger multimodal LLMs. HF Daily Papers' note
score 4