Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence
The paper proposes a closed-loop agent that tunes prompts, seeds, and CFG scales to make image-to-video outputs follow instructions more reliably.
Its first stage uses a multimodal LLM to revise prompts against semantic checks and artifact-detection questions. A second stage applies Bayesian optimization over generation settings, guided partly by a new Video-Text Adherence score. The authors report human preference win rates up to 69% for videos produced by the agentic method over baseline outputs. ArXiv · AI/CL/LG's note
Its first stage uses a multimodal LLM to revise prompts against semantic checks and artifact-detection questions. A second stage applies Bayesian optimization over generation settings, guided partly by a new Video-Text Adherence score. The authors report human preference win rates up to 69% for videos produced by the agentic method over baseline outputs. ArXiv · AI/CL/LG's note
score 4