Visual prompt engineering for video models
Editing the image input itself improved video-model reasoning across tasks.
The paper calls the method VIPE: automatically modifying task images before asking a video model to reason over them. One example turns an abstract physics scene into a photorealistic version using an image editing model. The authors report that this visual prompt engineering can outperform text prompt engineering and test-time scaling for video models. HF Daily Papers' note
The paper calls the method VIPE: automatically modifying task images before asking a video model to reason over them. One example turns an abstract physics scene into a photorealistic version using an image editing model. The authors report that this visual prompt engineering can outperform text prompt engineering and test-time scaling for video models. HF Daily Papers' note
score 5