World Embedding Benchmark
The benchmark tests whether video embeddings carry usable physics, not just matching text.
The authors introduce 8,000 controlled simulation videos across mechanics, dynamics, fluids, optics, and electromagnetism. Pre-trained omnimodal embedding models performed weakly on retrieval and near chance on within-family video-description matching, while lightweight probes could still recover some physical properties from frozen embeddings. Physics-specific contrastive training improved alignment tasks but hurt quantitative property regression. Retrieved reference videos also improved MiniMax-H3 video generation fidelity in their experiments. ArXiv · AI/CL/LG's note
The authors introduce 8,000 controlled simulation videos across mechanics, dynamics, fluids, optics, and electromagnetism. Pre-trained omnimodal embedding models performed weakly on retrieval and near chance on within-family video-description matching, while lightweight probes could still recover some physical properties from frozen embeddings. Physics-specific contrastive training improved alignment tasks but hurt quantitative property regression. Retrieved reference videos also improved MiniMax-H3 video generation fidelity in their experiments. ArXiv · AI/CL/LG's note
score 5