World Embedding Benchmark
The benchmark tests whether video embeddings actually carry usable physical information.
The paper introduces 8,000 controlled simulation videos across mechanics, fluids, dynamics, optics, and electromagnetism. Its tasks separate text-video physical alignment from the ability to recover quantitative properties. Existing omnimodal embedding models perform weakly on retrieval and near chance on within-family matching, though probes can still extract useful physical signals. Physics-specific contrastive training improves alignment tasks but hurts property regression, and retrieved reference videos improved MiniMax-H3 generations in the authors’ experiments. HF Daily Papers' note
The paper introduces 8,000 controlled simulation videos across mechanics, fluids, dynamics, optics, and electromagnetism. Its tasks separate text-video physical alignment from the ability to recover quantitative properties. Existing omnimodal embedding models perform weakly on retrieval and near chance on within-family matching, though probes can still extract useful physical signals. Physics-specific contrastive training improves alignment tasks but hurts property regression, and retrieved reference videos improved MiniMax-H3 generations in the authors’ experiments. HF Daily Papers' note
score 5