OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation
OmniVBench tests whether reference-to-video models preserve and route specific reference factors, not just whether the output looks broadly consistent.
The paper introduces a benchmark covering 7 task families and 18 fine-grained tasks across content, motion, style, structure, narrative, and multi-reference settings. Its evaluation uses 12,172 case-specific checklist items to check preservation, disentanglement, binding, and instruction following. The authors also release an Omni-R2V Dataset with 340K processed training samples drawn mainly from professional video footage. Their tests of open- and closed-source models show clear remaining gaps across task families and evaluation dimensions. HF Daily Papers' note
The paper introduces a benchmark covering 7 task families and 18 fine-grained tasks across content, motion, style, structure, narrative, and multi-reference settings. Its evaluation uses 12,172 case-specific checklist items to check preservation, disentanglement, binding, and instruction following. The authors also release an Omni-R2V Dataset with 340K processed training samples drawn mainly from professional video footage. Their tests of open- and closed-source models show clear remaining gaps across task families and evaluation dimensions. HF Daily Papers' note
score 4