BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
BVB tests video understanding by making agents rebuild real videos as animated Blender scenes.
The benchmark runs agents through the same Mini-BVB sandbox and cost limit, then renders their Blender reconstructions for scoring. It measures both preserved spatiotemporal facts through Dual VQA and perceptual match through Latent Similarity. Across 51 configurations from 10 model families, the best model scored 88.6 on Latent Similarity but preserved only 53.7% of source-correct spatiotemporal answers. Extra reasoning improved visual similarity, but did not close the factual gap. HF Daily Papers' note
The benchmark runs agents through the same Mini-BVB sandbox and cost limit, then renders their Blender reconstructions for scoring. It measures both preserved spatiotemporal facts through Dual VQA and perceptual match through Latent Similarity. Across 51 configurations from 10 model families, the best model scored 88.6 on Latent Similarity but preserved only 53.7% of source-correct spatiotemporal answers. Extra reasoning improved visual similarity, but did not close the factual gap. HF Daily Papers' note
score 5