NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
NARU tests whether video models can follow Japanese long-form narrative shifts and read cultural context, not just spot events.
The benchmark covers 1,481 questions from 155 videos totaling 146.8 hours. Its questions span four narrative dimensions and five cultural dimensions, built through a hierarchical annotation pipeline and two native-speaker verification stages with 68 annotators. Eight tested model configurations still showed major weaknesses in long-range narrative integration and culturally grounded reasoning. HF Daily Papers' note
The benchmark covers 1,481 questions from 155 videos totaling 146.8 hours. Its questions span four narrative dimensions and five cultural dimensions, built through a hierarchical annotation pipeline and two native-speaker verification stages with 68 annotators. Eight tested model configurations still showed major weaknesses in long-range narrative integration and culturally grounded reasoning. HF Daily Papers' note
score 4