DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
DramaChain Bench tests short-drama generation as a full production pipeline, not just isolated video clips.
The benchmark covers script, storyboard, keyframe imagery, shot-level video, and final short-drama assembly. Its dataset includes 5,785 items, each scored by three professional annotators, producing 17,488 valid scores and 255,925 traceable defect records. The paper says those annotations show upstream errors cascading into final episode quality. Its automated DramaChain Agentic Judge reproduces human model rankings with a mean PLCC of 0.918.
HF Daily Papers' note
The benchmark covers script, storyboard, keyframe imagery, shot-level video, and final short-drama assembly. Its dataset includes 5,785 items, each scored by three professional annotators, producing 17,488 valid scores and 255,925 traceable defect records. The paper says those annotations show upstream errors cascading into final episode quality. Its automated DramaChain Agentic Judge reproduces human model rankings with a mean PLCC of 0.918.
HF Daily Papers' note
score 4