Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
The study isolates training data quality by using procedurally generated moving letters instead of messy real-world video.
Moving Alphabet lets the authors control fonts, colors, sizes, positions, motion, and captions, then deliberately corrupt metadata to test model behavior. They report that balanced video content and duration are critical for generalization. Caption quality affects both performance and training efficiency. Guidance and fine-tuning on cleaner data help after bad captions, but do not fully undo weak pre-training data. HF Daily Papers' note
Moving Alphabet lets the authors control fonts, colors, sizes, positions, motion, and captions, then deliberately corrupt metadata to test model behavior. They report that balanced video content and duration are critical for generalization. Caption quality affects both performance and training efficiency. Guidance and fine-tuning on cleaner data help after bad captions, but do not fully undo weak pre-training data. HF Daily Papers' note
score 4