Megadose AI progress, ranked daily.

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

· HF Daily Papers ·
LAION says it has released an open multimodal video corpus built from 80 million downloaded videos totaling 10 million hours.

The paper says the source pool began with 1.3 billion platform-specific video URLs collected from CommonCrawl. The authors use scene detection to cut clips and generate synthetic video and audio captions for pre-training across video, audio, and image tasks. They report competitive benchmark results, with performance improving as training or model scale increases. HF Daily Papers' note

score 7

Categories: OSS & Tools, Research