Megadose AI progress, ranked and analyzed.

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

· HF Daily Papers ·
AV-GRPO reframes joint audio-video RL as separate modality problems to improve quality, alignment, and synchronization.

The paper says direct reinforcement-learning post-training is hard because audio and video rewards get entangled, making credit assignment and fair synchronization scoring difficult. Its framework uses modality-anchored rollouts, frozen-tower optimization, and modality-specific objectives to reduce that coupling. The authors also introduce 5DAV, a decoupled training dataset with controllable difficulty across five dimensions. In experiments on JavisBench and VABench, they report gains over LTX-2.3 under both LoRA and full fine-tuning. HF Daily Papers' note

score 5

Categories: Research