AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
AV-GRPO reframes joint audio-video RL as separate modality problems to improve quality, alignment, and synchronization.
The paper says direct reinforcement-learning post-training is hard because audio and video rewards get entangled, making credit assignment and fair synchronization scoring difficult. Its framework uses modality-anchored rollouts, frozen-tower optimization, and modality-specific objectives to reduce that coupling. The authors also introduce 5DAV, a decoupled training dataset with controllable difficulty across five dimensions. In experiments on JavisBench and VABench, they report gains over LTX-2.3 under both LoRA and full fine-tuning. HF Daily Papers' note
The paper says direct reinforcement-learning post-training is hard because audio and video rewards get entangled, making credit assignment and fair synchronization scoring difficult. Its framework uses modality-anchored rollouts, frozen-tower optimization, and modality-specific objectives to reduce that coupling. The authors also introduce 5DAV, a decoupled training dataset with controllable difficulty across five dimensions. In experiments on JavisBench and VABench, they report gains over LTX-2.3 under both LoRA and full fine-tuning. HF Daily Papers' note
score 5