Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing
ATR judges whether silent mouth movement matches a dubbing line in both wording and timing.
The paper says existing visual speech and video-language models can miss temporal errors even after fine-tuning. ATR first aligns lip-frame representations with phonetic units, then gives an LLM local alignment evidence plus a global score for the final judgment. On a seven-language benchmark, it reports mean AUC gains of about 50% or more over Qwen3.5 SFT baselines, with transfer to three unseen MuAViC languages. It also improves two dubbing-specific tasks: reranking dub lines and assigning scripts to clips. HF Daily Papers' note
The paper says existing visual speech and video-language models can miss temporal errors even after fine-tuning. ATR first aligns lip-frame representations with phonetic units, then gives an LLM local alignment evidence plus a global score for the final judgment. On a seven-language benchmark, it reports mean AUC gains of about 50% or more over Qwen3.5 SFT baselines, with transfer to three unseen MuAViC languages. It also improves two dubbing-specific tasks: reranking dub lines and assigning scripts to clips. HF Daily Papers' note
score 4