AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model
AnyTalk turns a static 3D character into a lip-synced speaker without character-specific animation data.
The method fine-tunes a pre-trained video diffusion model on rendered images of the target character paired with “no motion” audio embeddings. It then converts the generated talking-head video back into 3D animation by optimizing blendshape parameters. The paper says this works across different face meshes and blendshape setups, with a distilled `AnyTalk_RT` version for real-time use. HF Daily Papers' note
The method fine-tunes a pre-trained video diffusion model on rendered images of the target character paired with “no motion” audio embeddings. It then converts the generated talking-head video back into 3D animation by optimizing blendshape parameters. The paper says this works across different face meshes and blendshape setups, with a distilled `AnyTalk_RT` version for real-time use. HF Daily Papers' note
score 4