UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
UniSwap puts face and voice replacement into one streaming model instead of stitching separate systems together.
The paper says the framework takes a source video, a reference image, and a reference voice clip, then swaps in both the target appearance and vocal timbre while keeping the source motion, scene, content, and timing. Its training setup uses a swap-and-reconstruct pipeline to work around the lack of aligned cross-identity examples. The authors also describe streaming adaptations for cached generation, reducing sampling from 30 to 3 denoising steps per block. They report strong synchronization, identity preservation, efficient streaming, and stable long-form output. HF Daily Papers' note
The paper says the framework takes a source video, a reference image, and a reference voice clip, then swaps in both the target appearance and vocal timbre while keeping the source motion, scene, content, and timing. Its training setup uses a swap-and-reconstruct pipeline to work around the lack of aligned cross-identity examples. The authors also describe streaming adaptations for cached generation, reducing sampling from 30 to 3 denoising steps per block. They report strong synchronization, identity preservation, efficient streaming, and stable long-form output. HF Daily Papers' note
score 5