BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives
BLARM turns a monocular video and a static 3D mesh into a temporally coherent animated mesh without a rig.
The method represents motion as learned rigid components blended through predicted skinning weights, rather than direct vertex motion. Its deformation latents are conditioned on video features with factorized spatial-temporal attention. The paper says this keeps the animation compact, stable, and interpretable while following the source video. HF Daily Papers' note
The method represents motion as learned rigid components blended through predicted skinning weights, rather than direct vertex motion. Its deformation latents are conditioned on video features with factorized spatial-temporal attention. The paper says this keeps the animation compact, stable, and interpretable while following the source video. HF Daily Papers' note
score 5