HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement
HOMIE targets personalized video generation where people must keep identity while interacting accurately with objects.
The paper says existing methods still struggle with subject fidelity and interaction patterns, especially for abstract objects such as logos. HOMIE uses multimodal model features inside self-attention to align reference-level semantics with VAE tokens without retraining text encoders. It also adds modality-reference embeddings to separate MLLM and VAE tokens and connect intra-subject reference images. The authors report state-of-the-art results across HOCVP tasks. HF Daily Papers' note
The paper says existing methods still struggle with subject fidelity and interaction patterns, especially for abstract objects such as logos. HOMIE uses multimodal model features inside self-attention to align reference-level semantics with VAE tokens without retraining text encoders. It also adds modality-reference embeddings to separate MLLM and VAE tokens and connect intra-subject reference images. The authors report state-of-the-art results across HOCVP tasks. HF Daily Papers' note
score 4