Megadose AI progress, ranked and analyzed.

FATE: Frame-Level Audio-Visual Temporal Embedding

· HF Daily Papers ·
FATE keeps audio and video aligned at the frame level so a model can represent both what happened and when.

The paper says existing audio-visual embeddings tend to preserve meaning while losing timing, while synchronization models capture offsets without reusable semantic representations. FATE keeps frame-level sequences, aligns them on the physical timeline, and compares strictly matched audio-video frame pairs. The authors train it with both cross-video semantic contrast and within-video temporal contrast. They report gains on temporal and semantic retrieval, zero-shot event localization competitive with fully supervised methods, and stronger agreement with human judgments as a generation evaluation metric. HF Daily Papers' note

score 5

Categories: Research