Megadose Built for builders and researchers.

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

· HF Daily Papers ·
The paper’s claim is a live-commerce model that aligns speech, video, product imagery, on-screen text, and user questions in one representation.

TLive-Omni is built for noisy, long e-commerce streams where product facts are split across multiple modalities. The authors introduce Per-vGrid to pair timestamped video grids with matching audio for temporal alignment. Training moves from supervised omni-modal perception to instruction following, then Faithful-RFT to improve answer faithfulness and response quality under real-time constraints. Experiments are reported as strong on live-commerce tasks, with generalization to broader benchmarks. Source: HF Daily Papers' note

score 4

Categories: Research