TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
The paper’s claim is a live-commerce model that aligns speech, video, product imagery, on-screen text, and user questions in one representation.
TLive-Omni is built for noisy, long e-commerce streams where product facts are split across multiple modalities. The authors introduce Per-vGrid to pair timestamped video grids with matching audio for temporal alignment. Training moves from supervised omni-modal perception to instruction following, then Faithful-RFT to improve answer faithfulness and response quality under real-time constraints. Experiments are reported as strong on live-commerce tasks, with generalization to broader benchmarks. Source: HF Daily Papers' note
TLive-Omni is built for noisy, long e-commerce streams where product facts are split across multiple modalities. The authors introduce Per-vGrid to pair timestamped video grids with matching audio for temporal alignment. Training moves from supervised omni-modal perception to instruction following, then Faithful-RFT to improve answer faithfulness and response quality under real-time constraints. Experiments are reported as strong on live-commerce tasks, with generalization to broader benchmarks. Source: HF Daily Papers' note
score 4