Megadose Built for builders and researchers.

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

· HF Daily Papers ·
SemanTok makes shorter semantic prefixes do more of the work in autoregressive video generation.

The paper introduces a flexible video tokenizer that feeds frozen DINO features into the encoder and trains lightweight heads to reconstruct those features from each retained token prefix. The authors report that a 201M-parameter SemanTok AR model matches or beats a VideoFlexTok AR model 3.4 times larger. They also say SemanTok keeps semantic alignment on out-of-distribution classes and improves decoder alignment across noise levels, including pure noise. Short prefixes are cheaper to predict and improve generation fidelity, while later tokens carry more pixel detail. HF Daily Papers' note

score 5

Categories: Research