SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
SemanTok makes shorter semantic prefixes do more of the work in autoregressive video generation.
The paper introduces a flexible video tokenizer that feeds frozen DINO features into the encoder and trains lightweight heads to reconstruct those features from each retained token prefix. The authors report that a 201M-parameter SemanTok AR model matches or beats a VideoFlexTok AR model 3.4 times larger. They also say SemanTok keeps semantic alignment on out-of-distribution classes and improves decoder alignment across noise levels, including pure noise. Short prefixes are cheaper to predict and improve generation fidelity, while later tokens carry more pixel detail. HF Daily Papers' note
The paper introduces a flexible video tokenizer that feeds frozen DINO features into the encoder and trains lightweight heads to reconstruct those features from each retained token prefix. The authors report that a 201M-parameter SemanTok AR model matches or beats a VideoFlexTok AR model 3.4 times larger. They also say SemanTok keeps semantic alignment on out-of-distribution classes and improves decoder alignment across noise levels, including pure noise. Short prefixes are cheaper to predict and improve generation fidelity, while later tokens carry more pixel detail. HF Daily Papers' note
score 5