QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation
QuadTok uses a quadtree tokenizer to spend image tokens where the detail is, instead of forcing every region into the same grid.
The authors say the tokenizer saves about 10% of tokens on ImageNet versus a fixed 256-token grid, with comparable reconstruction fidelity. The saving transfers zero-shot to COCO at about 9%. Their 947M GPT-style generator, conditioned on a quadtree topology, reports 2.08 gFID on ImageNet 256x256. The paper also claims zero-shot spatial control from the quadtree’s preserved spatial structure. HF Daily Papers' note
The authors say the tokenizer saves about 10% of tokens on ImageNet versus a fixed 256-token grid, with comparable reconstruction fidelity. The saving transfers zero-shot to COCO at about 9%. Their 947M GPT-style generator, conditioned on a quadtree topology, reports 2.08 gFID on ImageNet 256x256. The paper also claims zero-shot spatial control from the quadtree’s preserved spatial structure. HF Daily Papers' note
score 5