Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
Chimera claims major efficiency gains for long-context visual diffusion by mixing linear attention, latent attention, convolutions, and sparse experts.
The paper says full attention is becoming too expensive for high-resolution images, long videos, and multimodal inputs. Its 11B-parameter Chimera model activates 2B parameters and is scaled with a module-wise recipe called HeteroP. In experiments, the full system reaches 7.3x compute efficiency over a matched full-attention Wan-2.1 2B baseline on pretraining diffusion loss. It also reports zero-shot extrapolation from 5-second training clips to 30-second videos, with 6.5% FID degradation in the final five seconds. Source: HF Daily Papers' note.
The paper says full attention is becoming too expensive for high-resolution images, long videos, and multimodal inputs. Its 11B-parameter Chimera model activates 2B parameters and is scaled with a module-wise recipe called HeteroP. In experiments, the full system reaches 7.3x compute efficiency over a matched full-attention Wan-2.1 2B baseline on pretraining diffusion loss. It also reports zero-shot extrapolation from 5-second training clips to 30-second videos, with 6.5% FID degradation in the final five seconds. Source: HF Daily Papers' note.
score 6