Megadose AI progress, ranked and analyzed.

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

· HF Daily Papers ·
FuseReg trains RAEs to tolerate different encoder-layer mixes instead of locking reconstruction and generation to one fusion choice.

The paper says random subset training penalizes sensitivity to disagreement between layers. On ImageNet-256 with DINOv3-L, one FuseReg decoder handles full, sparse, and single-layer fusions without retraining, while beating fixed-fusion decoders on PSNR. Swapping in that decoder cuts unguided gFID by 27% with the same RAEv2 DiT-XL generator. Extending the same regularization into diffusion training cuts unguided gFID by 29% on DiT-Base. Source: HF Daily Papers' note

score 4

Categories: Research