Megadose Built for builders and researchers.

Embedding Prediction Helps Image Generation

· ArXiv · AI/CL/LG ·
NEPA conditions a diffusion transformer on predicted clean-image embeddings that are recomputed at each denoising step.

The paper tests this on class-conditional ImageNet 256x256 generation. Its NEPA model predicts the next continuous embeddings after the condition and noisy image, then feeds those predictions into a DiT generator. That adds a second network during sampling, but the authors report NEPA-DiT-XL reaches 1.32 FID when combined with REPA, using about one-third of REPA’s training compute. Source: ArXiv · AI/CL/LG's note.

score 4

Categories: Research