UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
A single multimodal model is trained to improve itself by learning from its own critique-conditioned version.
UniEvo-VL has the model play both student and teacher: the student sees the plain prompt, while the teacher also sees a self-critique. The training minimizes divergence between their diffusion distributions over the student’s own sampled trajectories. Built on Qwen-image-2512, it reports GenEval improvement from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53. The paper also notes uneven gains, including mixed text-rendering outcomes. HF Daily Papers' note
UniEvo-VL has the model play both student and teacher: the student sees the plain prompt, while the teacher also sees a self-critique. The training minimizes divergence between their diffusion distributions over the student’s own sampled trajectories. Built on Qwen-image-2512, it reports GenEval improvement from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53. The paper also notes uneven gains, including mixed text-rendering outcomes. HF Daily Papers' note
score 4