UniEvo-VL Self-Distillation Lifts Qwen-Image-2512 GenEval from 0.747 to 0.808
UniEvo-VL demonstrates that a multimodal model can improve its own image generation by distilling from its internal self-critiques. Gains are measurable on GenEval and GenEval2 but remain non-uniform across tasks. The result tightens the link between judge capability and self-evolution ceiling.
The framework replaces external teachers with internal privileged critiques. The student receives the base prompt while the teacher receives the same prompt plus its own prior critique. Training aligns per-state diffusion distributions across the student's sampling paths without additional parameters or separate models.
On the reported benchmark the 0.061 GenEval gain and 2.56-point Soft-TIFA lift occur after a single self-distillation pass. Mixed text-rendering results indicate the improvement is task-dependent rather than uniform. Stronger external judges such as GPT5.6-Luna raise the observed ceiling, confirming that base judge quality bounds the self-evolution limit.
The approach continues the line of test-time compute self-improvement seen in earlier diffusion and language-model work. It removes the requirement for a larger fixed teacher yet inherits any biases present in the model's initial critique distribution. Operational deployment therefore depends on whether the observed ceiling remains above user thresholds after repeated iterations.
Next steps include scaling the loop to video and audio modalities and measuring retention of reflection sensitivity after multiple self-distillation cycles.
Leskovec lab: UniEvo-VL loops will exceed external-teacher GenEval of 0.82 within two additional self-distillation rounds by March 2027.
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2609.38721)
- [2]Supporting Source(https://arxiv.org/abs/2409.11340)