NeFut Logo NeFut
中 Admin Login

[CS.AI] UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

Published at: 2026-10-01 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #Open Source

Modern multimodal models unify generation and understanding, enabling them to learn from their own feedback. Motivated by this unified capability, we introduce UniEvo-VL, a framework that lets a multimodal model self‑evolve by leveraging constructive self‑correction feedback during test‑time computation. Instead of relying on a separate, often larger teacher, UniEvo-VL uses a single model as both teacher and student under different contexts. The student receives only the plain question, while the teacher conditions on a privileged critique. Training minimizes the per‑state divergence between their denoising diffusion distributions over the student’s sampling trajectories.

We build on the open‑source Qwen‑image‑2512 and evaluate with GenEval and GenEval2 Soft‑TIFA. UniEvo‑VL raises GenEval from 0.747 to 0.808 and GenEval2 Soft‑TIFA from 32.97 to 35.53. Experiments with stronger external critics such as GPT5.6‑Luna suggest that multimodal models with strong judging abilities can reach a higher self‑evolving ceiling.

Mixed text‑rendering results indicate that improvements are not uniform across all tasks, showing task‑dependent behavior. This study sheds light on the hot recursive self‑improvement line, aiming to enhance user experience when using multimodal models without external supervision.

Review

Original Source: https://arxiv.org/abs/2609.38721

[h] Back to Home