NeFut Logo NeFut
中 Admin Login

[CS.AI] CoRe: Co‑Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models

Published at: 2026-09-30 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #optimization

Latent Reward Models (LRMs) score intermediate states of video diffusion models directly in latent space, enabling efficient alignment. Optimizing against a fixed reward, however, quickly leads to latent reward hacking: the predicted reward stays high while perceptual fidelity and motion consistency deteriorate.

Our analysis identifies distributional escape as the root cause. After a few hundred updates the generator moves beyond the training support of the reward model, so the scores no longer reflect true video quality.

To address this, we introduce CoRe, a co‑evolving reward framework that treats latent‑space alignment as a dynamic interaction between generator and reward model. Instead of optimizing against a static proxy, CoRe continuously refits the reward model on the generator's current samples while anchoring it to real‑video preferences, preventing the generator from gaining reward by drifting away from the data distribution.

Experiments on Wan2.1‑T2V‑1.3B show that CoRe consistently improves generation quality over both the pretrained baseline and prior alignment methods, while avoiding the quality collapse observed with fixed‑reward optimization.

Review: CoRe’s joint evolution of generator and reward model offers a robust solution to reward hacking, making latent‑space alignment more reliable for high‑quality video synthesis.

Original Source: https://arxiv.org/abs/2609.36245

[h] Back to Home