NeFut Logo NeFut
Admin Login

[CS.AI] RoCo-ACE: Rollout-Conditioned Online Distillation for Knowledge Injection

Published at: 2026-07-29 22:00 Last updated: 2026-07-30 03:24
#AI #Machine Learning #Neural

Abstract

Knowledge injection updates pretrained MLLMs with new factual or domain-specific knowledge, but fitting full authoritative answers can cause drift in non-updated behavior. Online distillation mitigates this drift by training on model-generated rollouts, yet uniform reference-conditioned distillation provides coarse supervision: it can under-emphasize reference-supported rollout tokens and supervise omitted facts only indirectly.

We introduce RoCo-ACE, a rollout-conditioned online distillation objective for knowledge injection. RoCo uses same-rollout reference-free/reference-conditioned likelihood contrast to reallocate additional distillation weight to reference-supported rollout tokens, while ACE adds sparse reference-side anchored correction for authoritative anchors omitted from the rollout without full-answer imitation.

Across three knowledge-injection settings, six retention benchmarks, multiple baselines, and multiple base models, RoCo-ACE achieves the best injected-knowledge accuracy among compared methods while keeping evaluated retention close to the base model.

Blogger's Review: RoCo-ACE effectively addresses the drift issue in knowledge injection through its innovative rollout-conditioned distillation approach, demonstrating the potential to enhance knowledge accuracy while maintaining model behavior stability, providing new insights for future multilingual model optimization.

Original Source: https://arxiv.org/abs/2607.24771

[h] Back to Home