NeFut Logo NeFut
中 Admin Login

[CS.AI] Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#algorithm #Machine Learning #optimization

Hindsight relabeling improves sample efficiency in reinforcement learning by retroactively replacing a transition’s goal with the outcome actually achieved. Extending this idea to preference‑conditioned multi‑objective RL (MORL) typically involves relabeling a transition with the preference direction the agent realized rather than the one originally requested. We evaluated this extension on the continuous‑control MO‑Gymnasium suite using four off‑policy, preference‑conditioned algorithms (covering two critic backbones and two preference‑sampling schemes). The results show that the extension is often detrimental: out of 36 algorithm‑environment settings, 19 degrade by up to four standard deviations, only one improves, and the rest remain unchanged.

The degradation is not caused by noisy relabels—denoising the target yields negligible recovery, and neither prioritized sampling nor buffer‑structural changes mitigate the issue. Instead, repeated relabeling collapses the critic’s coverage onto the narrow region of the preference space that the agent happened to visit. We name this failure mode Preference Coverage Collapse. To quantify it we introduce abandoned preference mass (APM), a value‑aware statistic that captures the harm ($\rho = -0.73$), whereas a pure coverage count does not.

We propose `her_mix`, a single‑parameter convex combination that pulls the achieved preference direction back toward the requested one. Using a fixed parameter across all algorithms and environments, `her_mix` restores 16 of the 19 harmed settings to baseline performance, preserves and even improves the sole setting where relabeling helped, and reduces abandoned preference mass from $69\%$ to $6\%$. The key insight is that protecting coverage over the preference simplex—not merely filtering noisy relabels—makes hindsight relabeling safe for MORL.

Review: The paper convincingly diagnoses a hidden failure mode of hindsight relabeling in multi‑objective settings and offers a simple yet effective remedy, advancing the robustness of MORL methods.

Original Source: https://arxiv.org/abs/2609.26918

[h] Back to Home