NeFut Logo NeFut
中 Admin Login

[CS.AI] MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#algorithm #AI #Machine Learning

Rubric‑based reinforcement learning extends reward‑driven optimization to open‑ended tasks by assigning partial credit to each response requirement. However, rubric judges sometimes award high scores even when the required information or action is missing, a failure mode called Vacuous Credit. This award persists after the missing information is removed and can even flip the sign of a response's GRPO advantage.

MetaRubric addresses this issue by alternating evidence‑aware policy optimization with response‑guided rubric adaptation. For each prompt a counterfactual counterpart is created by changing one task‑relevant fact. During policy optimization, credit is given only when the response contains sufficient evidence to satisfy the corresponding rubric criterion. After each optimization stage, the current policy's responses guide revisions to both the original and counterfactual criteria, preserving the original rubric's meaning under each set of facts. Criterion weights are also adjusted at stage boundaries to better correct observed policy errors.

Across several model backbones, MetaRubric improves PubMedQA accuracy by 6.00–20.40 percentage points over static‑judge GRPO, with additional gains on HealthBench‑Hard and two multimodal medical benchmarks.

Review

Original Source: https://arxiv.org/abs/2610.02824

[h] Back to Home