NeFut Logo NeFut
中 Admin Login

[CS.AI] OpenJev-RLCD: A Working RLCD Implementation

Published at: 2026-10-01 22:00 Last updated: 2026-10-06 12:11
#algorithm #AI #Machine Learning

Decision models such as Jev output answer probabilities, but these are only useful when calibrated. Existing open‑source reproductions rely on supervised fine‑tuning (SFT) plus temperature scaling, while reinforcement learning from verifiable rewards (RLVR) makes reasoning models overconfident. We present a working implementation of reinforcement learning for calibrated decisions (RLCD) for reasoning models: the model first samples a rationale, then we score the answer distribution it commits to using a strictly proper scoring rule.

A variance identity shows that scoring the mixture of several samples rewards disagreeing rationales, and that RLVR is exactly this mixture objective without its diversity term. Optimized naively, the per‑rationale objective either switches reasoning off or is drowned out by policy‑gradient noise, which leads to a two‑stage recipe: first calibrate, then reinforce.

With Qwen3‑1.7B on two reasoning tasks (3 seeds, paired tests), RLCD matches or beats SFT, RFT/STaR and temperature‑scaled GRPO in accuracy and outperforms all of them in selective prediction. On the GSM8K answer‑verification benchmark a single query decides $\text{\gvTwoCovFive}\%$ of items at $\le5\%$ error, versus $\text{\gvGrpoCovFive}\%$ for GRPO. When uncertainty stems from annotator disagreement, RLCD provably cannot surpass cross‑entropy.

Code and results are available at https://github.com/ZimmyGao/openjev-rlcd.

Review

Original Source: https://arxiv.org/abs/2609.38850

[h] Back to Home