NeFut Logo NeFut
中 Admin Login

[CS.AI] Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#Machine Learning #LLM #Artificial Intelligence

Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most approaches are generative judges that require a decoding pass for each criterion, while classifiers that read token probabilities (e.g., Llama Guard) output a single fixed label per call. Jev is a model trained with reinforcement learning for calibrated decisions (RLCD) that can answer many typed questions about a single input with calibrated probabilities in one call. Its ability to detect alignment failures had not been measured. We therefore built RLCDAlignBench, which evaluates Jev on ten alignment failure types: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking. The benchmark spans 44 tests across five target models, labeled by each benchmark's scorer and, for two tests, by humans. Many failures are relational, defined against a reference such as the user's belief or an injected instruction, so the response alone does not reveal the issue. Our key idea is to separate what Jev is asked from what it sees: the wording and answer type of the question on one side, the fields of the input on the other. A single generic question achieves a median AUROC of 0.886 zero‑shot and beats supervised baselines on most benchmarks. Question wording matters little, while context matters more, mainly through fields that encode the label. Jev matches the reference scorer's agreement with human labels, surfaces label defects in existing benchmarks, and costs 63× less than LLM‑judge scorers. Code and data are released at https://github.com/sumleo/RLCDAlignBench.

Review: This work demonstrates that an RLCD‑trained model can serve as an efficient zero‑shot detector of diverse alignment failures, offering a high‑performing and cost‑effective alternative to existing judges, and provides valuable insights for future AI alignment safety research.

Original Source: https://arxiv.org/abs/2609.29429

[h] Back to Home