Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, output a probability, a choice or a score for a typed question about an input state, and downstream software acts on the answer without human reading.
The robustness of these models has not been systematically measured. Existing adversarial benchmarks evaluate generated text or executed commands, while a typed RLCD model still returns a well‑formed answer even when perturbed, making assessment difficult.
Measurement is challenging because identical requests may yield different answers, most available labels are produced by the model itself, and the API performs hidden preprocessing on each request.
Our key idea is to compare each attacked decision with the model's own clean decision rather than with external labels, and to treat the change observed in an identical re‑run as the baseline.
Building on this, we introduce JevAdvBench, the first adversarial benchmark for RLCD models to our knowledge. The benchmark comprises 812 typed questions across 66 scenarios and a black‑box attack suite of 9,744 single‑edit variants, each confirmed by billed input tokens that the edit reached the model.
On jev-1.13.0, rewording stays within 1.2 percentage points of the re‑run baseline, and fields outside the schema never reach the model. In contrast, appending an unverified opinion to the state flips 12.1% of decisions, statistically tied with the strongest injected command (10.1%), and pushes 38% of confident answers below the 0.8 confidence threshold that triggers human review.
Consequently, applications built on RLCD models should treat the state as an untrusted, argued input.
Review: JevAdvBench offers a practical security evaluation framework for RLCD systems, exposing how even minor textual edits can cause substantial decision shifts and underscoring the need for rigorous input sanitization and auditing in deployment.