Safety monitors protect language‑model agents that interact with external tools and environments, yet overly conservative monitoring generates many false alarms, exhausting review resources and eroding trust. False and genuine alarms are often interleaved in raw monitor scores, so a reliable cutoff still demands extensive manual verification. In practice, a small set of verified‑safe, non‑alarmed trajectories may be available while all alarms remain unlabeled, naturally casting false‑alarm auditing as a positive‑unlabeled (PU) ranking problem.
The main difficulty stems from monitor‑induced selection bias: observed safe references are those accepted by the monitor, whereas the hidden safe alarms we aim to recover are precisely the ones the monitor incorrectly flags. Consequently, the observed positives are a poor proxy for the positives to be retrieved.
We address this with a two‑stage framework. The first stage, Trust‑aware PU Supervision, adapts safe references toward the alarm domain and shields plausible false alarms from excessive negative pressure. The second stage, Reliability‑gated Rank Distillation, aggregates consistent ordering preferences from multiple PU reference models into a single student model. Finally, Consensus‑guided Structural Refinement improves the student ranking using hierarchical safe‑reference support, alarm relations, and predicted reference consensus.
The framework requires no alarm safety labels for fitting and leaves the underlying monitor unchanged. Across mainstream safety monitors, it achieves a macro AUPRC of $0.6444$, outperforming eight PU baselines by $5.27\sim16.98$ absolute percentage points; compared with the strongest baseline PULDA, it recovers 33.3% more false alarms at a 5% review budget.
Review