NeFut Logo NeFut
Admin Login

[CS.AI] Do Frontier Models Seek Safety Evidence Before Acting?

Published at: 2026-09-17 22:00 Last updated: 2026-09-18 00:46
#AI #Machine Learning #LLM

We introduce the SAFE benchmark to study whether models actively acquire safety‑relevant evidence before making deployment decisions. Each decision offers optional evidence that varies in retrieval cost, probability of occurrence, severity, and presentation format.

Experiments on GPT‑5.5, o3, Claude Opus 4.8 and Claude Sonnet 4.6 reveal distinct evidence‑acquisition policies: Opus inspects almost by default, o3 skips most often and is highly threshold‑sensitive, while GPT‑5.5 and Sonnet fall in between. Inspection rates rise sharply with severity, drop with retrieval cost, and are only weakly affected by probability – raising the stated likelihood from 10% to 70% changes inspection by at most 21 percentage points.

Across models, Stage 1 rationales are dominated by expected‑value reasoning. A cost‑obligation decomposition shows that avoidance is driven mainly by retrieval friction and explicit threats to deployment payoff rather than by remediation duties created by knowing the risk.

Counterfactual interventions expose a mismatch: evidence framing can strongly shift decisions near the inspection boundary yet is rarely mentioned in explanations, whereas probability is frequently cited despite its limited causal impact.

These findings suggest that deployment‑time safety depends not only on how models react to known risks, but also on whether they acquire the evidence needed to determine that acting is safe.

Review

Original Source: https://arxiv.org/abs/2609.17865

[h] Back to Home