NeFut Logo NeFut
中 Admin Login

[CS.AI] From Local Evidence to Safety Verdicts: Causal Tracing in Vision-Language Models

Published at: 2026-10-07 22:00 Last updated: 2026-10-08 01:25
#algorithm #AI #Machine Learning

Vision-language models often need to combine an image with a textual prompt to detect a safety risk that is invisible to either modality alone. To locate where this joint safety judgment becomes accessible inside the model, we introduce SSU-Bench, a dataset of matched safe and unsafe image‑text pairs created via single‑item prompt edits or image edits with annotated target regions.

We evaluate three vision‑language models by transferring internal states between paired inputs and measuring the resulting change in the safety verdict. Across both types of counterfactual, interventions at the altered input positions are most effective in early decoder layers, while interventions at the final input token only become effective in later layers. Directions estimated from other examples produce comparable late‑layer effects.

A linear readout of the final‑token representation predicts the model’s own verdict, including incorrect ones, and cross‑model analysis reveals similar patterns of counterfactual change. These findings highlight a recurring transition point where interventions can influence a joint safety verdict and distinguish a readable model decision from a correct safety judgment.

Review

Original Source: https://arxiv.org/abs/2610.07514

[h] Back to Home