NeFut Logo NeFut
中 Admin Login

[CS.AI] Bounded Reachability & Jailbreak Detection via Contraction-Constrained State Space Models

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #LLM #Safety

Safety heads are lightweight classifiers attached to pretrained language models to flag potentially harmful inputs before text generation. While extensive empirical studies have examined their detection performance, formal robustness guarantees remain largely unexplored. This work investigates safety heads built on State Space Models (SSMs) and asks whether they can be certified to produce identical predictions for all inputs within a bounded perturbation in the embedding space. We prove that certification hinges on a single condition: the $l_\infty$ norm of the state transition matrix $A$ must satisfy $$\|A\|_\infty \le 1.$$ If this inequality holds, any perturbation confined to an $\epsilon$‑ball in the embedding space cannot alter the safety head’s output, thereby providing a formal defense against jailbreak attacks. The finding offers a clear mathematical guideline for constructing language models with provable safety.

Review

Original Source: https://arxiv.org/abs/2610.02853

[h] Back to Home