NeFut Logo NeFut
中 Admin Login

[CS.AI] Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Training‑free safeguards often rely on a reusable safety signal—such as an unsafe direction or a global toxic subspace—and apply it uniformly across all prompts. A controlled geometric analysis of this global‑unsafety assumption reveals a consistent coverage‑selectivity trade‑off: compact unsafe subspaces cannot cover heterogeneous unsafe semantics, while broader subspaces increasingly distort benign prompts that lie near the safety boundary. Motivated by this finding, we introduce CALM (Counterfactual Adaptive Local Modulation), a training‑free safeguard that replaces uniform global removal with prompt‑local counterfactual correction. CALM first matches unsafe‑benign anchors to route each prompt to its active unsafe categories, then minimally edits only the violating token representations toward the safe side and suppresses positively aligned unsafe residual components. Extensive evaluation shows that CALM markedly improves unsafe content suppression while preserving the utility of benign prompts, demonstrating that local counterfactual correction offers a more selective alternative to global unsafe signal removal.

Review

Original Source: https://arxiv.org/abs/2610.02300

[h] Back to Home