NeFut Logo NeFut
中 Admin Login

[CS.AI] Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis

Published at: 2026-09-28 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #LLM

Existing attacks and defenses for diffusion‑based large language models (dLLMs) usually target isolated vulnerabilities and lack a unified explanatory framework. We interpret safety alignment as shaping the denoising energy landscape: a well‑aligned model creates an energy barrier separating safe outputs from harmful ones, steering harmful queries toward the safe region.\ \ Current jailbreak attacks reduce to two strategies for bypassing this barrier: obscuring the query’s safety disposition at initialization, or intervening mid‑trajectory to force the denoising path across the barrier.\ \ Leveraging the fact that masked diffusion models minimise kinetic energy $E_k$ during denoising, we derive three complementary, training‑free detection signals:\

  1. A step‑0 ratio that reads the initial safety disposition from the logit distribution before generation starts;\
  2. Two trajectory‑velocity signals that monitor kinetic energy in complementary subspaces of the logit space.\ \ An attack must either reveal its intent at initialization, caught by the step‑0 ratio, or expend kinetic energy to cross the barrier, leaving a trace in at least one subspace’s velocity signal. Consequently, the three signals cover each other’s blind spots in the energy budget by construction.\ \ We evaluate across three dense dLLMs (LLaDA‑8B, LLaDA‑1.5, Dream‑7B) and a sparse mixture‑of‑experts dLLM (LLaDA‑MoE‑7B), confirming the complementarity of the signals. In stress tests of known attacks, every configuration that evades detection also fails to produce harmful content, suggesting that detection and barrier‑crossing thresholds are hard to separate.\ \ Review: This work offers a unified energy‑landscape perspective that explains the fundamental mechanism of jailbreaks and proposes lightweight kinetic‑energy‑based detection methods. The experiments demonstrate robustness across multiple models, providing a promising direction for securing dLLMs.
Original Source: https://arxiv.org/abs/2609.30841

[h] Back to Home