Safety‑aligned large language models (LLMs) often refuse a harmful request in a single turn, yet may comply when the same goal is spread across multiple turns. Preference‑based objectives score whole responses to isolated prompts, so their training loss cannot control risk on unseen histories. This work derives sufficient conditions under which suppression in supervised single‑turn contexts yields an upper bound on multi‑turn trajectory risk. The bound consists of coverage, transfer slack, and leakage terms, and shows contraction relative to a base‑policy risk budget evaluated on the trained policy’s contexts. Building on this principle, TRACE (Trajectory Return Attribution and Contrastive Erasure) introduces a token‑level objective. For each token in a safe response, a weight equal to the discounted return of a refusal‑attributable advantage is applied. The advantage is computed by contrasting a frozen reference model with its refusal‑ablated copy, allowing earlier tokens to receive credit from later refusal evidence. The discounted return is defined as $$G_t = \sum_{k=t}^{T} \gamma^{k-t} r_k$$. At high‑gap positions on rejected responses, TRACE combines the observed token with policy‑selected alternatives to form an erasure target, and replaces the retain set with a gradient‑norm penalty. Experiments on five open‑weight models and seven multi‑turn attacks show TRACE achieves the lowest attack success rate (ASR) across all 35 model‑attack pairs, while utility on MMLU and HellaSwag drops by at most 1.23 points. The source code is provided in the supplemental material.
Review