Large Reasoning Models (LRMs) are often fine‑tuned with reinforcement learning (RL) to improve chain‑of‑thought (CoT) generation before emitting a final answer. Typical RL rewards are based solely on the final answer, offering little direct supervision for intermediate reasoning, which can cause deceptive safety alignment: the reasoning trace and the final answer convey conflicting safety signals. To evaluate this systematically, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses the safety of reasoning traces and final answers, quantifying their inconsistency. Across various LRMs and benchmarks we observe that deceptive safety alignment is pervasive under standard prompting and is dramatically amplified by prefilling attacks. Hidden‑representation analysis reveals that models discriminate safety more strongly at the final‑answer stage than during intermediate reasoning. To bridge this gap we propose SARA (Safety‑Aware Reasoning Alignment), an RL‑based approach that rewards both safety‑aware reasoning and safe final answers, encouraging early detection of harmful intent and enforcing reasoning‑answer consistency. Experiments show that SARA markedly reduces deceptive safety alignment in both standard and adversarial settings while preserving helpfulness and utility. The implementation is publicly available at https://github.com/xzhou98/SARA.
Review: The paper identifies a critical temporal safety mismatch in LRMs, introduces a concrete metric (DSAR) to measure it, and offers a practical RL solution (SARA) that aligns process and outcome safety, advancing trustworthy AI development.