Large Reasoning Models (LRMs) depend on explicit chain‑of‑thought (CoT) reasoning and large context windows to achieve strong performance on complex tasks, but these features also open new attack surfaces. We show that prepending counter‑aligned few‑shot conversations containing explicit CoT traces can systematically steer an LRM's reasoning, causing unsafe generations on harmful queries and unwarranted refusals on benign ones. This attack is formalized as SRCF (Steering Reasoning via Counter-Aligned Few-shot Conversations) and operates solely through a flexible conversational interface without requiring access to model parameters or gradients. The key insight is that SRCF exploits an adversarial generalization issue that induces representation drift, shifting the embeddings of both benign and harmful inputs in a similar direction. Motivated by this, we propose a post‑training defense, ARCF (Aligning Reasoning via Counter-Aligned Few-Shot Conversations), which exposes models to counter‑aligned conversational contexts while enforcing aligned targets. ARCF is compatible with existing post‑training methods and consistently improves safety and helpfulness without degrading utility.
Review: The paper uncovers a practical vulnerability in LRMs through conversational manipulation and offers an easily deployable mitigation, advancing the safety of large‑scale reasoning models.