As large language models (LLMs) become autonomous agents that can alter real‑world states, guaranteeing safety across multi‑step workflows has turned into a critical problem. Recent work has shifted from single‑turn to multi‑turn evaluation, yet two major limitations remain:
- Step‑level methods treat each action in isolation, missing the accumulation of risk;
- Trajectory‑level evaluations are performed post‑hoc, offering no chance for timely intervention.
To overcome these issues, we formalize Decoupled Proactive Safety Monitoring along three dimensions: whether to intervene, when to intervene, and what the risk is. Building on this framework, we introduce PASTABench, a benchmark containing 1,139 multi‑turn trajectories that span 5 risk categories and 13 sub‑categories.
We also define the Optimal Intervention Window (OIW), anchored by annotated Earliest‑Signal and Trigger turns, to quantify the timeliness of interventions. Evaluation of 16 LLMs shows that proactive intervention is still largely unsolved, with the best model achieving only a 40.74% optimal‑timing intervention rate. Fine‑grained diagnosis reveals pervasive lexical overfitting: smaller models obtain competitive safety scores by being hypersensitive to keywords, but their proactive capability collapses when hazard vocabulary is neutralized, indicating a lack of genuine risk understanding.
Review: PASTABench provides a systematic benchmark and timeliness metric for assessing LLM proactive safety, exposing fundamental shortcomings in current models’ risk perception and intervention timing, and pointing to clear directions for future improvement.