Process Reward Models (PRMs) provide fine‑grained signals for evaluating intermediate reasoning states, achieving notable test‑time scaling and reinforcement learning performance, yet their training heavily relies on costly process annotations.\ \ Existing PRMs model reasoning prefixes independently, offering no explicit way to leverage the final outcome to guide the learning of intermediate states, which limits the synergy between process and outcome supervision.\ \ We introduce Reasoning State Propagation (RSP), which assigns a binary validity state to each reasoning prefix and models transitions between successive states along the reasoning trajectory. Specifically, RSP predicts a break probability (a valid state becoming invalid) and a repair probability (an invalid state returning to valid), and propagates these transitions to connect intermediate states with the final state.\ \ During training, process annotations continue to supervise the validity of intermediate states, while outcome labels supervise the final state and back‑propagate learning signals through the state transitions, achieving a complementary supervision scheme.\ \ Experiments span reasoning search, response selection, and reinforcement learning. Across all tasks, RSP consistently outperforms representative PRM baselines, delivering an average improvement of 5.6% over Qwen2.5‑Math‑PRM in beam search and 2.1% in reinforcement learning.\ \ Review: By explicitly modeling state transitions, RSP effectively back‑propagates outcome information to earlier reasoning steps, reducing dependence on extensive process annotations and offering a more efficient learning framework for large‑scale reasoning tasks.