Adaptive large language model (LLM) reinforcement‑learning post‑training adjusts multiple actuators online—rollout temperature, batch size, clipping, KL regularization, verifier allocation, and update budget. Three coupled problems remain: (1) a future‑risk model trained on behavior trajectories may not estimate the risk of the controller that will actually be deployed; (2) a score calibrated on logged state‑action pairs can become miscalibrated after selective action choice; (3) independent per‑resource minimum costs do not generally guarantee a feasible multi‑resource continuation. FSPO (Feedback‑State Policy‑consistent Optimizer) tackles all three jointly. It learns a policy‑consistent risk‑to‑go model whose Bellman target follows the same frozen controller used for future decisions, together with a long‑horizon utility model. Decision‑Conditioned Trajectory Calibration (DCTC) calibrates risk on cross‑fitted trajectories generated by provisional controllers’ actions. Pareto Resource Continuation Certificate (PRCC) admits an action only when a non‑dominated cumulative reservation remains feasible over the remaining horizon. Under a matched GRPO resource envelope, FSPO reaches 66.11% held‑out and 59.43% OOD accuracy, surpassing the strongest adaptive baseline PB2 (64.47% / 57.03%). Across three paired training seeds, gains of +2.42 and +3.19 percentage points over a contextual bandit are observed on held‑out and OOD evaluations. With high behavior‑deployment mismatch, policy‑consistent risk lowers selected‑decision ECE from 0.108 to 0.053; DCTC reduces it from 0.039 to 0.022 at matched acceptance; PRCC eliminates false‑feasible admissions on an 18‑action catalog ($0.197\rightarrow0.000$); and enabling all three components cuts trajectory failure from $0.181$ to $0.083$ in a factorial ablation.
Review: FSPO’s unified risk modeling, cross‑trajectory calibration, and Pareto‑based resource constraints deliver more reliable and efficient control for budget‑constrained LLM RL post‑training, offering interpretable safety guarantees for real‑world deployment.