Reinforcement learning for trading faces a fundamental trade‑off between capturing upside and controlling drawdown. Pure profit‑maximizing policies tend to collapse into passive long exposure on upward‑drifting assets, while heavily risk‑penalized rewards become overly defensive during volatile periods, missing opportunities. To address this dilemma, we introduce PPO‑HRAP, a hybrid policy that combines Proximal Policy Optimization (PPO) with an interpretable regime‑aware prior.
Method Overview
- Observation space: includes market features (prices, volatility, etc.) and portfolio‑state variables (positions, cash, etc.).
- Reward function: $$ R = \underbrace{\log\,\frac{V_{t+1}}{V_t}}_{\text{portfolio log return}} - \lambda_{dd}\,\mathbf{1}_{\text{VIX}>\theta}\,\Delta\text{DD} - \lambda_{exp}\,(e_t - e^{*})^2 - \lambda_{turn}\,\text{Turnover}_t $$ where $\Delta\text{DD}$ is the daily drawdown increase, $e_t$ the actual exposure, and $e^{*}$ the target exposure derived from the current volatility regime.
- Action blending: the PPO actor output $a^{\text{PPO}}_t$ is combined with the regime‑derived target exposure $a^{\text{regime}}_t$ using a weight $\alpha$, yielding the executed action $a_t = (1-\alpha) a^{\text{PPO}}_t + \alpha a^{\text{regime}}_t$.
Experimental Results
On the held‑out 2020‑2022 SPY test window, PPO‑HRAP achieves:
- Total return 27.62% (annualized 8.48%)
- Sharpe 0.6447, Sortino 0.8588, Calmar 0.4592
- Maximum drawdown reduced from 34.10% (Buy‑and‑Hold) to 18.47%
Across five random seeds, mean total return $0.2725 \pm 0.0109$ and mean Sharpe $0.6219 \pm 0.0565$, indicating stable performance. Single‑run cross‑asset tests on QQQ and DIA show PPO‑HRAP ranking first in both total return and Sharpe for all three assets.
Limitations and Future Work
While the hybrid approach markedly improves risk‑adjusted returns, it still incurs relatively high turnover, and cross‑asset robustness is supported only by single‑run evidence. Future research should aim to reduce transaction costs and strengthen generalization across different assets.
Review: PPO‑HRAP demonstrates that blending learned actions with a volatility‑aware regime prior can effectively balance profit and risk, offering an interpretable and practical solution for real‑world trading systems. However, addressing turnover and extending robustness beyond SPY remain critical next steps.