NeFut Logo NeFut
中 Admin Login

[CS.AI] PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading

Published at: 2026-10-03 22:00 Last updated: 2026-10-06 12:11
#Machine Learning #optimization #Artificial Intelligence

Reinforcement learning for trading faces a fundamental trade‑off between capturing upside and controlling drawdown. Pure profit‑maximizing policies tend to collapse into passive long exposure on upward‑drifting assets, while heavily risk‑penalized rewards become overly defensive during volatile periods, missing opportunities. To address this dilemma, we introduce PPO‑HRAP, a hybrid policy that combines Proximal Policy Optimization (PPO) with an interpretable regime‑aware prior.

Method Overview

Experimental Results

On the held‑out 2020‑2022 SPY test window, PPO‑HRAP achieves:

Across five random seeds, mean total return $0.2725 \pm 0.0109$ and mean Sharpe $0.6219 \pm 0.0565$, indicating stable performance. Single‑run cross‑asset tests on QQQ and DIA show PPO‑HRAP ranking first in both total return and Sharpe for all three assets.

Limitations and Future Work

While the hybrid approach markedly improves risk‑adjusted returns, it still incurs relatively high turnover, and cross‑asset robustness is supported only by single‑run evidence. Future research should aim to reduce transaction costs and strengthen generalization across different assets.

Review: PPO‑HRAP demonstrates that blending learned actions with a volatility‑aware regime prior can effectively balance profit and risk, offering an interpretable and practical solution for real‑world trading systems. However, addressing turnover and extending robustness beyond SPY remain critical next steps.

Original Source: https://arxiv.org/abs/2610.01325

[h] Back to Home