The study formulates solar PV policy design as a sequential decision problem and integrates reinforcement learning (RL) with a stochastic agent‑based model (ABM) that simulates yearly adoption under uncertainty. A policymaker agent selects annual incentives—capital grants, subsidised loan rates, and feed‑in tariffs—over a 16‑year horizon. Adoption and cost are balanced through a scalarised reward with weight $w_{\text{cost}}$.\
Policies are learned using PPO, SAC, and TD3 and evaluated in stochastic simulations. The highest‑adoption policy (TD3, $w_{\text{cost}}=0.5$) yields about 4,145 adopters at a cost of €41.73 M; the lowest‑cost policy (PPO, $w_{\text{cost}}=2.0$) reduces expenditure to €7.27 M with 2,682 adopters; the balanced policy (PPO, $w_{\text{cost}}=1.6$) achieves 3,495 adopters at €22.47 M. Consistent trade‑off patterns across algorithms indicate robustness of the adoption‑cost relationship. Compared with static baseline policies, the RL framework explores a broader policy space, demonstrating its potential for adaptive policy design under uncertainty.\
Review