Online POMDP planners typically minimize the expected cumulative cost, which can hide dangerous states when the belief places significant probability on high‑cost regions. Existing risk‑averse approaches embed a static or dynamic Conditional Value‑at‑Risk (CVaR) into the value function, thus capturing trajectory‑level risk, but they suffer two drawbacks: (i) they retain the immediate cost as an expectation over the belief, leaving the risk within the belief unaddressed; (ii) by altering the value function they require bespoke algorithms and cannot reuse standard expectation‑based planners. We instead apply CVaR directly to the immediate cost over the belief at each step while keeping the objective as the expected cumulative return. Consequently the problem retains a standard MDP form and any expectation‑based POMDP planner becomes risk‑sensitive simply by changing the cost computation. The approach inherits finite‑time guarantees for policy evaluation and sparse sampling, with estimation error independent of the risk level. Our main theoretical contribution is a finite‑time bound on the gap between the particle‑belief MDP surrogate and the original POMDP, yielding an end‑to‑end guarantee from the true POMDP value to the algorithmic estimate. In the risk‑neutral limit the formulation collapses to conventional expectation‑based planning.
$$\text{CVaR}_\alpha(X)=\min_{\eta}\left\{\eta+\frac{1}{1-\alpha}\mathbb{E}\big[(X-\eta)_+\big]\right\}$$
Review