Conservative bandits must improve an incumbent policy while respecting a prescribed performance budget. When the incumbent is uncertain, the usual approach bounds candidate and baseline rewards separately, which double‑charges the shared estimation error. We introduce Reserve‑C4B, which builds certificates directly on the baseline‑relative contrast. A shared confidence set yields an exact expression for the avoidable penalty and provides a tighter admissibility test for any fixed history. A reserve ledger separates statistical evidence from the allowed performance deficit, and a prefix‑refresh extension re‑certifies accumulated decisions under the current confidence set without discarding previously earned credit. For linear rewards, self‑normalized confidence sets give simultaneous validity over time and for adaptively generated candidates, ensuring that the resulting policy satisfies a conditional‑mean performance constraint with high probability. Reproducible experiments isolate the effects of certificate coupling, prefix refresh, and historical information, showing large reductions in baseline fallback while exposing the limitations of frozen certificates.
Review