NeFut Logo NeFut
中 Admin Login

[CS.AI] Reliable Self-Evolution with Imperfect Proxy Rewards

Published at: 2026-10-06 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

Large language model (LLM) based self‑evolving search is regarded as a promising route to scientific discovery. In many domains, evaluating every candidate with high‑fidelity methods is prohibitively expensive, so systems must rely on cheap but imperfect proxy rewards. Such proxies can assign high scores to infeasible candidates, creating false positives that contaminate both the final output and the feedback for subsequent generations. To address this, we introduce Conformal Interval‑Driven Self‑Evolution (CISE), which builds candidate‑specific reward intervals using conditional conformal inference and updates interval widths via iteration‑wise online density‑ratio estimation. Evolutionary feedback uses conservative interval‑based rewards, and a candidate is returned only when all required property intervals lie entirely within their feasible regions. Under explicit independence and covariate‑shift assumptions we derive fixed‑iteration coverage guarantees. We evaluate CISE on three self‑evolving search tasks in materials science; all candidates returned by CISE are true positives under high‑fidelity evaluation, whereas baselines return more candidates but include false positives. These findings highlight the advantage of a smaller, more precise shortlist when downstream validation budgets are limited. The implementation is available at https://github.com/MLAI-Yonsei/CISE.git

Review

Original Source: https://arxiv.org/abs/2610.02975

[h] Back to Home