NeFut Logo NeFut
中 Admin Login

[CS.AI] Self-Evolving Algorithm-Design Agents: Overcoming In-Context Evolutionary Stagnation via Population-Curated Policy Optimization

Published at: 2026-10-01 22:00 Last updated: 2026-10-06 12:11
#algorithm #AI #optimization

Large language models are increasingly acting as algorithm-design agents, tackling complex real-world tasks by creating and refining algorithms. Most successful agents rely on pure in-context evolutionary frameworks, yet they quickly hit performance plateaus in domains that demand specialized knowledge. Parametric adaptation can internalize such knowledge, but conventional training requires abundant domain-specific corpora, while high-quality algorithms are scarce in complex design scenarios.

We propose a sample-efficient parametric self‑evolution approach that lets agents explore and learn from algorithms they generate themselves. First, we formalize in-context evolutionary stagnation and introduce the Improvement Chain proposition, proving that learning a sequence of self‑generated algorithms locally increases the probability of neighboring algorithms. Motivated by this local‑transfer view, we present Population‑Curated Policy Optimization (PCPO), which leverages a global population and a hybrid policy update scheme to retain and reuse high‑quality, diverse self‑generated algorithms, gradually shifting the policy toward stronger solutions.

In the task of learning‑rate schedule design for global placement in electronic design automation, training on only four chip cases, PCPO outperforms state‑of‑the‑art in-context evolutionary methods such as OpenEvolve and ShinkaEvolve across sixteen chip cases on average. With an 8B‑parameter base model, PCPO achieves performance comparable to frontier closed‑source models like GPT‑5.5. PCPO also cuts inference‑time token cost by internalizing grounded domain knowledge and applying prompt distillation.

Moreover, PCPO yields significant speedups on four GPU kernel designs, achieving an average $8.27\times$ improvement over the PyTorch Eager baseline.

Review

Original Source: https://arxiv.org/abs/2609.38757

[h] Back to Home