NeFut Logo NeFut
中 Admin Login

[CS.AI] Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #Artificial Intelligence

Artificial intelligence is advancing rapidly, with systems taking larger roles in reasoning, decision‑making, scientific discovery, and autonomous development. As AI begins to take part in its own improvement—through model training, experience accumulation, agent evolution, and automated AI development—the prospect of recursive self‑improvement (RSI) becomes increasingly relevant. This transition raises a fundamental safety question: how can safety be maintained when the system, its experience, and even the process that produces its successors keep changing?

To address this we introduce the perspective of Evolutionary Safety, which studies safety under persistent and recursive self‑improvement. It concerns not only safety at a single moment but how safety properties change, persist, accumulate, and propagate throughout evolution.

We identify recurring risk manifestations: intent drift, error accumulation, experience contamination, safety‑property erosion, evaluator drift, and risk propagation.

Based on these we build a taxonomy spanning five dimensions: persistent agent state, model state, evaluation and environmental feedback, computational substrate, and meta‑level update mechanisms.

Within this framework we examine how evolutionary risks can be discovered and evaluated across different states, update mechanisms, trajectories, and lineages, and we derive governance principles covering modification, selection, authorization, provenance, and recovery.

Finally we outline open problems aimed at preserving safety guarantees as AI systems become increasingly persistent, adaptive, and recursively self‑improving. Project resources and proposed evaluation systems are publicly available.

Review

Original Source: https://arxiv.org/abs/2609.31186

[h] Back to Home