NeFut Logo NeFut
中 Admin Login

[CS.AI] Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#algorithm #AI #Machine Learning

Task performance does not dictate which intervention mechanism an agent retains. We define causal retention as the ability of a frozen learned state to answer a mechanism‑probe map that is fixed independently of training, covering dimensions such as action, context, direct target, value, and delay. For finite structural causal model classes, the optimal probe error equals a Bayes decision risk. This risk vanishes exactly when every learning‑interface fiber lies within a single probe‑answer fiber; any post‑processing of that interface inherits the same lower bound. A posterior‑coverage theorem characterizes budgeted retesting, and an exact edit decomposition shows that the shifted set is the unique support of an error‑free target update. Causal Core implements these conditions via evidence‑gated writing, readout filtering, temporal credit, hidden‑context setup, and local diagnostic updates. Experiments span finite causal systems, continuous simulators, the official TD‑MPC2 world model, and Qwen2.5‑7B‑Instruct. A frozen Qwen last‑layer probe reaches 0.958 balanced accuracy on source mechanisms but drops to 0.583 under changed delays; the gated mechanism state attains 1.000 accuracy while accepting only 0.056 of synchronized‑readout candidates. In TD‑MPC2, five target states per actuator raise effect‑sign accuracy from 0.057 to 0.948 without degrading stable responses. Hence, causal retention is distinct from task sufficiency and source‑domain decodability.

Review

Original Source: https://arxiv.org/abs/2609.30650

[h] Back to Home