NeFut Logo NeFut
中 Admin Login

[CS.AI] Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training

Published at: 2026-09-30 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Sparse‑autoencoder decomposability and concentrated feature attribution are often taken as signs that a model's computation is easier to reverse‑engineer. This work directly tests whether such representational simplicity truly corresponds to a smaller or more tractable causal circuit.

Starting from the same pretrained GPT‑2 Small checkpoint, we apply matched standard continual training and adversarial continual training. Both conditions must retain competence on the indirect object identification (IOI) task and pass independent robustness verification before any mechanistic comparison.

We evaluate three complementary axes: SAE decomposability, the number of SAE features engaged in task attribution, and the size of faithful circuits recovered from the raw computational graph. Circuit size is measured as the number of causal edges needed to reproduce model behavior at a fixed faithfulness level.

Results show that the robust model is more SAE‑decomposable and uses fewer SAE features for task attribution. Circuit size is regime‑dependent: on competence‑matched IOI, the standard model leads or ties below 85% faithfulness, but at higher thresholds (90% and 95%) the robust model requires substantially fewer edges. This pattern holds for the primary pair, generalizes across a seven‑point faithfulness sweep, and replicates on a second corpus.

The conclusion is that representational or attributional simplicity does not automatically translate into overall circuit simplicity; the advantage emerges only at high faithfulness thresholds. To our knowledge, this is the first controlled empirical test of whether representational cleanliness yields causal simplicity at the circuit level.

Review

Original Source: https://arxiv.org/abs/2609.35890

[h] Back to Home