Neural networks can exploit feature superposition to encode more concepts than the raw dimensionality, yet cross‑feature interference limits the linear accessibility of simultaneously active features.
By casting linear accessibility as a compressed‑sensing problem, we derive high‑probability bounds for fixed support sets under sub‑Gaussian noise. The analysis shows that a dimension $d$ satisfying $d=O_{\varepsilon}(k\log m)$ is sufficient for reliable linear recovery, dramatically improving on previous worst‑case quadratic bounds $O(k^2)$.
We then validate these bounds across a range of system parameters using Gaussian‑tail approximations, finding close agreement between theory and simulation.
The results quantify the geometric constraints of the linear representation hypothesis and provide a concrete framework for assessing sparse autoencoders, compositional generalization, and neural interpretability.
Review