arXiv · 2609.09556
High-probability guarantees for linear accessibility in feature superposition
Abstract
Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear accessibility of simultaneously active features. By framing linear accessibility as a compressed sensing problem, we derive high-probability bounds for fixed supports under subgaussian noise, proving the sufficient dimension scales linearly ($d=O_{\varepsilon}(k \log m)$) rather than prior worst-case quadratic limits. We characterize the asymmetry between active and inactive interference and the trade-off between interference and observation-noise budgets. We then validate these bounds across system parameters through Gaussian-tail approximations. We also introduce IHT-SAE, which uses learned iterative refinement to improve feature recovery beyond the limits of linear availability. These results quantify the geometric constraints of the linear representation hypothesis, providing a framework for evaluating sparse autoencoders, compositional generalization, and neural interpretability.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Enrico Vompa. 2026-09-17. High-probability guarantees for linear accessibility in feature superposition. https://arxiv.org/abs/2609.09556
Cite the original work for its findings. Save a collection to share your selection of sources.