Search arXivSearch

arXiv · 2609.08615

Why shared attention vectors fail: a case for outcome-indexed tuning

Abstract

Dimensional attention in learning is often implemented as a globally shared attention vector, where each stimulus dimension corresponds to a single scalar. These scalars are learned by models through gradient-descent on error, where predictive features acquire more salience. We show that under multi-outcome learning, where models predict more than one outcome, this shared vector becomes unstable; it collapses to its bounds and prevents the models from learning meaningful attentional tunings for learning and generalization. We address this by introducing an outcome-indexed attentional matrix that converts globally shared attentional tuning into an outcome-indexed representation. We present an analysis of the unstable shared vectors and derive the conditions under which it holds. Empirically, three synthetic experiments benchmark the proposed attention matrices and show that they converge to meaningful representations, something shared attention vectors fail to do. These results suggest that outcome-indexed attentional matrices are a general fix for gradient-based attentional processes, which improves models of learning under multi-outcome conditions.

Explore related subjects

Keep this discovery

BibTeXRIS

Lenard Dome. 2026-09-08. Why shared attention vectors fail: a case for outcome-indexed tuning. https://arxiv.org/abs/2609.08615

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Perforated Backpropagation: A Neuroscience Inspired Extension to Artificial Neural Networks

The neurons of artificial neural networks were originally invented when much less was known about biological neurons than is known today. Our work explores a modification to the core neuron unit to make it more parallel to a biological neuron. The modification is made with the knowledge that biological dendrites are not simply passive activation funnels, but also compute complex non-linear functions as they transmit activation to the cell body. The paper explores a novel system of ``perforated'' backpropagation empowering the artificial neurons of deep neural networks to achieve better performance coding for the same features they coded for in the original architecture. After an initial network training phase, additional ``dendrite'' nodes are added to the network and separately trained with a different objective: to correlate their output with the remaining error of the original neurons. The trained dendrites are then frozen, and the original neurons are further trained, now taking into account the additional error signals provided by the dendrites. The cycle of training the original neurons and then adding and training dendrites can be repeated several times until satisfactory performance is achieved. Our algorithm was successfully added to modern state-of-the-art PyTorch networks across multiple domains, improving upon original accuracies and allowing for significant model compression without a loss in accuracy.

cs.NE

Geometric Reliability of Neural Population Codes: Sampling Calibration and Within-Session Nonstationarity

Trial-to-trial variability limits how reliably neural population geometry can be estimated, while comparisons across populations depend on neuron and trial counts, response quality, and clustered sampling. We quantified within-session geometric reliability using Shesha, the Spearman correlation between representational dissimilarity matrices estimated from independent trial subsets, in all 39 Steinmetz Neuropixels sessions and in olfactory bulb and piriform cortex recordings from Bolding and Franks. Steinmetz analyses matched neurons and repetitions, compared observed reliability with a stationary residual-bootstrap expectation, and used mouse-level or mouse-clustered inference. Mean matched reliability was 0.0402 across 312 area-by-session recordings. Regional differences and reliability above the stationary benchmark did not survive correction. Temporal effects received the strongest support: interleaving early and late trials increased reliability relative to blocked allocation ($Δ=0.02666$, $q=0.001953$), and RDM similarity declined with within-session lag (mean mouse-level slope $=-0.01912$, $q=0.001953$; $n=10$ mice). Outer-cross-fitted reliability was not associated with choice-direction coupling or stimulus or response-direction decoding after correction. Olfactory comparisons remained descriptive because few paired sessions and no animal identities were available. In held-out simulations, associative recurrence outperformed feedforward subspace denoising but not divisive normalization. Representational geometry became less reproducible with temporal separation within a session, and comparisons across neural populations require sampling calibration and independent inference.

q-bio.NC

Formation of structural attractors in neuromorphic systems

This paper examines the theory of Invariant Structural Learning (ISL), which proposes a non-optimization approach to concept formation. Learning is interpreted as convergence to structural attractors in a hypergraph space, rather than as the minimization of a global loss function. The paper presents the ISL model, including its mathematical formalization, computational verification, and a hypothetical neurobiological interpretation. The mathematical section introduces the formal apparatus of the structural reduction process and proves its finite convergence, the existence and uniqueness of class structural attractors, and the self-organization of attractor maps. The computational section demonstrates the feasibility of the proposed approach on classical image recognition tasks, utilizing the proposed learning mechanism without backpropagation and with extremely small training datasets. Finally, the neurobiological section formulates hypotheses regarding the possible implementation of structural attractors in dendritic trees, neural coding as a projection of internal attractor dynamics, and the development of neural architectures supporting the proposed learning concept. These hypotheses are discussed in the context of modern experimental data in the fields of dendritic computations, synaptic plasticity, and the structural organization of neural circuits. The proposed neurobiological mechanisms are presented as testable hypotheses rather than established biological facts. The results demonstrate the mathematical consistency and computational feasibility of the proposed model, while the neurobiological hypotheses outline potential directions for its experimental verification.

cs.AI