arXiv · 2610.11798
Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must
Abstract
Softmax attention, at the heart of Transformers, has demonstrated remarkable capabilities. Yet its underlying mechanisms remain only partially understood. Recent theoretical work studies Gaussian prompts, where the infinite-prompt limit reduces softmax attention to a linear map, but also removes the query-dependent selection that distinguishes it from linear attention. This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies. We show that softmax attention can represent and learn, via gradient-based methods, optimal solutions to a range of statistical tasks, including supervised classification and denoising. Our results highlight two complementary capabilities of softmax attention: it can recover linear tasks as effectively as its simpler linear counterpart, while also exploiting query-dependent context selection to solve nonlinear tasks beyond the reach of linear attention.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Simon Gabet, Etienne Boursier, Claire Boyer. 2026-10-08. Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must. https://arxiv.org/abs/2610.11798
Cite the original work for its findings. Save a collection to share your selection of sources.