arXiv · 2610.05488
Underscoring the Problem: Why Softpick Fails at Initialization
Abstract
Softmax attention gives every token a nonzero weight, which in trained models concentrates into attention sinks and massive activations that widen the dynamic range low-precision inference must cover. Softpick removes this constraint by rectifying scores, eliminating sinks and lowering hidden-state kurtosis, but its advantage fades at scale. We reframe this failure as a normalization problem. Softpick's denominator splits into positive- and negative-shifted sums $D^+$ and $D^-$, used identically in the forward and backward pass, preventing their roles from being isolated. We separate them into a family of operators that independently choose each denominator. The failure originates at initialization: every layer contains rows where $D^+$ is exactly zero, while near-dead rows produce gradient norms above $10^{12}$ regardless of the backward denominator. Only Softpick and a stop-gradient variant, which keeps $D^+ + D^-$ forward but backpropagates through $D^+$ alone, train from scratch. At 230M parameters, the stop-gradient operator matches Softpick on quantization, has fewer dead heads, and retrieves passkeys more reliably, trailing only on peak attention-weight kurtosis.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Aryan Sood, Jaikaran Singh, Ishaan Bansal. 2026-10-04. Underscoring the Problem: Why Softpick Fails at Initialization. https://arxiv.org/abs/2610.05488
Cite the original work for its findings. Save a collection to share your selection of sources.