Search arXivSearch

arXiv · 2602.05950

Breaking Symmetry Bottlenecks in GNN Readouts

Abstract

Graph neural networks (GNNs) are widely used for learning on structured data, yet their ability to distinguish non-isomorphic graphs is fundamentally limited. These limitations are usually attributed to message passing; in this work we show that an independent bottleneck arises at the readout stage. Using finite-dimensional representation theory, we prove that all linear permutation-invariant readouts, including sum and mean pooling, factor through the Reynolds (group-averaging) operator and therefore project node embeddings onto the fixed subspace of the permutation action, erasing all non-trivial symmetry-aware components regardless of encoder expressivity. This yields both a new expressivity barrier and an interpretable characterization of what global pooling preserves or destroys. To overcome this collapse, we introduce projector-based invariant readouts that decompose node representations into symmetry-aware channels and summarize them with nonlinear invariant statistics, preserving permutation invariance while retaining information provably invisible to averaging. Empirically, swapping only the readout enables fixed encoders to separate WL-hard graph pairs and improves performance across multiple benchmarks, demonstrating that readout design is a decisive and under-appreciated factor in GNN expressivity.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mouad Talhi, Arne Wolf, Anthea Monod. 2026-02-05. Breaking Symmetry Bottlenecks in GNN Readouts. https://arxiv.org/abs/2602.05950

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Online Regularized Statistical Learning in Reproducing Kernel Hilbert Space With Non-Stationary Data

We study recursive regularized learning algorithms in the reproducing kernel Hilbert space (RKHS) with non-stationary online data streams. We introduce the concept of a random Tikhonov regularization path and decompose the tracking error of the algorithm's output for the regularization path into random difference equations in RKHS. We show that the tracking error vanishes in mean square and almost surely if the regularization path is slowly time-varying. Then, leveraging the monotonicity of inverse operators and the spectral decomposition of compact operators, and introducing the RKHS persistence of excitation condition, we develop a dominated convergence method to prove the mean square and almost sure consistency between the regularization path and the unknown function to be learned. Especially, for independent and non-identically distributed data streams, the mean square and almost sure consistency between the algorithm's output and the unknown function is achieved if the input data's marginal probability measures are slowly time-varying and the average measure over each fixed-length time period is uniformly above a strictly positive finite Borel measure.

cs.LG

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models

Direct Preference Optimization (DPO) improves relative preference by increasing the margin between chosen and rejected responses, but this objective does not specify how probability mass should be redistributed to achieve that margin. Consequently, preference separation can improve even as probability mass collapses away from both preferred and unrelated responses, a phenomenon associated with the squeezing effect. We show that this pathology is localized: continued negative updates become destructive when rejected responses enter low-probability valleys, where further suppression induces harmful redistribution through the coupled softmax geometry. We introduce Gradient-Gated Preference Optimization (Gate-DPO), which uses the current probability geometry of each rejected response to selectively attenuate its contribution, limiting excessive rejected suppression while delaying preference-loss saturation. Gate-DPO is objective-agnostic and composes with existing preference methods. Across four architectures, two preference datasets, and multiple objectives, gating consistently reduces squeezing and improves chosen-response likelihood, while exhibiting robust reductions in rejected-response collapse across architectures and objectives. Mass-dynamics analysis further shows that large preference margins can conceal substantial suppression of unrelated responses, while gating achieves healthier redistribution. Together, our results show that preference margin alone is insufficient to characterize optimization health and that selectively controlling probability dynamics provides a simple, general mechanism for stabilizing preference optimization.

cs.LG

Noise-Debiased Thermodynamic Variance for Local Learning Coefficient Probes

Local learning coefficient (LLC) probes offer a singularity-aware view of neural-network training, but mean-energy methods require a local loss baseline that is ambiguous at transient checkpoints. Thermodynamic variance avoids this input; under mini-batch evaluation, however, direct variance mixes cross-state loss fluctuations with same-state noise. We operationalize this route with the \emph{Shift-Invariant Variance Estimator} (SIVE), which estimates and subtracts the latter component using repeated evaluations. Conditional on any fixed retained path, unclipped SIVE is unbiased for noiseless path variance without requiring MCMC stationarity. The finite-scale diagnostic remains indexed by localization scale $h$---even a locally linear loss has tether-dependent variance---while interpretation as a Real Log Canonical Threshold (RLCT) requires additional stationary low-temperature conditions. Toy experiments recover calibrated finite-scale targets. At the primary localization scale, all five MNIST MLP trajectories exhibit a mid-training trough followed by a rebound in SIVE, while Raw Variance decreases from Epoch 40 to 100 in every trajectory. Across four localization scales, the joint early-drop/late-rise criterion is met in 19 of 20 trajectory--scale pairs. At Epoch 40, the estimated observation-noise correction accounts for $77.5\%$ of Raw Variance. Same-state debiasing thus reveals a reproducible turning structure masked by time-varying observation noise.

cs.LG