Search arXiv⌕ Search

arXiv · 2610.04264

Adaptive Bregman Alternating Projections for Feasible Gromov-Wasserstein Learning

Abstract

The Gromov-Wasserstein (GW) problem compares structured distributions without requiring a shared feature space or known correspondences, but its nonconvex objective and coupled marginal constraints make computation challenging. Bregman alternating projected gradient (BAPG) uses inexpensive alternating row and column updates, yet its fixed-penalty relaxation leaves a persistent feasibility gap. We propose Adaptive KL-BAPG (A-KL-BAPG), which combines a finite fixed-penalty burn-in with a guarded increasing-penalty phase. At each tail iteration, the method reuses BAPG's alternating updates and backtracks a delayed-power step until a Sinkhorn-inspired projective-diameter safeguard is satisfied. We prove finite termination of the backtracking at each iteration and show that the feasibility gap vanishes asymptotically. We further establish a best-iterate $O(1/\log N)$ bound for the weighted squared corrected residual and, under a support regularity condition, the existence of a stationary accumulation point for the original GW problem. This distinguishes A-KL-BAPG from fixed-penalty BAPG, whose stationarity guarantees are given for the relaxed problem. Experiments show that A-KL-BAPG achieves a favorable balance of accuracy, objective value, feasibility, and stationarity relative to BAPG variants, projection-based methods, and task-specific baselines. For synthetic and real graph alignment problems, it closely matches the accuracy and objective value of fixed-penalty KL-BAPG while reducing the marginal feasibility gap by 62-99% and the projected stationarity residual by 28-98%. Heterogeneous domain adaptation experiments show a similar pattern: A-KL-BAPG maintains comparable target accuracy and objective values while achieving better feasibility and stationarity than fixed-penalty KL-BAPG.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aoran Zhang, César A. Uribe. 2026-10-03. Adaptive Bregman Alternating Projections for Feasible Gromov-Wasserstein Learning. https://arxiv.org/abs/2610.04264

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Cooperative Sheaf Neural Networks

Sheaf diffusion has recently emerged as a promising design pattern for graph representation learning due to its inherent ability to handle heterophilic data and avoid oversmoothing. Meanwhile, cooperative message passing has also been proposed as a way to enhance the flexibility of information diffusion by allowing nodes to independently choose whether to propagate/gather information from/to neighbors. A natural question ensues: is sheaf diffusion capable of exhibiting this cooperative behavior? Here, we provide a negative answer to this question. In particular, we show that existing sheaf diffusion methods fail to achieve cooperative behavior due to the lack of message directionality. To circumvent this limitation, we introduce the notion of cellular sheaves over directed graphs and characterize their in- and out-degree Laplacians. We leverage our construction to propose Cooperative Sheaf Neural Networks (CSNNs). Theoretically, we characterize the receptive field of CSNN and show it allows nodes to selectively attend (listen) to arbitrarily far nodes while ignoring all others in their path, potentially mitigating oversquashing. Our experiments show that CSNN presents overall better performance compared to prior art on sheaf diffusion as well as cooperative graph neural networks.

cs.LG↗

GeoFunFlow: Geometric function flow matching for joint probabilistic inference of physical fields and complex geometries

Inverse problems governed by partial differential equations (PDEs) arise widely in science and engineering, but are often ill-posed and limited by sparse, noisy observations. In many applications, measurements reveal only part of the physical state, while the domain geometry may also be unknown even though it shapes the observed response. Joint field and geometry inference across varying computational domains and discretizations remains challenging, whereas many existing machine learning approaches are designed for known geometries and deterministic field reconstruction. Here, we introduce GeoFunFlow, a probabilistic framework that unifies field reconstruction on known domains and joint field and geometry inference on unknown domains. GeoFunFlow combines a geometric function autoencoder (GeoFAE) with flow matching in the latent space to model a joint distribution over physical fields and geometries. GeoFAE establishes a common representation across spatial discretizations that captures the relationship between physical fields and domain geometries, with unknown geometry represented by a signed distance function. The resulting representation allows observations to guide both field reconstruction and geometry recovery, while latent rectified flow enables efficient conditional sampling and spatially resolved uncertainty quantification. A calibration procedure further provides geometry uncertainty estimates with interpretable empirical coverage. Across seven benchmarks spanning porous media flow, fluid mechanics, and optical tomography, GeoFunFlow accurately recovers fields and geometries across complex, variable, and unknown domains while quantifying spatially resolved conditional uncertainty.

cs.LG↗

Truncated Kernel Stochastic Gradient Descent with General Losses and Spherical Radial Basis Functions

In this paper, we propose a novel kernel stochastic gradient descent (SGD) algorithm for large-scale supervised learning with general losses. Compared to traditional kernel SGD, our algorithm improves efficiency and scalability through an adaptive regularization strategy. By leveraging the infinite series expansion of spherical radial basis functions, this strategy projects the stochastic gradient onto a finite-dimensional hypothesis space, which is adaptively scaled according to the bias-variance trade-off, thereby enhancing generalization performance. To handle the gradient nonlinearity arising from general losses, we develop a new generalization framework combining an inequality-based characterization of the kernel-induced covariance operator with optimization techniques. We prove that both the last iterate and the suffix average converge at minimax-optimal rates, and we further establish optimal strong convergence in the reproducing kernel Hilbert space. Our framework accommodates a broad class of classical loss functions, including least-squares, Huber, and logistic losses. Moreover, the proposed algorithm significantly reduces computational complexity and achieves optimal storage complexity by incorporating coordinate-wise updates from linear SGD, thereby avoiding the costly pairwise operations typical of kernel SGD and enabling efficient processing of streaming data. Finally, extensive numerical experiments provide empirical support for the theoretical results and the computational advantages of our algorithm.

cs.LG↗