Search arXivSearch

arXiv subjects

Mathias Dus

Publications and source records attributed to Mathias Dus.

7 recordsLinked to original sources

Radon Measure Representations for Infinite-Width Neural Networks with Singular Activations

The theoretical foundation of infinite-width shallow neural networks relies heavily on continuous integral representations and Barron spaces. Recently, harmonic analysis-specifically the Radon and Ridgelet transforms-has emerged as a powerful tool to invert these representations and compute the optimal network weights. However, a major analytical bottleneck remains: standard neural network activation functions exhibit severe spectral singularities at the frequency origin. To bypass this divergence, existing frameworks either restrict the theory to specific activation families or mathematically quotient out the network's affine components, which inherently limits their practical scope. In this paper, we overcome these limitations by introducing a purely distributional framework for the generalized Radon transform R $\sigma$ that operates on a broad class of tempered distribution activation functions. By defining a regularized spectrum formulation g($\rho$) = (i$\rho$) $\alpha$ $\sigma$($\rho$), we rigorously absorb the origin singularities without truncating the underlying functional space. Building upon this exact reconstruction, we show that under a definite parity assumption on the activation, there exists an exact linear isometry between Barron functions and their optimal weight measures.

math.FA

Wasserstein Formulation of Reinforcement Learning. An Optimal Transport Perspective on Policy Optimization

We present a geometric framework for Reinforcement Learning (RL) that views policies as maps into the Wasserstein space of action probabilities. First, we define a Riemannian structure induced by stationary distributions, proving its existence in a general context. We then define the tangent space of policies and characterize the geodesics, specifically addressing the measurability of vector fields mapped from the state space to the tangent space of probability measures over the action space. Next, we formulate a general RL optimization problem and construct a gradient flow using Otto's calculus. We compute the gradient and the Hessian of the energy, providing a formal second-order analysis. Finally, we illustrate the method with numerical examples for low-dimensional problems, computing the gradient directly from our theoretical formalism. For high-dimensional problems, we parameterize the policy using a neural network and optimize it based on an ergodic approximation of the cost.

cs.LG

A review of compactness methods for cross-diffusion systems seen as Wasserstein gradient flows

A comprehensive methodology for establishing the existence of gradient flows for cross-diffusion systems with respect to suitable energies is proposed. The approach is based on the construction of piecewise-in-time constant approximations via the Jordan-Kinderlehrer-Otto scheme. Compactness of the approximate sequence is obtained using either the flow interchange technique or the five gradient inequality. These methods are illustrated for both parabolic and hyperbolic-parabolic Busenberg-Travis systems, as well as for several of their variants. This paper reviews the results from the literature and discusses additional properties.

math.AP

Grassmannian Geometry and Global Convergence of Variable Projection for Neural Networks

Training deep neural networks and Physics-Informed Neural Networks (PINNs) often leads to ill-conditioned and stiff optimization problems. A key structural feature of these models is that they are linear in the output-layer parameters and nonlinear in the hiddenlayer parameters, yielding a separable nonlinear least-squares formulation. In this work, we study the classical variable projection (VarPro) method for such problems in the context of deep neural networks. We provide a geometric formulation on the Grassmannian and analyze the structure of critical points and convergence properties of the reduced problem. When the feature map is parametrized by a neural network, we show that these properties persist except in rank-deficient regimes, which we address via a regularized Grassmannian framework. Numerical experiments for regression and PINNs, including an efficient solver for the heat equation, illustrate the practical effectiveness of the approach.

math.OC

Comparison between tensor methods and neural networks in electronic structure calculations

This article compares the tensor method density matrix renormalization group (DMRG) with two neural network based methods -namely FermiNet and PauliNet) for determining the ground state wavefunction of the many-body electronic Schr{\"o}dinger problem. We provide numerical simulations illustrating the main features of the methods and showing convergence with respect to some parameter, such as the rank for DMRG, and number of pretraining iterations for neural networks. We then compare the obtained energy with the methods for a few atoms and molecules, for some of which the exact value of the energy is known for the sake of comparison. In the last part of the article, we propose a new kind of neural network to solve the Schr{\"o}dinger problem based on the training of the wavefunction on a simplex, and an explicit permutation for evaluating the wavefunction on the whole space. We provide numerical results on a toy problem for the sake of illustration.

math.OC

Two-layers neural networks for Schr{\"o}dinger eigenvalue problems

The aim of this article is to analyze numerical schemes using two-layer neural networks withinfinite width for the resolution of high-dimensional Schr{\"o}dinger eigenvalue problems with smoothinteraction potentials and Neumann boundary condition on the unit cube in any dimension. Moreprecisely, any eigenfunction associated to the lowest eigenvalue of the Schr{\"o}dinger operator is a unitL 2 norm minimizer of the associated energy. Using Barron's representation of the solution witha probability measure defined on the set of parameter values and following the approach initiallysuggested by Bach and Chizat [1], the energy is minimized thanks to a constrained gradient curvedynamic on the 2-Wasserstein space of the set of parameter values defining the neural network. Weprove the existence of solutions to this constrained gradient curve. Furthermore, we prove that,if it converges, the represented function is then an eigenfunction of the considered Schr{\"o}dingeroperator. At least up to our knowledge, this is the first work where this type of analysis is carriedout to deal with the minimization of non-convex functionals.

math.AP

Numerical solution of Poisson partial differential equations in high dimension using two-layer neural networks

The aim of this article is to analyze numerical schemes using two-layer neural networks with infinite width for the resolution of the high-dimensional Poisson-Neumann partial differential equations (PDEs) with Neumann boundary conditions. Using Barron's representation of the solution with a measure of probability, the energy is minimized thanks to a gradient curve dynamic on the $2$ Wasserstein space of parameters defining the neural network. Inspired by the work from Bach and Chizat, we prove that if the gradient curve converges, then the represented function is the solution of the elliptic equation considered. Numerical experiments are given to show the potential of the method.

math.NA