Search arXivSearch

SEARCH · Search arXiv

Results for “math.IT”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

4,376 records · Page 7Linked to original sources

Representation Costs in Data Science: Foundations and the Quasi-Banach Spaces of Deep Neural Networks

We develop a general framework for analyzing representation costs induced by parameter-space regularizers in data-fitting methods. For an arbitrary parametric method, we define its representation cost and native function space, prove existence, and identify conditions under which parameter-space and function-space problems have equal infimal values and minimizers transfer between them. This framework yields representer theorems and recovers classical formulations---including kernel methods and RKHSs, wavelets and Besov spaces, and shallow neural networks and variation spaces---as special cases. Our main new results concern depth-$L$ feedforward ReLU networks with weight-decay regularization. For these networks, we prove that the representation cost is a power of a quasi-seminorm and that, under suitable hypotheses, the native space is a quasi-Banach space with nonconvex unit ball when $L > 2$. These results identify a novel depth-dependent quasi-Banach geometry induced by weight decay.

math.FA

Discrepancy of geometric incidences

We study the combinatorial (red-blue) discrepancy of finite point sets with respect to hyperplanes and, more generally, bounded-complexity affine algebraic sets. We prove that every $n$-point set in a real Euclidean space admits a red-blue coloring for which every affine algebraic set of dimension at most $D$ and degree at most $k$ has discrepancy at most $n^{\frac12-\frac{1}{2(D+1)}-\varepsilon}$ for some $\varepsilon=\varepsilon(D,k)>0$. This gives a polynomial improvement over the straightforward VC-dimension bound $\tilde O(n^{\frac12-\frac{1}{2(D+1)}})$. In the opposite direction, we construct $n$-point sets in $\mathbb R^d$ whose discrepancy with respect to hyperplanes is $\tildeΩ(n^{\frac12-\frac{1}{d+1}}),$ extending the point-line discrepancy lower bound of Chazelle and Lvov. We present further applications of our methods in communication complexity, concerning separation between randomized communication cost and deterministic communication cost with access to equality oracle.

math.CO

Three Infinite Classes of APN Permutations on $Z_n$

For any permutation of a nontrivial finite abelian group, the differential uniformity is at least two; permutations attaining this bound are called almost perfect nonlinear (APN). We construct three infinite classes of APN permutations on the cyclic group $\mathbb{Z}_n$ using Singer cycles, binomials inducing projective permutations, and completed reciprocals combined with parity and quadratic characters. The respective domain orders are $q+1$ for prime powers $q>2$, $(3^d-1)/2$ for integers $d\ge2$, and $2p$ for primes $p>5$ with $p\equiv5\pmod6$. Each class contains an infinite subclass of composite orders outside the standard forms $r-1$, $r-2$, $r-3$, and $r-4$, where $r$ is a prime power. These forms arise in the Welch--Costas, Panario--Sakzad--Stevens--Wang, and Golomb constructions. To the best of our knowledge, these are the first infinite APN constructions on $\mathbb{Z}_n$ reported since 2011 that yield infinitely many composite orders outside these standard forms.

math.CO

Data-efficient Kernel Methods for Learning Hamiltonian Systems

Hamiltonian dynamics describe a wide range of physical systems. As such, data-driven simulations of Hamiltonian systems are important for many scientific and engineering problems. In this work, we propose kernel-based methods for identifying and forecasting Hamiltonian systems directly from trajectory data. We present two approaches: a 2-step method that reconstructs trajectories before learning the Hamiltonian, and a 1-step method that jointly infers both. Across several benchmark systems, including mass-spring dynamics, a nonlinear pendulum, and the Henon-Heiles system, we demonstrate that our framework achieves accurate, data-efficient predictions and outperforms 2-step kernel-based baselines, particularly in scarce-data regimes, while preserving the Hamiltonian structure. Moreover, we prove a priori error estimates, ensuring reliability of the learned models. We also provide a more general, problem-agnostic numerical framework that goes beyond Hamiltonian systems and can be used for data-driven learning of arbitrary dynamical systems.

math.NA

Generalized infinite dimensional Alpha-Procrustes based geometries

This work extends the recently introduced Alpha-Procrustes family of Riemannian metrics for symmetric positive definite (SPD) matrices by incorporating generalized versions of the Bures-Wasserstein (GBW), Log-Euclidean, and Wasserstein distances. While the Alpha-Procrustes framework has unified many classical metrics in both finite- and infinite- dimensional settings, it previously lacked the structural components necessary to realize these generalized forms. We introduce a formalism based on unitized Hilbert-Schmidt operators and an extended Mahalanobis norm that allows the construction of robust, infinite-dimensional generalizations of GBW and Log-Hilbert-Schmidt distances. Our approach also incorporates a learnable regularization parameter that enhances geometric stability in high-dimensional comparisons. Preliminary experiments reproducing benchmarks from the literature demonstrate the improved performance of our generalized metrics, particularly in scenarios involving comparisons between datasets of varying dimension and scale. This work lays a theoretical and computational foundation for advancing robust geometric methods in machine learning, statistical inference, and functional data analysis.

stat.ML

Learning Fast Monomial Orders for Gröbner Basis Computations

The efficiency of Gröbner basis computation, the standard engine for solving systems of polynomial equations, depends on the choice of monomial ordering. Despite a near-continuum of possible monomial orders, most implementations rely on static heuristics such as GrevLex, guided primarily by expert intuition. We address this gap by casting the selection of monomial orderings as a reinforcement learning problem over the space of admissible orderings. Our approach leverages domain-informed reward signals that accurately reflect the computational cost of Gröbner basis computations and admits efficient Monte Carlo estimation. Experiments on benchmark problems from systems biology and computer vision show that the resulting learned policies consistently outperform standard heuristics, yielding substantial reductions in computational cost. Moreover, we find that these policies resist distillation into simple interpretable models, providing empirical evidence that deep reinforcement learning allows the agents to exploit non-linear geometric structure beyond the scope of traditional heuristics.

cs.SC

Optimal error estimates for the half-way bounce-back lattice Boltzmann method for the Stokes equations

We give a mathematical proof of the optimal convergence rates for the D2Q9 BGK lattice Boltzmann method with the half-way bounce-back rule for the incompressible Stokes equations in a flat channel. The convergence rates are second-order for the velocity and first-order for the pressure as the lattice spacing $h$ tends to zero, in agreement with formal analyses and numerical experiments, whereas the available rigorous convergence theorems only yield an $O(h^{1/2})$ bound for the velocity error. A key step in the proof is a decomposition of the leading boundary consistency error into macroscopic and kinetic components. These components are absorbed by suitably constructed Stokes and discrete Knudsen layer correctors, respectively. Incorporating these correctors into the prediction function used in previous rigorous analyses, we obtain a refined prediction function with consistency errors of sufficiently high order. Combined with the known weighted $L^2$-stability estimate, this gives the optimal convergence rates.

math.NA

Conformal Uncertainty Quantification Guarantees for Neural Operators

Neural operators provide fast surrogate models for approximating operators between function spaces, but their predictions often lack uncertainty quantification. We develop a split conformal framework to guarantee that a calibrated pointwise band around the neural operator output contains the true solution on at least a $1-γ$ fraction of the evaluation domain, with probability at least $1-α$ over test and calibration inputs, where $α,γ\in(0,1)$. Our method reduces a normalized residual field to its spatial $(1-γ)$-quantile and computes a scaling factor using a held-out calibration dataset. We prove marginal coverage guarantees for measurable residual fields defined on arbitrary probability spaces, covering both continuum domains and fixed discretizations. Under mild assumptions on the data distribution, we show that the coverage conditional on the calibration set follows a Beta distribution, which we verify with numerical experiments on Darcy flow and Navier--Stokes equations, where our calibration yields bands consistently tighter than existing corrections while retaining the target coverage.

math.NA

Geometric Optics Approximation Sampling: A Reflector-Induced Transport Map Framework

In this paper, we propose Geometric Optics Approximation Sampling (GOAS), a reflector-induced transport-map framework for sampling from target measures. Once a reflecting surface is constructed, the associated transport map is explicitly determined by the physical law of reflection. As a concrete realization, we develop a supporting-hyperellipsoid construction that requires only a discrete approximation of the target measure and does not require gradient information of the target density. The formulation accommodates both density-based and sample-based target representations. A softmin smoothing technique is introduced to obtain a smooth approximate transport map from this piecewise hyperellipsoidal construction. We establish well-posedness and stability of the reflector-induced push-forward measure and derive quantitative error estimates in the maximum mean discrepancy metric, and convergence of continuous statistical observables, including fixed-order moments. Numerical experiments on an analytically tractable example, strongly non-Gaussian targets, sample-based target approximations, and Bayesian inverse problems demonstrate the accuracy and flexibility of GOAS.

math.NA

Sub-Gaussian Concentration and Entropic Normality of the Maximum Likelihood Estimator

It is well known that, under standard regularity conditions, the maximum likelihood estimator (MLE) satisfies a central limit theorem and converges in distribution to a Gaussian random variable as the sample size grows. This paper strengthens this classical result by developing several stronger forms of asymptotic normality for the normalized MLE. With additional assumptions on the score, we first establish sub-Gaussian tail bounds and convergence of all moments for the normalized estimation error. We then prove an entropic central limit theorem for a smoothed version of the estimator, showing convergence in relative entropy to the limiting Gaussian law. When the Fisher information of the normalized estimate is bounded, or its density has bounded first derivative, we further show that the smoothing can be removed, yielding entropic normality of the MLE itself. The proofs develop auxiliary tools that may be of independent interest, including exponential consistency bounds, high-moment estimates, and entropy-control arguments for the estimator.

cs.IT

Comments on the recent improvements of the MRRW bounds

The asymptotic McEliece--Rodemich--Rumsey--Welch bound (1977) limits the largest attainable rate of binary codes as a function of the relative distance. After a nearly half-century hiatus, this result was recently improved in two concurrent works, by OpenAI and by O. Alrabiah and V. Guruswami. The two arguments look entirely different, a Delsarte certificate on the one hand, a classical-quantum channel and the pretty good measurement on the other, and they yield the same bound. The purpose of this note is to explain why: in both proofs, a subspace is attached to every codeword and moved with it, and the bound counts how many such subspaces fit in the ambient space, exactly in the first case and in the probabilistic sense of typicality in the second. We also present the OpenAI proof in the language and context of coding theory, as an extension of the spectral method in which the single vector attached to a codeword is replaced by a subspace.

cs.IT

Entropy lower bounds and sum-product phenomena

Various lower bounds are established for the entropy of sums, products and their combinations. First, we derive a prime-field analogue of a version of the entropy power inequality established by Tao over torsion-free groups. Next, we prove an entropy sum-product statement: For independent and identically distributed random variables $X,X'$, the maximum of ${\bf H}(X+X')$ and ${\bf H}(XX')$ is bounded below by a linear combination of the entropy and the min-entropy (Rényi entropy of order~$\infty$) of $X$. This result, obtained by bounding entropies of the form ${\bf H}\bigl( X(Y+Z)\bigr)$ from above and below, is valid over arbitrary fields $F$. Over $F={\bf R}$, a slightly stronger inequality is derived. Finally, a weak version of a purely Shannon-entropic sum-product result is developed: If the entropic additive doubling of a random variable $X$ over an arbitrary field is $O(1)$, then its multiplicative doubling is at least proportional to ${\bf H}(X)$.

math.CO

Accelerated Primal-Dual Proximal Gradient Splitting Methods for Convex-Concave Saddle-Point Problems

In this paper, based a novel primal-dual dynamical model with adaptive scaling parameters and Bregman divergences, we propose new accelerated primal-dual proximal gradient splitting methods for solving bilinear saddle-point problems with optimal nonergodic convergence rates. For the first, using the spectral analysis, we show that a naive extension of acceleration to a quadratic game is unstable. Motivated by this, we present an accelerated primal-dual gradient flow which combines acceleration with careful velocity correction. To work with non-Euclidean distances, we also equip our continuous model with general Bregman divergences and prove the exponential decay of a Lyapunov function. Then, new primal-dual splitting methods are developed based on proper semi-implicit Euler schemes of the continuous model, and the theoretical convergence rates are nonergodic and optimal with respect to the matrix norms, Lipschitz constants and convexity parameters. Moreover, we introduce efficient restart variants to further improve the proposed methods and provide some numerical results to validate the practical performance.

math.OC

Machine learning of continuous and discrete variational ODEs with convergence guarantee and uncertainty quantification

The article introduces a method to learn dynamical systems that are governed by Euler--Lagrange equations from data. The method is based on Gaussian process regression and identifies continuous or discrete Lagrangians and is, therefore, structure preserving by design. A rigorous proof of convergence as the distance between observation data points converges to zero and lower bounds for convergence rates are provided. Next to convergence guarantees, the method allows for quantification of model uncertainty, which can provide a basis of adaptive sampling techniques. We provide efficient uncertainty quantification of any observable that is linear in the Lagrangian, including of Hamiltonian functions (energy) and symplectic structures, which is of interest in the context of system identification. The article overcomes major practical and theoretical difficulties related to the ill-posedness of the identification task of (discrete) Lagrangians through a careful design of geometric regularisation strategies and through an exploit of a relation to convex minimisation problems in reproducing kernel Hilbert spaces.

math.NA

Further Comments on Yablo's Construction

We continue our analysis of Yablo's coding of the liar paradox by infinite acyclic graphs. The present notes are based on and continue the author's previous results on the problem. In particular, our approach is often more systematic than before.

math.CO

On two proofs of $d^2$ mixing of weighted Dikin walks

We study the mixing time of weighted Dikin walks for sampling from exponential distributions on polytopes and truncated positive-semidefinite (PSD) cones. Our first result gives a general total-variation mixing bound under strong self-concordance, $\barν$-symmetry, and mixed-trace regularity on the local metric. The key idea is to control the Metropolis--Hastings acceptance probability on a high-probability region rather than at every point. Applying this framework to the Lee--Sidford, Lewis-weight, and John metrics yields an $\widetilde O(d^2)$ mixing bound for sampling from polytopes, while applying it to a hybrid barrier yields an $\widetilde O(d^4)$ mixing bound for sampling from truncated PSD cones. Our second result establishes stronger $χ^2$-divergence guarantees and pointwise acceptance control using a new fourth-order bootstrap condition. For a suitably scaled Lee--Sidford metric, this yields an $\widetilde O(d^2)$ mixing bound in $χ^2$-divergence, improving on the previous $\widetilde O(d^{9/4})$ bound.

cs.DS

Variational Continuation for Double Pendulum Periodic Orbits

We present a Hessian-based approach to numerically continue periodic orbits in dynamical systems. A loop (periodic orbit candidate) is parametrized as a Fourier series; a loss function is defined based on the deviation of the loop from the physical differential equations. Unlike previous work relying on hand-derived Jacobians, our method automates the process by leveraging automatic differentiation, a common machine learning technique. The continuation direction can be determined by the flat directions of the loss landscapes (directions with zero eigenvalues), making the search of periodic orbits efficient and guided. Our method is integrator-free, precisely initializes oscillations around unstable fixed points, and efficiently detects orbit family intersections and subharmonic bifurcations. As a demonstration, we present full continuations of periodic double pendulum oscillations from fixed points, showing bifurcations along orbit families and categorizing branches of periodic orbits. In particular, we find periodic orbits where both pendulum masses are never simultaneously at rest, which to our knowledge has been missing in the literature.

cs.LG