Search arXivSearch

SEARCH · Search arXiv

Results for “math.IT”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

5,876 records · Page 10Linked to original sources

Accelerating Fourier--Motzkin elimination: redundancy removal and the choice of variable elimination order

Fourier-Motzkin elimination computes an inequality description of the projection of a polyhedron onto a subset of its coordinates by eliminating one variable at a time. It is used in several areas of optimisation and computer science, and it is a standard way of obtaining the entropic constraints of a causal structure, where the marginalisation over the latent variables produces such a projection. Its limitation is the growth of the intermediate systems of inequalities, which can be doubly exponential in the number of eliminated variables even though the projection itself grows only as a single exponential. In practice the computational overload of the method therefore depends on two choices: how the redundant inequalities are removed after each step, and the order in which the variables are eliminated. We consider both. We first show, by an explicit example, that Imbert's redundancy test cannot be interleaved with redundancy removal by linear programming. We show that the two methods, however, can be combined soundly if the derivation records used by Imbert's test are re-initialised after every step at which linear programming is used. We then propose a rule for choosing the elimination order of the variables that gives a significant computational advantage, however, at the cost of increased resource usage. We demonstrate this advantage on some random polytopes, where the rule reduces the running time by factors of between 6 and 25 compared with the same elimination under a fixed order. For entropic descriptions of causal structures, with more than 250 inequalities and more than 100 variables to eliminate, our rule keeps the number of inequalities handled at each step one to two orders of magnitude lower than a fixed order.

cs.CC

Generalized infinite dimensional Alpha-Procrustes based geometries

This work extends the recently introduced Alpha-Procrustes family of Riemannian metrics for symmetric positive definite (SPD) matrices by incorporating generalized versions of the Bures-Wasserstein (GBW), Log-Euclidean, and Wasserstein distances. While the Alpha-Procrustes framework has unified many classical metrics in both finite- and infinite- dimensional settings, it previously lacked the structural components necessary to realize these generalized forms. We introduce a formalism based on unitized Hilbert-Schmidt operators and an extended Mahalanobis norm that allows the construction of robust, infinite-dimensional generalizations of GBW and Log-Hilbert-Schmidt distances. Our approach also incorporates a learnable regularization parameter that enhances geometric stability in high-dimensional comparisons. Preliminary experiments reproducing benchmarks from the literature demonstrate the improved performance of our generalized metrics, particularly in scenarios involving comparisons between datasets of varying dimension and scale. This work lays a theoretical and computational foundation for advancing robust geometric methods in machine learning, statistical inference, and functional data analysis.

stat.ML

Learning Fast Monomial Orders for Gröbner Basis Computations

The efficiency of Gröbner basis computation, the standard engine for solving systems of polynomial equations, depends on the choice of monomial ordering. Despite a near-continuum of possible monomial orders, most implementations rely on static heuristics such as GrevLex, guided primarily by expert intuition. We address this gap by casting the selection of monomial orderings as a reinforcement learning problem over the space of admissible orderings. Our approach leverages domain-informed reward signals that accurately reflect the computational cost of Gröbner basis computations and admits efficient Monte Carlo estimation. Experiments on benchmark problems from systems biology and computer vision show that the resulting learned policies consistently outperform standard heuristics, yielding substantial reductions in computational cost. Moreover, we find that these policies resist distillation into simple interpretable models, providing empirical evidence that deep reinforcement learning allows the agents to exploit non-linear geometric structure beyond the scope of traditional heuristics.

cs.SC

Optimal error estimates for the half-way bounce-back lattice Boltzmann method for the Stokes equations

We give a mathematical proof of the optimal convergence rates for the D2Q9 BGK lattice Boltzmann method with the half-way bounce-back rule for the incompressible Stokes equations in a flat channel. The convergence rates are second-order for the velocity and first-order for the pressure as the lattice spacing $h$ tends to zero, in agreement with formal analyses and numerical experiments, whereas the available rigorous convergence theorems only yield an $O(h^{1/2})$ bound for the velocity error. A key step in the proof is a decomposition of the leading boundary consistency error into macroscopic and kinetic components. These components are absorbed by suitably constructed Stokes and discrete Knudsen layer correctors, respectively. Incorporating these correctors into the prediction function used in previous rigorous analyses, we obtain a refined prediction function with consistency errors of sufficiently high order. Combined with the known weighted $L^2$-stability estimate, this gives the optimal convergence rates.

math.NA

A quaternionic construction behind $841$-point kissing arrangement in ${\mathbb R}^{12}$

Recently, a new record kissing arrangement of $841$ points in $\mathbb R^{12}$ was obtained numerically by optimization (Takhanov-Assylbekov-Yun, 2026). The configuration was released as a coordinate file, without a mathematical description of its structure. The purpose of this paper is to provide such a description. The key observation is that the geometry becomes transparent once we regard $\mathbb R^{12}\cong \mathbb H^3$ as the Cartesian product of three copies of the quaternion algebra. We first introduce a new $840$-point kissing arrangement with a certain quaternionic structure. It consists of three mutually orthogonal regular $24$-cells, supported on the three quaternionic coordinate factors $\mathbb H\times\{0\}\times\{0\}$, $\{0\}\times\mathbb H\times\{0\}$, $\{0\}\times\{0\}\times\mathbb H$, together with two $384$-point families obtained by lifting affine sets of the form $$\{(u,v,w)\in (\mathbb F_2^2)^3\mid u+v+w=η\},$$ to quaternionic triples (whose components belong to the binary octahedral group $2O$) and then applying suitable component-wise rotations and weightings. A characteristic feature of this construction is a pronounced asymmetry among the three quaternionic factors. For the $816$ vectors obtained after removing the third $24$-cell, most of the squared norm is concentrated in the first two quaternionic coordinates, while the third coordinate carries systematically less mass. Thus, the third four-dimensional factor contains more available space than the first two. We then show that this $840$-point configuration provides a natural structural model for the numerical $841$-point record. Finally, we introduce a notion of the general quaternionic construction in dimensions divisible by $4$, and check that record kissing arrangements in ${\mathbb R}^{4k}$, $k\leq 5$, admit a quaternionic construction.

math.RA

Rough variational principles and applications to adjoint systems

We consider a Type-II variational principle driven by geometric rough path with split boundary conditions naturally suited to adjoint systems. From this rough variational principle we derive rough Hamilton's equations, establish their pathwise conservation laws and associated Hamilton--Jacobi equation. We then specialise the framework to rough adjoint systems, obtaining pathwise conservation and quasi-conservation laws that underpin adjoint sensitivity analysis with respect to initial conditions and parameters. On the discrete side, we construct a rough Galerkin discretisation of the rough Type-II variational principle and show that it generates a symplectic flow with discrete analogues of the continuous conservation laws. We establish it's equivalence to a class of Rough Symplectic Partitioned Runge--Kutta (RSPRK) methods and analyse its convergence and naturality properties. Lastly, we perform numerical experiments to validate the predicted convergence rates and demonstrate that RSPRK methods preserve the adjoint conservation laws to machine precision, yielding more accurate and stable gradients in optimisation problems than non-symplectic alternatives.

math.NA

Smoothed Picard Hamiltonian Monte Carlo

We develop a new low-accuracy sampler, called \emph{smoothed Picard Hamiltonian Monte Carlo}, which combines Gaussian smoothing, Picard iteration, and higher-order discretization. For a log-concave target $π\propto \exp(-V)$ in dimension $d$ satisfying $0 \prec αI \preceq \nabla^2 V \preceq βI$, with condition number $κ:= β/α$, smoothed Picard HMC returns a sample with $\sqrt α\,W_2(\cdot,π) \le \varepsilon$ using $\widetilde O(κ^2 + κ^{7/6} d^{1/6}/\varepsilon^{1/3})$ gradient queries. We also prove stronger $W_q$ bounds, and then develop an algorithmic framework, the recursive warm start generator, to upgrade these $W_q$ bounds to stronger divergence guarantees. This produces a warm start for the proximal bouncy particle sampler, introduced in a companion work, leading to a high-accuracy log-concave sampler with complexity $\widetilde O((κ^{7/6} d^{1/6} + κ^{1/2} d^{1/4})\mathrm{polylog}(1/\varepsilon))$.

math.ST

Conformal Uncertainty Quantification Guarantees for Neural Operators

Neural operators provide fast surrogate models for approximating operators between function spaces, but their predictions often lack uncertainty quantification. We develop a split conformal framework to guarantee that a calibrated pointwise band around the neural operator output contains the true solution on at least a $1-γ$ fraction of the evaluation domain, with probability at least $1-α$ over test and calibration inputs, where $α,γ\in(0,1)$. Our method reduces a normalized residual field to its spatial $(1-γ)$-quantile and computes a scaling factor using a held-out calibration dataset. We prove marginal coverage guarantees for measurable residual fields defined on arbitrary probability spaces, covering both continuum domains and fixed discretizations. Under mild assumptions on the data distribution, we show that the coverage conditional on the calibration set follows a Beta distribution, which we verify with numerical experiments on Darcy flow and Navier--Stokes equations, where our calibration yields bands consistently tighter than existing corrections while retaining the target coverage.

math.NA

An empirical study of various candidate selection and partitioning techniques in the DIRECT framework

Over the last three decades, many attempts have been made to improve the DIRECT (DIviding RECTangles) algorithm's efficiency. Various novel ideas and extensions have been suggested. The main two steps of DIRECT-type algorithms are selecting and partitioning potentially optimal rectangles. However, the most efficient combination of these two steps is an area that has not been investigated so far. This paper presents a study covering an extensive examination of various candidate selection and partitioning techniques within the same DIRECT algorithmic framework. Twelve DIRECT-type algorithmic variations are compared on 800 randomly generated GKLS-type test problems and 96 box-constrained global optimization problems from DIRECTGOLib v1.1 with varying complexity. Based on these studies, we have identified the most efficient selection and partitioning combinations leading to new, more efficient, DIRECT-type algorithms. All these algorithms are included in the latest version of DIRECTGO v1.1.0 and are publicly available.

math.OC

Geometric Optics Approximation Sampling: A Reflector-Induced Transport Map Framework

In this paper, we propose Geometric Optics Approximation Sampling (GOAS), a reflector-induced transport-map framework for sampling from target measures. Once a reflecting surface is constructed, the associated transport map is explicitly determined by the physical law of reflection. As a concrete realization, we develop a supporting-hyperellipsoid construction that requires only a discrete approximation of the target measure and does not require gradient information of the target density. The formulation accommodates both density-based and sample-based target representations. A softmin smoothing technique is introduced to obtain a smooth approximate transport map from this piecewise hyperellipsoidal construction. We establish well-posedness and stability of the reflector-induced push-forward measure and derive quantitative error estimates in the maximum mean discrepancy metric, and convergence of continuous statistical observables, including fixed-order moments. Numerical experiments on an analytically tractable example, strongly non-Gaussian targets, sample-based target approximations, and Bayesian inverse problems demonstrate the accuracy and flexibility of GOAS.

math.NA

Sub-Gaussian Concentration and Entropic Normality of the Maximum Likelihood Estimator

It is well known that, under standard regularity conditions, the maximum likelihood estimator (MLE) satisfies a central limit theorem and converges in distribution to a Gaussian random variable as the sample size grows. This paper strengthens this classical result by developing several stronger forms of asymptotic normality for the normalized MLE. With additional assumptions on the score, we first establish sub-Gaussian tail bounds and convergence of all moments for the normalized estimation error. We then prove an entropic central limit theorem for a smoothed version of the estimator, showing convergence in relative entropy to the limiting Gaussian law. When the Fisher information of the normalized estimate is bounded, or its density has bounded first derivative, we further show that the smoothing can be removed, yielding entropic normality of the MLE itself. The proofs develop auxiliary tools that may be of independent interest, including exponential consistency bounds, high-moment estimates, and entropy-control arguments for the estimator.

cs.IT

Comments on the recent improvements of the MRRW bounds

The asymptotic McEliece--Rodemich--Rumsey--Welch bound (1977) limits the largest attainable rate of binary codes as a function of the relative distance. After a nearly half-century hiatus, this result was recently improved in two concurrent works, by OpenAI and by O. Alrabiah and V. Guruswami. The two arguments look entirely different, a Delsarte certificate on the one hand, a classical-quantum channel and the pretty good measurement on the other, and they yield the same bound. The purpose of this note is to explain why: in both proofs, a subspace is attached to every codeword and moved with it, and the bound counts how many such subspaces fit in the ambient space, exactly in the first case and in the probabilistic sense of typicality in the second. We also present the OpenAI proof in the language and context of coding theory, as an extension of the spectral method in which the single vector attached to a codeword is replaced by a subspace.

cs.IT

On the well-posedness and efficient approximation for the classical Melan equation in suspension bridges

The classical Melan equation modeling suspension bridges is considered. We first study the explicit expression and global properties of the analytical solution for the simplified ``less stiff'' model, based on which we derive an a priori estimate for the original classical Melan equation and establish, under an explicit load condition, that it has a unique solution with nonnegative integral under downward live loads, thereby showing the uniqueness of the corresponding deflection curve in the engineering setting. We also develop an efficient iterative approximation method by taking the solution of the simplified ``less stiff'' model as the first iterate, and prove its geometric convergence with explicit error estimates. The applicability and computational efficiency of the method are demonstrated through calculations for two actual bridges, which also quantify the influence of the nonlinear nonlocal term on the solution and clarify the relationship between the simplified and original models. Several engineering observations are verified and explained, and some related open problems are suggested.

math.NA

Entropy lower bounds and sum-product phenomena

Various lower bounds are established for the entropy of sums, products and their combinations. First, we derive a prime-field analogue of a version of the entropy power inequality established by Tao over torsion-free groups. Next, we prove an entropy sum-product statement: For independent and identically distributed random variables $X,X'$, the maximum of ${\bf H}(X+X')$ and ${\bf H}(XX')$ is bounded below by a linear combination of the entropy and the min-entropy (Rényi entropy of order~$\infty$) of $X$. This result, obtained by bounding entropies of the form ${\bf H}\bigl( X(Y+Z)\bigr)$ from above and below, is valid over arbitrary fields $F$. Over $F={\bf R}$, a slightly stronger inequality is derived. Finally, a weak version of a purely Shannon-entropic sum-product result is developed: If the entropic additive doubling of a random variable $X$ over an arbitrary field is $O(1)$, then its multiplicative doubling is at least proportional to ${\bf H}(X)$.

math.CO

Accelerated Primal-Dual Proximal Gradient Splitting Methods for Convex-Concave Saddle-Point Problems

In this paper, based a novel primal-dual dynamical model with adaptive scaling parameters and Bregman divergences, we propose new accelerated primal-dual proximal gradient splitting methods for solving bilinear saddle-point problems with optimal nonergodic convergence rates. For the first, using the spectral analysis, we show that a naive extension of acceleration to a quadratic game is unstable. Motivated by this, we present an accelerated primal-dual gradient flow which combines acceleration with careful velocity correction. To work with non-Euclidean distances, we also equip our continuous model with general Bregman divergences and prove the exponential decay of a Lyapunov function. Then, new primal-dual splitting methods are developed based on proper semi-implicit Euler schemes of the continuous model, and the theoretical convergence rates are nonergodic and optimal with respect to the matrix norms, Lipschitz constants and convexity parameters. Moreover, we introduce efficient restart variants to further improve the proposed methods and provide some numerical results to validate the practical performance.

math.OC

From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis

We inspect the deductive connection between the neural scaling law and Zipf's law -- two statements discussed in machine learning and quantitative linguistics. The neural scaling law describes how the cross entropy rate of a foundation model -- such as a large language model -- changes with respect to the amount of training tokens, parameters, and compute. By contrast, Zipf's law posits that the distribution of tokens exhibits a power law tail. Whereas similar claims have been made in more specific settings, we show that the neural scaling law is a consequence of Zipf's law under certain broad assumptions that we reveal systematically. The derivation steps are as follows: We derive Heaps' law on the vocabulary growth from Zipf's law, Hilberg's hypothesis on the entropy scaling from Heaps' law, and the neural scaling from Hilberg's hypothesis. We illustrate these inference steps by a toy example of the Santa Fe process that satisfies all four statistical laws.

cs.IT

Machine learning of continuous and discrete variational ODEs with convergence guarantee and uncertainty quantification

The article introduces a method to learn dynamical systems that are governed by Euler--Lagrange equations from data. The method is based on Gaussian process regression and identifies continuous or discrete Lagrangians and is, therefore, structure preserving by design. A rigorous proof of convergence as the distance between observation data points converges to zero and lower bounds for convergence rates are provided. Next to convergence guarantees, the method allows for quantification of model uncertainty, which can provide a basis of adaptive sampling techniques. We provide efficient uncertainty quantification of any observable that is linear in the Lagrangian, including of Hamiltonian functions (energy) and symplectic structures, which is of interest in the context of system identification. The article overcomes major practical and theoretical difficulties related to the ill-posedness of the identification task of (discrete) Lagrangians through a careful design of geometric regularisation strategies and through an exploit of a relation to convex minimisation problems in reproducing kernel Hilbert spaces.

math.NA

Further Comments on Yablo's Construction

We continue our analysis of Yablo's coding of the liar paradox by infinite acyclic graphs. The present notes are based on and continue the author's previous results on the problem. In particular, our approach is often more systematic than before.

math.CO