Search arXivSearch

SEARCH · Search arXiv

Results for “math.ST”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

946 records · Page 7Linked to original sources

Optimal control of fractional diffusion with Dirac measures

We study a PDE-constrained optimization problem for an elliptic equation with the spectral fractional Laplacian and a linear combination of Dirac measures as the forcing term; the controls are the amplitudes of these singular sources. We prove existence and uniqueness of an optimal solution and derive first-order optimality conditions. We then propose a discretization based on finite elements. Since the set of admissible controls is finite dimensional, the control variable itself does not require discretization. We conclude by deriving a priori error bounds

math.OC

The Alexander-Hirschowitz theorem for neurovarieties

We study the dimension and identifiability of neurovarieties associated to polynomial neural networks. We give an independent geometric proof that the linear bounds $d_i\geq 2n_i-1$ on the activation degrees imply non defectiveness for any number of outputs, a dimension statement previously obtained from finite identifiability. The proof is based on a direct analysis of the differential of the parameterization. We also investigate secant and Grassmann-secant obstructions outside this range and prove global identifiability for multi-output architectures under the same degree bounds.

math.AG

Overcoming the spatial order barrier for nonlinear SPDEs with additive space-time white noise

We introduce a fully discrete numerical scheme for semilinear SPDEs with additive space-time white noise that overcomes the previous order barrier for the spatial convergence rate. The scheme achieves a strong convergence rate of $M^{-1+ε}$ in time and $N^{-3/2+ε}$ in space for any $ε>0$, where $M^{-1}$ and $N^{-1}$ are the temporal, respectively the spatial, meshsizes. This substantially improves the standard spatial error bounds of order $N^{-1/2}$ in the literature.

math.NA

On cancellative pairs of families of subsets

A pair $(\mathcal{A}, \mathcal{B})$ of families of subsets of $[n]$ is cancellative if whenever $A, A' \in \mathcal{A}, B \in \mathcal{B}$ satisfy $A \cup B=A' \cup B$, then $A=A'$, and whenever $A \in \mathcal{A}, B, B' \in \mathcal{B}$ satisfy $A \cup B=A \cup B'$, then $B=B'$. We show that for every cancellative pair $(\mathcal{A}, \mathcal{B})$, the inequality $|\mathcal{A}||\mathcal{B}| \le 2.25^n$ holds, matching Tolhuizen's $(2.25-o(1))^n$ lower bound construction.

math.CO

Tight Bounds for Linear and Non-Linear Contraction of Divergences via Duality

We develop a novel framework for bounding the contraction of information divergences, using duality and associated norms in Orlicz spaces. By working in the dual space, we obtain a principled approach to bounding both distribution-dependent strong data-processing inequality (SDPI) constants and \(F_φ\)-curves of divergences. Our bounds are either available in closed form or reducible to one-dimensional convex optimisation problems, in contrast to the infinite-dimensional optimisation problems that characterise SDPIs. These bounds depend on the densities of the reverse kernels with respect to a reference measure. To the best of our knowledge, they are the first universal closed-form bounds on distribution-dependent SDPI constants. We establish tightness for the \(χ^2\)-divergence on several important channel classes, including full-rank binary kernels. We apply our results to several settings. In particular, we derive bounds on the mixing times of Markov chains, including chains with heavy-tailed stationary distributions; obtain improved bounds on burn-in periods for Markov chain Monte Carlo; and strengthen concentration-of-measure bounds for dependent random variables.

cs.IT

TopoAlign: A Framework for Aligning Code to Math via Topological Decomposition

Large Language Models (LLMs) excel at both informal and formal (e.g. Lean 4) mathematical reasoning but still struggle with autoformalisation, the task of transforming informal into formal mathematical statements. Yet, the performance of current Math LLMs is constrained by the scarcity of large-scale corpora, particularly those containing pairs of informal and formal statements. Interestingly, the formal languages used in autoformalisation share structural similarities with programming languages, and code data is available at scale. However, current models trained on code do not transfer effectively to formal math, due to structural and syntactic differences between them. To address this, we propose TopoAlign, a framework that unlocks widely available code repositories as training resources for Math LLMs. TopoAlign decomposes code into docstrings, main functions, and dependency functions, and reassembles these components into analogues that structurally mirror formal statements. We train three state-of-the-art models, DeepSeek-Math, Qwen-3 and Herald, and evaluate them on the MiniF2F, Putnam, and ProofNet benchmarks. TopoAlign provides substantial gains for DeepSeek-Math, improving performance by 17.77% on BEq@10 and 68.82% on typecheck@10, and also measurably improves Herald by 0.12% on BEq@10 and 1.09% on typecheck@10 despite introducing no new mathematical knowledge.

cs.CL

Covering 1024 syndromes with 50 columns

We exhibit a binary linear $[50,40]_2$ code of covering radius $2$, so $\ell_2(10,2)\le 50$, one column below the Kaikkonen--Rosendahl length $51$ that has stood since 2003 and that still seeds the $R=2$ family of Davydov--Marcugini--Pambianco (arXiv:2511.02542). The new matrix admits a $(2,0)$-partition into ten blocks, so Construction $\mathrm{QM}_2^2$ propagates it to exhaustively verified codes of lengths $815$ and $1631$ at $r=18$ and $r=20$, and to the family $n=51\cdot 2^{r/2-5}-1$ of asymptotic density $2601/2048$. The matrices, verifiers, and source are at https://github.com/wustep/maths, pin problems/covering/share/2026-08-24/ at commit 736a38f.

cs.IT

Moments of crosscorrelation demerit factors of binary sequences

Families of sequences with low mutual aperiodic crosscorrelation assist the design of systems for multi-user asynchronous communications and multiple-input multiple-output radar. The crosscorrelation demerit factor of a pair of sequences is the sum of the squared magnitudes of their crosscorrelation values at every shift when the sequences are normalized to unit Euclidean norm, and the merit factor is the reciprocal of the demerit factor. For each positive integer $\ell$, we endow the $2^{2 \ell}$ pairs of binary sequences of length $\ell$ with uniform probability measure and study the distribution of their crosscorrelation demerit factors. Sarwate showed that the mean value is always $1$ regardless of length $\ell$. We develop a method for finding an exact formula for the $p$th central moment (for any positive integer $p$) as a function of $\ell$. Formulae for the variance and third central moment ($p=2$ and $3$) are then obtained by hand calculations, while the fourth through sixth central moments are obtained by computer-assisted calculations. Our theory also shows that all the central moments must be strictly positive for $p\geq 2$ and $\ell \geq 3$.

cs.IT

Distinguishing classes of intersection graphs of homothets or similarities of two convex disks

For smooth convex disks $A$, i.e., convex compact subsets of the plane with non-empty interior and with at most one tangent at every boundary point, we classify the classes $G^{\text{hom}}(A)$ and $G^{\text{sim}}(A)$ of intersection graphs that can be obtained from homothets and similarities of $A$, respectively. Namely, we prove that $G^{\text{hom}}(A)=G^{\text{hom}}(B)$ if and only if $A$ and $B$ are affine equivalent, and $G^{\text{sim}}(A)=G^{\text{sim}}(B)$ if and only if $A$ and $B$ are similar.

cs.CG

A fully globalized solver for discretized inverse elliptic coefficient problems with exact data

We consider finite-dimensional nonlinear inverse problems arising from finite element discretizations of elliptic inverse coefficient problems such as the Calderón problem with finitely many measurements and unknowns. Such inverse coefficient problems are notorious for their nonlinearity and ill-posedness, and numerical solvers tend to depend strongly on good initial values. In this work, we develop a new locally convergent algorithm with an explicit residual criterion that ensures convergence to the inverse problem solution, and a globalized variant that is guaranteed to automatically switch to the faster locally convergent algorithm after finitely many global search steps.

math.NA

Exposing Finite-Depth, Finite-Shot Guarantees for Constrained Quantum Optimization via Fejér Filtering

Constrained quantum optimization algorithms need quantitative guarantees that connect circuit resources to the probability of actually sampling feasible or optimal solutions in finitely many shots. We establish such a connection by exposing a positive sampling law in which mixer-driven exploration and spectral selection can be controlled separately. We show that after removing interference between distinct cost eigenspaces as an analytic device, the measurement distribution becomes the normalized product of a mixer-induced exploration envelope and a Fejér spectral weight, with the former describing how the mixer spreads probability over the encoded manifold and the latter enhancing the target cost phase while suppressing spectrally separated nontarget phases. In this model, finite-shot success becomes a tractable competition between target weight and off-target leakage, yielding an explicit lower bound on the probability of sampling an optimum. For the primary bound, we rescale the cost Hamiltonian to an integer-valued spectrum, placing the wrapped cost phases on a controlled lattice for Fejér filtering. We then define $δ$ as the minimum circular separation between the optimal phase and every nontarget phase. The single-shot success probability $q_0$ satisfies \[ q_0 \ge \frac{x}{1+x}, \qquad x = (p+1)^2 \sin^2\!\left(\fracδ{2}\right) C_β, \] where $p$ is the filter order and $C_β$ is the mixer-envelope mass on the optimal set, exposing a finite-resource compensation law in which weaker phase separation or smaller envelope mass can be compensated by increased filter order and additional shots. The same filtering principle exposes a feasibility guarantee when applied to penalty phases. We further prove analogous bounds for nonlattice spectra through off-target suppression, extending our results beyond exact lattice normalization.

quant-ph

Residual neural networks overcome the curse of dimensionality for semilinear heat equations

Rigorous results show that feedforward neural networks can overcome the curse of dimensionality in the numerical approximation of high-dimensional partial differential equations (PDEs), but comparatively little is known about residual neural networks (ResNets) in the nonlinear PDE setting. We prove that ResNets overcome the curse of dimensionality in the numerical approximation of solutions of semilinear heat equations with globally Lipschitz continuous, gradient-independent nonlinearities: under polynomial growth and network approximability hypotheses on the PDE data, there exist $η\in(0,\infty)$ and ResNets $Ψ_{d,\varepsilon}$, $d\in\mathbb{N}$, $\varepsilon\in(0,1]$, with at most $ηd^η\varepsilon^{-η}$ parameters whose realizations approximate the solution in dimension $d$ with an $L^2$-error of at most $\varepsilon$. The proof represents one deterministic realization of a multilevel Picard estimator by a ResNet whose shortcut connections transmit the spatial variable and a scalar accumulator, while the residual branches successively add the summands of the estimator. For ridge-sum initial conditions, admissible sigmoidal activations, and globally Lipschitz truncations of the nonlinearity, we obtain, for every $ξ>0$, the explicit bound $C_ξd^{4+ξ}\varepsilon^{-(3+ξ)}$ on the number of parameters.

math.NA

A Projected Semiexplicit Integrator for Dissipative Systems with Configuration-Dependent Kinetic Energy: Contact-Herglotz Formulation and Benchmarks

Contact Hamiltonian dynamics gives dissipative mechanics an intrinsic action variable, but explicit contact splittings reach only kinetic energies whose terms are exactly integrable: frozen-coordinate diagonal metrics (the spherical pendulum, a torus particle) are included, while dense metrics with momentum cross terms, with the double pendulum as flagship, are not. We introduce a projected Pihajoki-contact integrator for this non-separable setting, combining phase-space duplication, symmetric projection onto the physical diagonal, and constant-friction damping half-steps, with the action factor carried by an exact Herglotz update. As in the projected extended-phase-space framework it builds on, the construction needs no binding parameter, returns the copies to the diagonal at every step, and confines the nonlinear solve to the $2n$ projection variables. For constant friction the step rescales $ω=dη$ by the exact factor $e^{-γτ}$ when the projection is solved exactly (a classical conformally symplectic identity, realized here for this class), while time-symmetry, consistency, and smoothness yield an $O(τ^3)$ one-step contact-form residual, a bound not specific to the contact form. On the damped double pendulum, spherical pendulum, and torus particle the method is second-order accurate, reproduces the contact decay law, and controls long-time energy and contact drift in coarse or stiff regimes where the Tao baseline and the unprojected average lose the solution. A head-to-head with exact-contactomorphism splittings delimits the niche: where a frozen-coordinate splitting exists it preserves the contact form exactly and wins at matched cost; for the dense double-pendulum metric the realizable alternative is first-order with a prohibitive constant and the projected method prevails. The contact-form estimate is local, one-step, and constant-friction.

math-ph

Transversality Conditions for Boundary Constraints Defined by Differential Equations

What are the transversality conditions for an optimal control problem when the boundary conditions are defined by differential equations? This seemingly bizarre question is motivated by trajectory optimization problems in the $N$-body system. The question, however, is more fundamental and goes beyond problems in astrodynamics to nonintegrable dynamical systems in general. The main contribution of this paper is the development of generic initial- and final-time transversality conditions for optimal control problems whose boundary conditions are defined in terms of differential equations with side conditions. The mathematical definition of differential boundary conditions are part of the foundations developed in this paper. To support the new fundamentals, the concept of coordinated/uncoordinated clock times and weak adjoint covectors are introduced. In the case of uncoordinated clock times, the new transversality conditions reveal that there exists a special situation where a weak adjoint covector is orthogonal to the vector field of the boundary differential equation. This condition is sharply different from the classical statement of orthogonality with respect to the endpoint manifold. The theorems developed in this paper are generic. An application of the theorems to several cases in the three-body problem are described in separate papers.

math.OC

Deciding superellipticity and computing the Weierstrass normal form

Let \( \mathcal{S}_{g,n} \subset \mathcal{M}_g \) be the locus of curves of genus \( g \geq 2 \) admitting a model \( y^n = h(x) \) with \( h \) separable; such curves $C$ have a cyclic group \( C_n \leq \operatorname{Aut}(C) \) of order \( n \) with \( C/C_n \cong \mathbb{P}^1 \). % We give an algorithm which, given an absolutely irreducible plane model \( F(x,y) = 0 \) of a curve \( C \) over a field \( k_0 \) of characteristic zero, decides for which \( n \) the curve lies in \( \mathcal{S}_{g,n} \) and returns a model \( y^n = h(x) \) together with the birational transformation to it.

math.AG

SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics

As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources. However, choosing the right retriever remains difficult, as it is infeasible to directly isolate its effect on downstream performance. On the other hand, existing retrieval-specific benchmarks often fail to capture fine-grained mathematical relevance, penalizing relevant documents. We address this gap by introducing SABER-Math, the first fully automated benchmark for evaluating mathematical IR without expert annotation. Starting from 283K high-school-level math problems with solutions, SABER-Math builds challenging reranking tasks in three steps: (i) first, LLMs extract concise solution summaries and mathematical topics for each problem; (ii) then, per-query relevant documents are discovered using ontology topic-based and lexical solutions-summary-based similarities, and (iii) finally, a Swiss-style LLM preference tournament produces fine-grained relevance ratings for the documents. We evaluate lexical retrievers, specialized mathematical retrieval systems, and recent embedding models. We find that while modern embedding models substantially outperform classical and math-specific baselines, even the strongest systems struggle in symbol-heavy domains like Algebra and Calculus. Importantly, we show that general-purpose IR benchmarks such as MTEB do not reliably predict mathematical performance, especially for recent embedding models, highlighting the need for math-specific retrieval benchmarks.

cs.IR

A proof of Ross's conjecture for two-site moving-target search

A target moves between two sites according to a discrete-time Markov chain with a $2\times2$ transition matrix $M$. At each epoch one site is searched at positive cost, and a search may overlook a target that is present. Ross conjectured that an optimal policy is threshold in the posterior probability that the target is at site~1. MacPhee and Jordan proved the conjecture throughout the nonpositive-determinant ($\det M\le0$) regime and for part of the positive-determinant ($\det M>0$) regime, leaving the remaining cases open. We prove threshold optimality throughout the positive-determinant regime, completing Ross's conjecture for all parameter values.

math.PR

The Ramshaw-Mesina Hybrid Algorithm applied to the Navier Stokes Equations

In 1991, Ramshaw and Mesina proposed a novel synthesis of penalty methods and artificial compression methods. When the two were balanced they found the combination was 3-4 orders more accurate than either alone. This report begins the study of their interesting method applied to the Navier-Stokes equations. We perform stability analysis, semi-discrete error analysis, and tests of the algorithm. Although most of the results for implicit time discretizations of our numerical tests comply with theirs for explicit time discretizations, the behavior in damping pressure oscillations and violations of incompressibility are different from their findings and our heuristic analysis.

math.NA