Search arXivSearch

SEARCH · Search arXiv

Results for “math.AT”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

5,006 records · Page 4Linked to original sources

Nonlinear Dynamics In Optimization Landscape of Shallow Neural Networks with Tunable Leaky ReLU

In this work, we study the nonlinear dynamics of a shallow neural network trained with mean-squared loss and leaky ReLU activation. Under Gaussian inputs and equal layer width k, (1) we establish, based on the equivariant gradient degree, a theoretical framework, applicable to any number of neurons k>= 4, to detect bifurcation of critical points with associated symmetries from global minimum as leaky parameter $α$ varies. Typically, our analysis reveals that a multi-mode degeneracy consistently occurs at the critical number 0, independent of k. (2) As a by-product, we further show that such bifurcations are width-independent, arise only for nonnegative $α$ and that the global minimum undergoes no further symmetry-breaking instability throughout the engineering regime $α$ in range (0,1). An explicit example with k=5 is presented to illustrate the framework and exhibit the resulting bifurcation together with their symmetries.

math.OC

Strong convergence of finite element schemes for the stochastic Landau--Lifshitz--Bloch equation

The dynamics of magnetisation in a bounded ferromagnet in $\mathbb{R}^d$ ($d=1,2$) at high temperatures can be described by the stochastic Landau--Lifshitz--Bloch (sLLB) equation, which is a vector-valued quasilinear stochastic partial differential equation. In this paper, assuming adequate regularity of the initial data, we establish strong convergence in $L^2(Ω)$ of several semi-implicit and implicit fully discrete finite element schemes for the sLLB equation, together with explicit convergence rates. The analysis relies on localised error estimates and new exponential moment bounds for the exact solution. As a by-product, these moment bounds yield mean-square exponential stability of solutions and uniqueness of the invariant measure in one spatial dimension under a small noise assumption. We also sharpen existing convergence-in-probability results for the numerical schemes. Numerical experiments are presented to illustrate and support the theoretical findings.

math.NA

Group-averaged Markov chains II: tuning of group action in finite state space

We study group-averaged Markov chains obtained by augmenting a $π$-stationary kernel $P$ with orbit kernels induced by a group action. We analyse the Gibbs ($G$), Metropolis--Hastings ($M$), and Barker ($B$) kernels, their sandwiches $QPQ$, and mixtures $\tfrac{1}{2}(P+Q)$, where $Q\in\{G,M,B\}$. Under suitable conditions, $M^t$ and $B^t$ converge blockwise to $G$. The projection chains of $GPG$ and $P$ coincide, while every sandwich $QPQ$ has absolute spectral gap no smaller than that of reversible $P$. For $GPG$, we derive an additive asymptotic-variance bound, prove monotonicity for $G$-invariant observables, and identify it as the Kullback--Leibler (KL) information projection of $P$ onto the $G$-invariant kernels. For a fixed orbit partition, the spectral and KL properties of $GPG$ reduce to those of a lower-dimensional orbit-space chain. Among Gibbs projections with a prescribed number of orbits, we identify the partition minimizing KL divergence to stationarity and characterize exact stationarity. Finally, alternating group projections converge at a rate determined by singular values of an overlap matrix and, in structured cases, can yield exact sampling with logarithmically many group actions. These results motivate tuning heuristics and yield polynomial mixing for a Curie--Weiss example in a regime where Glauber dynamics is exponentially slow.

math.PR

Circular Chromatic Numbers, Signability, Relation Algebras, and Network Satisfaction Problems

In this paper, we characterize finite graphs with circular chromatic number less than 3 in terms of the existence of certain signings ($\mathbb Z_2$-labellings studied in the context of signed graphs). In fact, we construct a signed graph which is universal for all such signings -- called anti-triangle-signings in this paper -- of finite $\overline{K_3}$-free graphs, and is closely related to the generic circular triangle-free graph studied by Bodirsky and Guzmán-Pro. Moreover, our universal structure gives rise to a representation of the relation algebra $56_{65}$. We then use this representation to show that the network satisfaction problem described by this relation algebra belongs to NP. This concludes the full classification of the existence of a universal square representation, as well as the complexity of the corresponding network satisfaction problem, for relation algebras with at most four atoms.

math.CO

Exact Limits of Random Projections for Preserving Geometry: Distance Recovery, Nearest-Neighbor Rankings, and Covariance Shape in Gaussian Models

The Johnson-Lindenstrauss (JL) lemma guarantees that a random projection of $n$ points to $m=O(\varepsilon^{-2}\log n)$ dimensions preserves pairwise squared distances within relative error $\varepsilon$ with high probability, and this dimension order is asymptotically optimal. In high dimensions, however, distances concentrate around a baseline while key geometric information lies in much smaller fluctuations. We show that the JL bound can therefore be uninformative about retained geometry: an independent Gaussian replacement map can satisfy it even though the replacement cloud is independent of the original data. We then ask how well any decoder can recover a feature $f(D)$ of a squared distance $D$ from a linear sketch. Under squared-error loss, the optimal decoder is conditional expectation, so recovery defines a linear operator whose singular values quantify feature recovery. For isotropic Gaussian data ($Σ=σ^2 I_d$), we diagonalize this operator in closed form. For fixed $k$ with $m,d-m\to\infty$, its $k$th singular value satisfies $\ell_k\approx(m/ d)^{k/2}$. This yields three sharp consequences. A rank-$m$ sketch retains at most an $m/d$ fraction of the variance of any feature of one squared distance. If $m\to\infty$ and $m/d\to0$, the expected Kendall correlation is $\frac{2}π\sqrt{m/d}(1+o(1))$; for fixed $q$, nearest- neighbor agreement tends to $1/q$. Yet one projection can satisfy the JL bound while mean Kendall correlation vanishes when $\log n\ll m\ll d$. After removing scale, Haar-averaged retained covariance-shape information is $(m/d)^2$. Thus JL distance preservation does not quantify the geometry available for comparison or inference.

cs.LG

A fast and stable test to check if a weakly diagonally dominant matrix is a nonsingular M-matrix

We present a test for determining if a substochastic matrix is convergent. By establishing a duality between weakly chained diagonally dominant (w.c.d.d.) L-matrices and convergent substochastic matrices, we show that this test can be trivially extended to determine whether a weakly diagonally dominant (w.d.d.) matrix is a nonsingular M-matrix. The test's runtime is linear in the order of the input matrix if it is sparse and quadratic if it is dense. This is a partial strengthening of the cubic test in [J. M. Peña., A stable test to check if a matrix is a nonsingular M-matrix, Math. Comp., 247, 1385-1392, 2004]. As a by-product of our analysis, we prove that a nonsingular w.d.d. M-matrix is a w.c.d.d. L-matrix, a fact whose converse has been known since at least 1964. We point out that this strengthens some recent results on M-matrices in the literature.

math.NA

Topology Obstructs Pure Foundation Neural Quantum States

Foundation models for ground states in spin-1/2 systems are a promising method for problems ranging from quantum chemistry to identifying new phase diagrams. Nearly all such models are currently pure-states that condition on the Hamiltonian's parameters, whose Monte Carlo samples give energy estimates according to the variational principle. In this contribution, we show that this representation is topologically obstructed. For any gapped Hamiltonian family whose ground-state bundle is non-trivial, every continuous normalized state-vector model has zero fidelity with the ground state at some parameter value in the Hamiltonian family. For that value, the energy is at least one spectral gap, $Δ$, with an $O(Δ)$ gap in an open-neighbourhood of that point. We show that this is a sufficient no-go also in the case of degenerate ground-state manifolds, time dynamics, and periodic systems with mixed space-time topology, demonstrating these obstructions on one- and two-qubit systems. We discuss how this causes a spike in the fidelity susceptibility, giving a numerical signature of a phase-transition where there is none. We then show that operator-valued models canonically avoid these obstructions and preserve topological information, implying a structural necessity in representation for foundation neural quantum states.

quant-ph

Three Infinite Classes of APN Permutations on $Z_n$

For any permutation of a nontrivial finite abelian group, the differential uniformity is at least two; permutations attaining this bound are called almost perfect nonlinear (APN). We construct three infinite classes of APN permutations on the cyclic group $\mathbb{Z}_n$ using Singer cycles, binomials inducing projective permutations, and completed reciprocals combined with parity and quadratic characters. The respective domain orders are $q+1$ for prime powers $q>2$, $(3^d-1)/2$ for integers $d\ge2$, and $2p$ for primes $p>5$ with $p\equiv5\pmod6$. Each class contains an infinite subclass of composite orders outside the standard forms $r-1$, $r-2$, $r-3$, and $r-4$, where $r$ is a prime power. These forms arise in the Welch--Costas, Panario--Sakzad--Stevens--Wang, and Golomb constructions. To the best of our knowledge, these are the first infinite APN constructions on $\mathbb{Z}_n$ reported since 2011 that yield infinitely many composite orders outside these standard forms.

math.CO

Embracing exchange sequences and oriented matroid polyhedron diameter

We reduce the embracing exchange distance of bases of oriented matroids to the metric of oriented matroid polyhedra. This allows us to disprove recent conjectures of Caoduro, Khodamoradi, Paat, and Shepherd and of Bérczi and Nádor. On the other hand, we show that any two embracing bases of an oriented matroid of rank $r$ can be transformed into each other in at most $2r^{\log_2(r)+3}$ steps and in at most $r$ steps in a graphic oriented matroid or a Lawrence oriented matroid, thus confirming the conjecture in these cases.

math.CO

Bellman--Shoreline Search in Arbitrary Dimension: Exponential Vector Oscillators, Active Memory, Precession, and Effective Computability

We study online search for an unknown affine hyperplane in $\mathbb{R}^D$, for arbitrary fixed finite dimension. Building on a companion self-similar cell reduction and support-function formulation, we ask how the mechanism changes as the normal space grows from $\mathbb{S}^0$ to $\mathbb{S}^{D-1}$. In $D=1$, alternation and productivity yield an equal-ripple principle and the exact stationary constant $9$. In $D=2$, the analogous relative equilibrium is a logarithmic spiral whose bottleneck chord imposes tangency and selects the pitch. For exponential orbits $Γ(σ)=e^{κσ}ω(σ)$, we develop log-directional geometry, exponentially discounted memory, gauges, and recursive hyperspherical parametrizations. Without a shape ansatz, the bottleneck admits a certificate supported by at most $D$ historical suppliers, and at globally worst phases the current point lies on the active face. Within regular chambers we derive exact variation, tangency, pitch, age, and, in $D=3$, delay-system identities. Odd-dimensional obstructions, antipodal subclasses, and harmonic towers provide constraints and explicit candidate families but are not claimed globally optimal. Finally, the N-COMP theorem shows that $C_D^*$ is a computable real for every fixed finite $D$ and that algebraic polygonal $\varepsilon$-optimal cells can in principle be synthesized. Numerical screening through $D=10$ is kept separate from the proved results.

cs.CG

The Value of Depth in Message Passing on Sparse Graphs: A Kesten-Stigum Dichotomy

How deep does a graph neural network need to be on a sparse graph? We study its purest statistical form: node classification on the sparse contextual stochastic block model (CSBM) with average degree $Δ=O(1)$, whose local weak limit is a broadcast-labelled Poisson Galton-Watson tree. Prior work derived a message-passing classifier $h_\ell$ that aggregates from each vertex at distance $k\le\ell$ the attenuated evidence $2\operatorname{artanh}(γ^k t(X_v))$, with $γ$ the edge signal and $t$ a bounded likelihood-ratio transform of the feature. We prove that the value of depth is governed by a single number, the Kesten-Stigum ratio $κ=γ^2Δ$. Below the threshold ($κ<1$), the error sequence is Cauchy at a geometric rate, $|\mathcal{E}(\ell)-\mathcal{E}(\ell')|\le Cκ^{(\ell+1)/3}$ for all $\ell'>\ell$, so all layers beyond depth $O(\log(1/ε))$ change the error by less than $ε$; conversely, under mild regularity each sufficiently deep layer still flips the decision with probability at least $cκ^{\ell/2}$, the empirically sharp exponent. Above the threshold ($κ>1$), depth is geometrically productive: $\mathcal{E}(\ell)$ is driven to a branching-process floor of order at most $1/(κ-1)$ at any geometric rate $κ^{-s\ell}$, $s<1$ (this bound has content only for $κ>17$). No local classifier of any depth beats the universal floor $e^{-Δ}Φ(-ζ)$ set by isolated roots ($ζ$ the feature signal-to-noise ratio), while the first layer provably helps by an explicit total-variation amount. Simulations with an exact belief-propagation baseline on the same trees show that the pairwise rule's error curve is mildly non-monotone in $\ell$, so an optimal finite depth exists (an exact instance is certified in the appendix), while BP saturates strictly faster, at an effective per-layer ratio below $κ$ that we identify.

math.ST

Extremal Asymmetric Depth of Planar Graphs and Hidden Near-Mirror Symmetries of IPR Fullerenes

Although almost all graphs are asymmetric -- having no nontrivial global automorphisms -- they may still possess local symmetries in the form of isomorphisms between induced subgraphs, i.e., partial automorphisms. We study such local symmetries via asymmetric depth, defined in terms of the maximum rank of a nontrivial partial automorphism. We prove a tight upper bound on asymmetric depth in the class of planar graphs and identify the extremal graphs: duals of IPR fullerenes attain the maximum already on $47$ vertices. Our main structural result concerns the IPR fullerenes that are neither maximally asymmetric nor symmetric. In such a cage no purely local action realises a low asymmetric depth, and we show that the map which does realise it cannot be confined to a small part of the cage either: neither to a single face, nor behind an interface of at most $5-k$ edges, $k \le 3$ being the deficiency. A cage of asymmetric depth $2$ or $3$ is therefore not asymmetric in one place; it carries a broken symmetry invisible to its automorphism group. Such cages are rare -- under $2\%$ of the asymmetric IPR fullerenes at $n = 118$. In all $727$ of them the largest partial automorphism is a near-mirror reflection, which we state as an explicit conjecture. We also extend the asymmetric depth bound to graphs of higher genus.

math.CO

Bellman-sufficient Information Complexity

We introduce Bellman-sufficient information complexity for minimax analysis of sequential decision problems. A Bellman-sufficient state retains enough of the history to close the controlled recursion, while an index $Y=χ(Ω)$ specifies the decision-relevant information being charged. The upper bound is a log-penalized Bellman program; the lower bound is a Bellman--Fano comparison along an algorithm-dependent reference trajectory. If the two values match at a common localization scale and the stated admissibility, calibration, and growth conditions hold, they form an information-risk sandwich. UCB, E2D, and AMS/EBO control or relax the upper Bellman bracket in different ways. For the main application, we give a negative answer to a widely studied form of the GP--UCB minimax-optimality question. For every $0<α<1/4$, we construct one bounded continuous kernel whose minimax regret is $Θ(T^{1-α})$ along an infinite sequence of horizons, while two globally calibrated GP--UCB rules incur linear regret under one fixed truth. An epochwise finite-marginal action-index AIR Bellman policy, implemented through robust AIR/AMS/EBO control, attains the minimax order. The construction separates realized information from the cost of uniform optimism: many low-value directions inflate the exploration multiplier and change the trajectory. Through the canonical RKHS feature map, it also yields a finite-horizon polynomial minimax separation for the specified maximal-information-calibrated LinUCB rule. A reproducible experiment illustrates the mechanism.

cs.LG

Assessing Nonlinear Elimination Preconditioning for Trust-Region Phase-Field Fracture

Each quasi-static load step of phase-field fracture is a bound-constrained minimization of a nonconvex, coupled displacement-damage energy under an irreversibility bound on the damage. Monolithic Newton stalls once the nonlinearity localizes at the advancing crack front, and staggered (alternate-minimization) schemes converge slowly there. We present an on-demand nonlinear-elimination preconditioned trust-region Newton method: an energy Steihaug-Toint trust region, a primal-dual active set for irreversibility, and a bound-constrained field-split sweep that eliminates an algebraically-identified "hard set" spanning both fields before each step. The elimination is applied on demand -- triggered by the coupled Newton's own stalling and otherwise skipped -- so the method reduces to monolithic Newton at no surcharge where the step is already healthy. We find the robustness to come from the energy trust region: with that globalization fixed, monolithic Newton already completes every loading history without cutbacks, where residual-merit Newton death-spirals, alternate minimization stalls, and the full-field sweep loses robustness. Against that well-globalized baseline, the on-demand elimination cuts outer nonlinear iterations by 19-25% (brittle) and 17% (ductile), with always-on elimination reaching 26-28% and about $39\%$ at the ductile nucleation step. Measured machine-independently, as a full-mesh-equivalent assembly-work proxy rather than wall-clock, it is competitive with -- not faster than -- monolithic Newton (within about 10%), whereas an always-on sweep adds up to 30%. Nonlinear elimination is thus an iteration-reduction mechanism whose overhead the on-demand gate bounds, with no demonstrated total-work advantage over well-globalized monolithic Newton.

cs.CE

Semi-discrete quadratic Wasserstein energy and state-dependent Langevin exploration

We study the semi-discrete quadratic Wasserstein energy. The energy is nonsmooth at collisions of sites. We prove local Lipschitz continuity on the full configuration space, together with global semiconcavity, coercivity, and dissipativity; show that every global minimizer is interior and collision free; and establish $C^2$ regularity on the collision-free configuration space. The gradient is expressed through the barycenters of the balanced Laguerre cells, while the Hessian is given by an explicit facet formula and satisfies a global one-sided bound. We also solve the one-dimensional problem explicitly in each ordering chamber and give a two-site example on the unit square with non-minimizing Lloyd fixed points. For $d\ge 2$, we then formulate an entropy-regularized relaxed control of the Langevin temperature. The controlled dynamics is strongly well posed, nonexplosive, and collision free. Its value function is a classical interior solution of the exploratory Hamilton-Jacobi-Bellman equation; the Laplacian of the value function is locally $C^1$, which yields a locally Lipschitz optimal temperature feedback. Independently of this optimal-control result, for every fixed Borel temperature rule bounded away from zero, and every sufficiently small step size, the associated Gaussian Euler chain is geometrically ergodic with a full-support invariant law. The raw iterates do not converge, whereas the best-so-far energy converges almost surely to the global minimum and the running record approaches the set of global minimizers.

math.NA

Non-Adaptive 1-Bit Mean Estimation: Minimax Rates and the Sample-Interval Tradeoff

We study distributed one-dimensional mean estimation under a 1-bit communication constraint. Each agent observes one sample, drawn independently from an unknown distribution, and returns a single bit in response to a query $Q: \mathbb{R}\to\{0,1\}$ chosen by a central learner. The distribution has mean in $[-λ,λ]$ and $k$-th central moment at most $σ^k$, for a fixed $k>1$. The order-optimal two-stage protocol of Lau and Scarlett uses responses from the first batch to choose the second-batch queries, motivating the question of whether this single round of interaction is necessary. We answer this negatively: for every $k>1$, a non-adaptive protocol attains the adaptive 1-bit minimax rate (and concurrent works reached the same conclusion via different strategies). We further determine the minimax sample complexity among non-adaptive 1-bit estimators when every one-set $Q^{-1}(1)$ is restricted to a union of at most $s$ intervals. Relative to unrestricted non-adaptive 1-bit querying, this constraint adds a term of order $(λσ/(s\varepsilon^2))\log(1/δ)$, giving the full tradeoff between sample complexity and interval complexity to within $k$-dependent constant factors. As a corollary, we identify, order-wise, the minimum interval budget needed to retain the unrestricted 1-bit minimax sample rate.

stat.ML

Entropy lower bounds and sum-product phenomena

Various lower bounds are established for the entropy of sums, products and their combinations. First, we derive a prime-field analogue of a version of the entropy power inequality established by Tao over torsion-free groups. Next, we prove an entropy sum-product statement: For independent and identically distributed random variables $X,X'$, the maximum of ${\bf H}(X+X')$ and ${\bf H}(XX')$ is bounded below by a linear combination of the entropy and the min-entropy (Rényi entropy of order~$\infty$) of $X$. This result, obtained by bounding entropies of the form ${\bf H}\bigl( X(Y+Z)\bigr)$ from above and below, is valid over arbitrary fields $F$. Over $F={\bf R}$, a slightly stronger inequality is derived. Finally, a weak version of a purely Shannon-entropic sum-product result is developed: If the entropic additive doubling of a random variable $X$ over an arbitrary field is $O(1)$, then its multiplicative doubling is at least proportional to ${\bf H}(X)$.

math.CO

An Improved Bound for Smith's Longest Cycles Conjecture via a Forbidden Subdivision

Smith's conjecture asserts that in every $k$-connected graph with $k\geq 2$, any two longest cycles intersect in at least $k$ vertices. In this work, we establish an $Ω(k^{8/11})$ bound for this conjecture, improving upon the $Ω(k^{2/3})$ bound of Ma and Zhao. Our proof combines a Ramsey theoretic refinement of the traditional Turán-type approach with computer search.

math.CO