Search arXivSearch

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

58 recordsLinked to original sources

Numerical approximation to the invariant measure of McKean-Vlasov stochastic differential equations

Inspired by the stochastic particle method, this paper develops an easily implementable explicit scheme for McKean-Vlasov stochastic differential equations (MV-SDEs) with superlinear growth coefficients. We prove that the numerical solution of the interacting particle system (IPS) attains the optimal uniform-in-time strong convergence rate of order 1/2, and that it faithfully captures the long-term dynamics of MV-SDEs, including moment boundedness, stability, and ergodicity. In particular, the existence and uniqueness of an exchangeable numerical invariant probability measure for the IPS are established via an appropriately constructed operator semigroup. Concerning the approximation of the invariant measure, we derive a non-asymptotic error bound between the distribution of the one-particle numerical solution and the marginal distribution of the IPS's invariant measure; By the uniform-in-time propagation of chaos, we further obtain an asymptotic error bound between the one-particle marginal of the IPS's numerical invariant measure and the exact invariant measure of the MV-SDE. Numerical experiments are provided to validate the theoretical results.

math.PR

Exact affine conditioning beyond Gaussians: a unique characterization of the ensemble Kalman update

The analysis step of the stochastic ensemble Kalman filter, called the ensemble Kalman update (EnKU), is widely used for approximating posterior distributions in inverse problems and data assimilation. The EnKU approximates the posterior distribution $π_{X\mid Y=y_\star}$ by pushing forward the joint distribution $(X,Y)\simπ$ through an affine map $L^{\mathrm{EnKU}}_{π,y_\star}(x,y)$ that depends only on the covariance structure of $π$ and the observation $y_\star$. While the EnKU yields the exact posterior for Gaussian $π$ in the mean-field, this property alone does not uniquely determine the EnKU. In fact, there are infinitely many affine maps $L_{π, y_\star}$ that achieve such exact conditioning. In this paper, we offer a novel characterization of the EnKU among all such affine maps. We first exhaustively characterize the set ${E}^{\mathrm{EnKU}}$ of joint distributions for which the EnKU yields exact conditioning, showing that it is much larger than the set of Gaussians. Next, we show that except for a small class of highly symmetric distributions within ${E}^{\mathrm{EnKU}}$, the EnKU is the {unique} exact affine conditioning map. Further, we characterize the largest possible set of distributions ${F}$ for which a distribution-dependent, weakly observation-dependent, affine map exists, a class of transports that naturally includes the EnKU. We show that ${F}={E}^{\mathrm{EnKU}}\cup{S}_{\mathrm{nl-dec}}$ with a small symmetry class ${S}_{\mathrm{nl-dec}}$, meaning that for affine conditioning beyond the Gaussian setting, the EnKU has an exact set that is essentially maximally large.

math.ST

An Euler scheme for BSDEs via the Wiener chaos decomposition

The Euler scheme is a standard time discretization for BSDEs, but its implementation hinges on approximating conditional expectations and the associated martingale terms at each time step. We propose an implementation based on the Wiener chaos decomposition to approximate these quantities. In contrast to many numerical schemes that rely on a finite-dimensional Markovian representation, our approach accommodates arbitrary $\mathcal{F}_T$-measurable square-integrable terminal conditions. We provide a comprehensive convergence analysis under additional Malliavin regularity assumptions and illustrate the method on several numerical examples, including genuinely non-Markovian problems arising, for instance, in the pricing and hedging of contingent claims under rough-volatility models.

math.NA

Deep belief networks are exact

We prove that every strictly positive probability distribution on \(\{-1,1\}^n\) is represented exactly by a sigmoid belief network with finite parameters. This answers a question of Sutskever and Hinton. The proof upgrades their probability-sharing approximation to exact representation using Brouwer's fixed-point theorem.

cs.AI

Optimal mixing of the systematic scan dynamics via approximate tensorization of entropy

We study the mixing time of the systematic scan dynamics for high-dimensional discrete distributions. This Markov chain updates coordinates sequentially according to a fixed predetermined order, in contrast to the Glauber dynamics that updates coordinates selected uniformly at random. The systematic scan is often favored in practice because it exhibits strong empirical performance, but its theoretical analysis remains far less developed than that of Glauber dynamics. We take a step toward addressing this imbalance by showing that two standard functional notions of weak dependence between the coordinates of the distribution provide strong convergence guarantees for the systematic scan dynamics. First, we show that approximate tensorization of entropy implies optimal $O(\log n)$ mixing time for every scan order under standard marginal, connectivity, and bounded interaction degree assumptions about the distribution. Second, we show that approximate tensorization of variance yields a constant-factor contraction of the variance functional per scan, which in turn implies an optimal $O(1)$ relaxation time for the natural additive and multiplicative reversibilizations of the systematic scan dynamics. Compared with our entropy result, the variance bound improves the dependence on the maximum interaction degree from exponential to quadratic and requires weaker assumptions on the distribution. As concrete applications of our results, we establish optimal $O(\log n)$ mixing of the systematic scan dynamics for bounded-degree antiferromagnetic two-spin systems in the tree-uniqueness region and for the ferromagnetic $q$-state Potts model on square boxes in $\mathbb{Z}^2$ throughout its subcritical regime.

math.PR

Moments of crosscorrelation demerit factors of binary sequences

Families of sequences with low mutual aperiodic crosscorrelation assist the design of systems for multi-user asynchronous communications and multiple-input multiple-output radar. The crosscorrelation demerit factor of a pair of sequences is the sum of the squared magnitudes of their crosscorrelation values at every shift when the sequences are normalized to unit Euclidean norm, and the merit factor is the reciprocal of the demerit factor. For each positive integer $\ell$, we endow the $2^{2 \ell}$ pairs of binary sequences of length $\ell$ with uniform probability measure and study the distribution of their crosscorrelation demerit factors. Sarwate showed that the mean value is always $1$ regardless of length $\ell$. We develop a method for finding an exact formula for the $p$th central moment (for any positive integer $p$) as a function of $\ell$. Formulae for the variance and third central moment ($p=2$ and $3$) are then obtained by hand calculations, while the fourth through sixth central moments are obtained by computer-assisted calculations. Our theory also shows that all the central moments must be strictly positive for $p\geq 2$ and $\ell \geq 3$.

cs.IT

On the Equality of the ELBO to a Sum of Entropies at Stationary Points of Learning

The variational lower bound (a.k.a. ELBO or free energy) is the central objective for many established as well as for many novel algorithms for unsupervised learning. Such algorithms usually increase the bound until parameters have converged to values close to a stationary point of the learning dynamics. Here we show that (for a very large class of generative models) the variational lower bound is at all stationary points of learning equal to a sum of entropies. Concretely, for standard generative models with one set of latents and one set of observed variables, the sum consists of three entropies: (A) the (average) entropy of the variational distributions, (B) the negative entropy of the model's prior distribution, and (C) the (expected) negative entropy of the observable distribution. The obtained result applies under realistic conditions including: finite numbers of data points, at any stationary point (including saddle points) and for any family of (well behaved) variational distributions. The class of generative models for which we show the equality to entropy sums contains many standard as well as novel generative models including standard (Gaussian) variational autoencoders. The prerequisites we use to show equality to entropy sums are relatively mild. Concretely, the distributions defining a given generative model have to be of the exponential family, and the model has to satisfy a parameterization criterion (which is usually fulfilled). Proving equality of the ELBO to entropy sums at stationary points (under the stated conditions) is the main contribution of this work.

stat.ML

Entropy and Distributed Source Coding of Connected Soft Random Geometric Graphs

We consider the distributed compression of Soft Random Geometric Graphs (SRGGs) above the connectivity threshold. We establish the Slepian-Wolf rate region for the SRGG in the setting where there are a finite number of encoders compressing sections of the graph independently. To do so, we prove novel limit theorems and asymptotic equipartition properties for the SRGG and its entropy, which allow us to use random binning techniques for distributed compression.

cs.IT

A proof of Ross's conjecture for two-site moving-target search

A target moves between two sites according to a discrete-time Markov chain with a $2\times2$ transition matrix $M$. At each epoch one site is searched at positive cost, and a search may overlook a target that is present. Ross conjectured that an optimal policy is threshold in the posterior probability that the target is at site~1. MacPhee and Jordan proved the conjecture throughout the nonpositive-determinant ($\det M\le0$) regime and for part of the positive-determinant ($\det M>0$) regime, leaving the remaining cases open. We prove threshold optimality throughout the positive-determinant regime, completing Ross's conjecture for all parameter values.

math.PR

Semi-discrete quadratic Wasserstein energy and state-dependent Langevin exploration

We study the semi-discrete quadratic Wasserstein energy. The energy is nonsmooth at collisions of sites. We prove local Lipschitz continuity on the full configuration space, together with global semiconcavity, coercivity, and dissipativity; show that every global minimizer is interior and collision free; and establish $C^2$ regularity on the collision-free configuration space. The gradient is expressed through the barycenters of the balanced Laguerre cells, while the Hessian is given by an explicit facet formula and satisfies a global one-sided bound. We also solve the one-dimensional problem explicitly in each ordering chamber and give a two-site example on the unit square with non-minimizing Lloyd fixed points. For $d\ge 2$, we then formulate an entropy-regularized relaxed control of the Langevin temperature. The controlled dynamics is strongly well posed, nonexplosive, and collision free. Its value function is a classical interior solution of the exploratory Hamilton-Jacobi-Bellman equation; the Laplacian of the value function is locally $C^1$, which yields a locally Lipschitz optimal temperature feedback. Independently of this optimal-control result, for every fixed Borel temperature rule bounded away from zero, and every sufficiently small step size, the associated Gaussian Euler chain is geometrically ergodic with a full-support invariant law. The raw iterates do not converge, whereas the best-so-far energy converges almost surely to the global minimum and the running record approaches the set of global minimizers.

math.NA

Correlated initialization of deep residual networks

We study the large-depth behavior of residual networks whose weights are correlated across layers at initialization. Our results confirm and extend a conjecture of Marion et al. [2025], according to which correlated initializations should interpolate continuously between the Brownian stochastic differential equation arising from independent initialization and the ordinary differential equation arising from perfectly correlated initialization. When the initialization is obtained from the application of a feature function to a stationary Gaussian sequence with regularly varying correlation, we prove that there exists a unique critical scaling such that the infinite-depth limit is the solution of a Young differential equation driven by a Hermite process. Hermite processes reduce to the fractional Brownian motion if the feature function generating the initialization has Hermite rank one, which is the case for the identity function, for example. We show that the critical scaling and asymptotic limit are uniquely determined by the decay of correlations together with the Hermite rank of the feature function. Consequently, the correlation structure and Hermite rank of the initialization represent meaningful hyperparameters in the asymptotic regime. By contrast, under finite-variance iid initialization, the asymptotic driver is universally Brownian up to normalization regardless of the choice of distribution. Our proofs rely on a collection of novel results establishing a robust stability theory for Young differential equations in Banach spaces.

math.PR

Residual neural networks overcome the curse of dimensionality for semilinear heat equations

Rigorous results show that feedforward neural networks can overcome the curse of dimensionality in the numerical approximation of high-dimensional partial differential equations (PDEs), but comparatively little is known about residual neural networks (ResNets) in the nonlinear PDE setting. We prove that ResNets overcome the curse of dimensionality in the numerical approximation of solutions of semilinear heat equations with globally Lipschitz continuous, gradient-independent nonlinearities: under polynomial growth and network approximability hypotheses on the PDE data, there exist $η\in(0,\infty)$ and ResNets $Ψ_{d,\varepsilon}$, $d\in\mathbb{N}$, $\varepsilon\in(0,1]$, with at most $ηd^η\varepsilon^{-η}$ parameters whose realizations approximate the solution in dimension $d$ with an $L^2$-error of at most $\varepsilon$. The proof represents one deterministic realization of a multilevel Picard estimator by a ResNet whose shortcut connections transmit the spatial variable and a scalar accumulator, while the residual branches successively add the summands of the estimator. For ridge-sum initial conditions, admissible sigmoidal activations, and globally Lipschitz truncations of the nonlinearity, we obtain, for every $ξ>0$, the explicit bound $C_ξd^{4+ξ}\varepsilon^{-(3+ξ)}$ on the number of parameters.

math.NA

Diffuse Gaussian Truncation For Deterministic Approximate Counting

We give deterministic FPTASes for two dense counting problems on which the known deterministic algorithms, based on zero-free interpolation, run in quasipolynomial time. For fixed $0<γ<1/2$ and $0<θ\leq1$, the first approximates $\mathrm{haf}(A)$ for a symmetric matrix $A$ when its support graph $G$ has minimum degree at least $(1/2+γ)n$ and its nonzero entries lie in $[θ,1]$. It also approximates permanents under the analogous bipartite condition, including full-support matrices in $[θ,1]$. For fixed $β>0$ and $0<κ\leq1$, the second approximates the zero-field Ising partition function $Z(J)$ for zero-diagonal real symmetric matrices $J$ satisfying $\max_{i,j}|J_{ij}|\leqβ/n$ and $λ_{\max}(J)\leq1-κ$. No separate lower-eigenvalue condition is imposed. We further prove $\log\mathrm{haf}(A)=h_A(G)-n/2+O_{γ,θ}(1)$ and $Z(J)=2^n\det(I-J)^{-1/2}(1+O_{β,κ}(1/n))$. Here $h_A(G)$ is the maximum weighted fractional-matching entropy. For unweighted graphs, the first formula improves the Cuckler--Kahn error from $o(n)$ to $O_γ(1)$ on the fixed-margin class and extends it to weights in $[θ,1]$. Both algorithms use a common Gaussian truncation principle. Each problem becomes an integral of a product of a fixed entire function over Gaussian coordinates, with possibly indefinite moment matrix entries of order $1/n$. Cancelling the linear term and exactly resumming the quadratic term leaves a coordinate remainder vanishing to order at least three. Complex dilation handles small supports. For large supports, we bound the recombined tail by a large-deviation rate that beats the entropy of the subsets. The truncation error is at most $(CR/n)^{R/2}+e^{-cn}$. This faster-than-geometric decay permits $R\log(en/R)=O(\log n+\log(1/ε))$ and hence polynomial enumeration.

cs.DS

Tight Bounds for Linear and Non-Linear Contraction of Divergences via Duality

We develop a novel framework for bounding the contraction of information divergences, using duality and associated norms in Orlicz spaces. By working in the dual space, we obtain a principled approach to bounding both distribution-dependent strong data-processing inequality (SDPI) constants and \(F_φ\)-curves of divergences. Our bounds are either available in closed form or reducible to one-dimensional convex optimisation problems, in contrast to the infinite-dimensional optimisation problems that characterise SDPIs. These bounds depend on the densities of the reverse kernels with respect to a reference measure. To the best of our knowledge, they are the first universal closed-form bounds on distribution-dependent SDPI constants. We establish tightness for the \(χ^2\)-divergence on several important channel classes, including full-rank binary kernels. We apply our results to several settings. In particular, we derive bounds on the mixing times of Markov chains, including chains with heavy-tailed stationary distributions; obtain improved bounds on burn-in periods for Markov chain Monte Carlo; and strengthen concentration-of-measure bounds for dependent random variables.

cs.IT

Connections between the Föllmer process and the denoising diffusion probabilistic model

The Föllmer process is a Brownian motion conditioned to have a pre-specified distribution at time 1. This process can be interpreted as an ``augmented'' time-compressed version of the reverse stochastic differential equation (SDE) corresponding to the denoising diffusion probabilistic model (DDPM). While this fact has been indirectly used to analyze DDPM sampling errors via discretization of the reverse SDE, the connection between direct discretization of the Föllmer process and the DDPM sampler has not yet been fully explored. This paper clarifies this point while surveying relevant results from the literature. We show that discretized Föllmer processes give natural hyper-parameter settings of the DDPM sampler while accommodating a broader class of variance schedules than discretized reverse SDEs. Moreover, this allows us to systematically recover state-of-the-art results on DDPM sampling error bounds, along with slight improvements.

stat.ML

Shortcomings and capacities of real-constrained neural networks in complex spaces

We find the asymptotic ratio between the storage capacities when enforcing real pre-activations in a complex hypothesis class as opposed to complex ones in the same class. We use weights drawn from the complex Gaussian, which converge asymptotically in norm to the square root of dimension almost surely. Our methods depend on Gardner volume-type comparisons at critical capacity. Our proof relies on an application of the Harish-Chandra-Itzykson-Zuber (HCIZ) formula, nonstandard in literature. With the HCIZ formula, we may obtain a more robust approximation for the final asymptotic ratio. This strategy is applicable to our work specifically since we integrate over the unitary and orthogonal compact manifolds, facilitated via the Weyl integration formula and the Haar measure.

cs.LG

Multi-Turn LLM Conversations under the Least-Recently-Used Policy: Mean-Field Asymptotics and Hit Ratio Approximation

The major workloads in modern large language model (LLM) serving systems have shifted from single-shot LLM calls to multi-turn conversations, where new responses are generated based on the whole conversation history across all previous turns. The hit ratio, i.e., the average fraction of KV caches accessed directly from existing caches stored in high-bandwidth memory (HBM), is hence a crucial metric that governs system performance. Estimating the hit ratio is a highly nontrivial task due to the complex system dynamics, where the KV cache prefixes grow with turns and some must be evicted due to finite memory capacity. We formulate the system as a multi-turn conversation model under the least-recently-used (LRU) policy. Through a mean-field asymptotic framework, we prove that as the conversation arrival rate and the memory capacity grow proportionally to infinity, the hit ratio converges to a closed-form limit. Based on the characterization of the limit, we further propose a practical hit ratio estimator, and validate its accuracy by real LLM serving experiments on the Qwen3-8B model implemented on Ascend NPUs. Our results provide a theoretical foundation for the analysis of multi-turn LLM serving systems and a practical guideline for memory capacity provisioning.

cs.PF

Two Adjoint Perspectives on Fokker-Planck Optimization: A Microscopic-Macroscopic Correspondence

The Fokker-Planck equation admits both a macroscopic Eulerian description through probability densities and a microscopic Lagrangian description through stochastic trajectories. Consequently, optimization problems constrained by the Fokker-Planck equation can be formulated from either perspective. Surprisingly, the corresponding adjoint equations appear to be fundamentally different: the macroscopic adjoint is governed by the backward Kolmogorov equation, whereas the microscopic adjoint evolves pathwise along stochastic trajectories. In this note, we reconcile these two formulations by establishing their correspondence in the continuum setting. We further show that, although their discrete gradients no longer coincide after discretization, both provide consistent numerical approximations of the continuum gradient. Explicit convergence rates are established for both discretization strategies.

math.NA