Search arXivSearch

arXiv · 2409.07689

Entropy Contractions in Markov Chains: Half-Step, Full-Step and Continuous-Time

Abstract

This paper considers the speed of convergence (mixing) of a finite Markov kernel $P$ with respect to the Kullback-Leibler divergence (entropy). Given a Markov kernel one defines either a discrete-time Markov chain (with the $n$-step transition kernel given by the matrix power $P^n$) or a continuous-time Markov process (with the time-$t$ transition kernel given by $e^{t(P-\mathrm{Id})}$). The contraction of entropy for $n=1$ or $t=0+$ are characterized by the famous functional inequalities, the strong data processing inequality (SDPI) and the modified log-Sobolev inequality (MLSI), respectively. When $P=KK^*$ is written as the product of a kernel and its adjoint, one could also consider the ``half-step'' contraction, which is the SDPI for $K$, while the ``full-step'' contraction refers to the SDPI for $P$. The work [DMLM03] claimed that these contraction coefficients (half-step, full-step, and continuous-time) are generally within a constant factor of each other. We disprove this and related conjectures by working out a number of different counterexamples. In particular, we construct (a) a continuous-time Markov process that contracts arbitrarily faster than its discrete-time counterpart; and (b) a kernel $P$ such that $P^{m+1}$ contracts arbitrarily better than $P^m$. Hence, our main conclusion is that the four standard inequalities comparing five common notions of entropy and variance contraction are generally not improvable. In the process of analyzing the counterexamples, we survey and sharpen the tools for bounding the contraction coefficients and characterize properties of extremizers of the respective functional inequalities. As our examples range from Bernoulli-Laplace model, random walks on graphs, to birth-death chains, the paper is also intended as a tutorial on computing MLSI, SDPI and other constants for these types of commonly occurring Markov chains.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pietro Caputo, Zongchen Chen, Yuzhou Gu, Yury Polyanskiy. 2024-09-12. Entropy Contractions in Markov Chains: Half-Step, Full-Step and Continuous-Time. https://arxiv.org/abs/2409.07689

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The extremal process of a cascading family of branching Brownian motion

We study the asymptotic behaviour of the extremal process of a cascading family of branching Brownian motions. This is a particle system on the real line such that each particle has a type in addition to his position. Particles of type $1$ move on the real line according to Brownian motions and branch at rate $1$ into two children of type $1$. Furthermore, at rate $α$, they give birth to children too of type $2$. Particles of type $2$ move according to standard Brownian motion and branch at rate $1$, but cannot give birth to descendants of type $1$. We obtain the asymptotic behaviour of the extremal process of particles of type $2$.

math.PR

Breuer-Major Theorems for Hilbert Space-Valued Random Variables

Let $\{X_k\}_{k\in\mathbb Z}$ be a stationary Gaussian process with values in a separable Hilbert space $\mathcal H_1$, and let $G:\mathcal H_1\to\mathcal H_2$ be a measurable map into another separable Hilbert space $\mathcal H_2$. We derive a central limit theorem for the centered normalized partial sums of the Hilbert space-valued subordinated process $\{G[X_k]\}_{k\in\mathbb Z}$. Our result holds under either of two sets of sufficient conditions, formulated in terms of the transformation $G$ and the temporal and cross-sectional dependence structure of $\{X_k\}_{k\in\mathbb Z}$. These conditions coincide in finite dimensions but lead to genuinely different phenomena in the infinite-dimensional setting. The proof relies on the recently developed Fourth Moment Theorem on Hilbert spaces, leveraging tools from the infinite-dimensional Malliavin-Stein framework. We also provide continuous-time and quantitative versions of the central limit theorem. In a series of examples, we recover and strengthen limit theorems for a wide array of statistics relevant in functional data analysis, and present, as an application of our result, a novel limit theorem in the framework of neural operators.

math.PR

Controlled rough SDEs, pathwise stochastic control and dynamic programming principles

We study stochastic optimal control of rough stochastic differential equations (RSDEs). This is in the spirit of the pathwise control problem (Lions--Souganidis 1998, Buckdahn--Ma 2007; also Davis--Burstein 1992), with renewed interest and recent works drawing motivation from filtering, SPDEs, and reinforcement learning. Results include regularity of rough value functions, validity of a rough dynamic programming principles and new rough stability results for HJB equations, removing excessive regularity demands previously imposed by flow transformation methods. Measurable selection is used to relate RSDEs to "doubly stochastic" SDEs under conditioning. In contrast to previous works, Brownian statistics for the to-be-conditioned-on noise are not required, aligned with the "pathwise" intuition that these should not matter upon conditioning. Depending on the chosen class of admissible controls, the involved processes may also be anticipating. The resulting stochastic value functions coincide in great generality for different classes of controls. RSDE theory offers a powerful and unified perspective on this problem class.

math.PR