Search arXivSearch

arXiv · 1505.00749

A central limit theorem for temporally non-homogenous Markov chains with applications to dynamic programming

Abstract

We prove a central limit theorem for a class of additive processes that arise naturally in the theory of finite horizon Markov decision problems. The main theorem generalizes a classic result of Dobrushin (1956) for temporally non-homogeneous Markov chains, and the principal innovation is that here the summands are permitted to depend on both the current state and a bounded number of future states of the chain. We show through several examples that this added flexibility gives one a direct path to asymptotic normality of the optimal total reward of finite horizon Markov decision problems. The same examples also explain why such results are not easily obtained by alternative Markovian techniques such as enlargement of the state space.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alessandro Arlotto, J. Michael Steele. 2015-12-06. A central limit theorem for temporally non-homogenous Markov chains with applications to dynamic programming. https://doi.org/10.1287/moor.2016.0784

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The extremal process of a cascading family of branching Brownian motion

We study the asymptotic behaviour of the extremal process of a cascading family of branching Brownian motions. This is a particle system on the real line such that each particle has a type in addition to his position. Particles of type $1$ move on the real line according to Brownian motions and branch at rate $1$ into two children of type $1$. Furthermore, at rate $α$, they give birth to children too of type $2$. Particles of type $2$ move according to standard Brownian motion and branch at rate $1$, but cannot give birth to descendants of type $1$. We obtain the asymptotic behaviour of the extremal process of particles of type $2$.

math.PR

Breuer-Major Theorems for Hilbert Space-Valued Random Variables

Let $\{X_k\}_{k\in\mathbb Z}$ be a stationary Gaussian process with values in a separable Hilbert space $\mathcal H_1$, and let $G:\mathcal H_1\to\mathcal H_2$ be a measurable map into another separable Hilbert space $\mathcal H_2$. We derive a central limit theorem for the centered normalized partial sums of the Hilbert space-valued subordinated process $\{G[X_k]\}_{k\in\mathbb Z}$. Our result holds under either of two sets of sufficient conditions, formulated in terms of the transformation $G$ and the temporal and cross-sectional dependence structure of $\{X_k\}_{k\in\mathbb Z}$. These conditions coincide in finite dimensions but lead to genuinely different phenomena in the infinite-dimensional setting. The proof relies on the recently developed Fourth Moment Theorem on Hilbert spaces, leveraging tools from the infinite-dimensional Malliavin-Stein framework. We also provide continuous-time and quantitative versions of the central limit theorem. In a series of examples, we recover and strengthen limit theorems for a wide array of statistics relevant in functional data analysis, and present, as an application of our result, a novel limit theorem in the framework of neural operators.

math.PR

Controlled rough SDEs, pathwise stochastic control and dynamic programming principles

We study stochastic optimal control of rough stochastic differential equations (RSDEs). This is in the spirit of the pathwise control problem (Lions--Souganidis 1998, Buckdahn--Ma 2007; also Davis--Burstein 1992), with renewed interest and recent works drawing motivation from filtering, SPDEs, and reinforcement learning. Results include regularity of rough value functions, validity of a rough dynamic programming principles and new rough stability results for HJB equations, removing excessive regularity demands previously imposed by flow transformation methods. Measurable selection is used to relate RSDEs to "doubly stochastic" SDEs under conditioning. In contrast to previous works, Brownian statistics for the to-be-conditioned-on noise are not required, aligned with the "pathwise" intuition that these should not matter upon conditioning. Depending on the chosen class of admissible controls, the involved processes may also be anticipating. The resulting stochastic value functions coincide in great generality for different classes of controls. RSDE theory offers a powerful and unified perspective on this problem class.

math.PR