Search arXivSearch

arXiv · 1504.03530

Partially Observable Risk-Sensitive Markov Decision Processes

Abstract

We consider the problem of minimizing a certainty equivalent of the total or discounted cost over a finite and an infinite time horizon which is generated by a Partially Observable Markov Decision Process (POMDP). The certainty equivalent is defined by $U^{-1}(EU(Y))$ where $U$ is an increasing function. In contrast to a risk-neutral decision maker this optimization criterion takes the variability of the cost into account. It contains as a special case the classical risk-sensitive optimization criterion with an exponential utility. We show that this optimization problem can be solved by embedding the problem into a completely observable Markov Decision Process with extended state space and give conditions under which an optimal policy exists. The state space has to be extended by the joint conditional distribution of current unobserved state and accumulated cost. In case of an exponential utility, the problem simplifies considerably and we rediscover what in previous literature has been named information state. However, since we do not use any change of measure techniques here, our approach is simpler. A small numerical example, namely the classical repeated casino game with unknown success probability is considered to illustrate the influence of the certainty equivalent and its parameters.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nicole Bäuerle, Ulrich Rieder. 2016-11-29. Partially Observable Risk-Sensitive Markov Decision Processes. https://doi.org/10.1287/moor.2016.0844

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The extremal process of a cascading family of branching Brownian motion

We study the asymptotic behaviour of the extremal process of a cascading family of branching Brownian motions. This is a particle system on the real line such that each particle has a type in addition to his position. Particles of type $1$ move on the real line according to Brownian motions and branch at rate $1$ into two children of type $1$. Furthermore, at rate $α$, they give birth to children too of type $2$. Particles of type $2$ move according to standard Brownian motion and branch at rate $1$, but cannot give birth to descendants of type $1$. We obtain the asymptotic behaviour of the extremal process of particles of type $2$.

math.PR

Breuer-Major Theorems for Hilbert Space-Valued Random Variables

Let $\{X_k\}_{k\in\mathbb Z}$ be a stationary Gaussian process with values in a separable Hilbert space $\mathcal H_1$, and let $G:\mathcal H_1\to\mathcal H_2$ be a measurable map into another separable Hilbert space $\mathcal H_2$. We derive a central limit theorem for the centered normalized partial sums of the Hilbert space-valued subordinated process $\{G[X_k]\}_{k\in\mathbb Z}$. Our result holds under either of two sets of sufficient conditions, formulated in terms of the transformation $G$ and the temporal and cross-sectional dependence structure of $\{X_k\}_{k\in\mathbb Z}$. These conditions coincide in finite dimensions but lead to genuinely different phenomena in the infinite-dimensional setting. The proof relies on the recently developed Fourth Moment Theorem on Hilbert spaces, leveraging tools from the infinite-dimensional Malliavin-Stein framework. We also provide continuous-time and quantitative versions of the central limit theorem. In a series of examples, we recover and strengthen limit theorems for a wide array of statistics relevant in functional data analysis, and present, as an application of our result, a novel limit theorem in the framework of neural operators.

math.PR

Controlled rough SDEs, pathwise stochastic control and dynamic programming principles

We study stochastic optimal control of rough stochastic differential equations (RSDEs). This is in the spirit of the pathwise control problem (Lions--Souganidis 1998, Buckdahn--Ma 2007; also Davis--Burstein 1992), with renewed interest and recent works drawing motivation from filtering, SPDEs, and reinforcement learning. Results include regularity of rough value functions, validity of a rough dynamic programming principles and new rough stability results for HJB equations, removing excessive regularity demands previously imposed by flow transformation methods. Measurable selection is used to relate RSDEs to "doubly stochastic" SDEs under conditioning. In contrast to previous works, Brownian statistics for the to-be-conditioned-on noise are not required, aligned with the "pathwise" intuition that these should not matter upon conditioning. Depending on the chosen class of admissible controls, the involved processes may also be anticipating. The resulting stochastic value functions coincide in great generality for different classes of controls. RSDE theory offers a powerful and unified perspective on this problem class.

math.PR