Search arXivSearch

arXiv subjects

Erhan Bayraktar

Publications and source records attributed to Erhan Bayraktar.

At least 19 recordsLinked to original sources

Convex order and preservation of convexity for Bayesian posterior updates

We study how the response of a Bayesian posterior statistic to future observations changes as information accumulates. For a non-decreasing function $T$, define $\Pi_n^T=\E[T(\Theta)\vert \mathcal F_n]$, where $\Theta$ has an arbitrary prior and the observations come from a one-parameter exponential family. Conditioning on the same current value of $\Pi^T$, we show that the posterior statistic after additional observations is larger in convex order when the current posterior is based on fewer observations. We also prove preservation of convexity: the expected value of a convex function of the future posterior statistic is convex in the current posterior statistic. Together, these two properties provide structural tools for establishing time-monotonicity results in dynamic Bayesian decision and optimal stopping problems. If the exponential family contains an infinitely divisible distribution, the results extend to a continuous-time observation model through a family of L\'evy processes.

math.ST

Mean-field optimal stopping with endogenous quantile cutoffs

We study a mean-field optimal stopping problem with an endogenous population-level shutdown. All remaining agents stop when the survival mass falls below a prescribed threshold. We recast the discontinuous objective as the singular, nonconvex constraint that the survival mass lie in $\{0\}\cup[\alpha,1]$. We prove the equivalence of strong and weak values via an approximation and the existence of an optimal rule via compactness and penalization. We also prove a dynamic programming principle. The value is continuous away from the critical boundary but may be discontinuous at the boundary itself. Under strict initial feasibility, finite-population values converge to the mean-field value. In the same regime, the laws of near-optimal empirical measures are tight and every mean-field optimizer admits a recovery sequence. At the threshold, however, finite-population convergence may fail.

math.OC

Kullback-Leibler Mirror-Prox for Measure-Valued Variational Inequalities and Mean-Field Equilibria

We study the computation of static mean-field equilibria on a compact state space by formulating the equilibrium condition as a variational inequality over probability measures. We propose an entropic variant of Korpelevich's extragradient algorithm---the Kullback--Leibler Mirror-Prox method---in which Euclidean projections are replaced by relative-entropy proximal steps. Each half-step is therefore an explicit exponential reweighting of the current measure, implemented on a finite state-space discretization. Under Lasry--Lions monotonicity and continuity assumptions, we prove convergence of mesh-refined ergodic averages and obtain finite-iteration Minty-residual and approximate-equilibrium bounds that jointly quantify iteration and discretization errors. Under strong monotonicity, we derive metric convergence rates for the last, best, and averaged iterates. We also develop a KL-type Tikhonov regularization that selects the equilibrium minimizing relative entropy with respect to a reference measure. The framework applies to potential and nonpotential cost operators and does not require differentiability or convexity of the cost in the individual state.

math.OC

Mean-Field Doubly Reflected Forward-Backward SDEs with Optional Barriers and $L^p$-Data

We study mean-field doubly reflected forward-backward stochastic differential equations with two optional barriers satisfying a strong Mokobodzki condition. For $L^p$-data, $p\in(1,2]$, we prove existence and uniqueness on sufficiently short time horizons when the coefficients may depend on the joint law of $(X,Y,Z)$. Under an additional monotonicity condition and using an exponentially weighted norm, we also obtain a global-in-time result for $p=2$. The setting is motivated by recursive mean-field Dynkin games and game-option valuation with irregular payoff barriers.

math.PR

Quantitative Particle Approximation for Controlled Nonlinear Filtering

We estimate convergence rates of value functions for particle approximations of a controlled nonlinear filtering problem. The state is a McKean--Vlasov diffusion on the flat torus, driven by hidden idiosyncratic noise and observed common noise. The filter---the conditional law of the state given the observations---serves as the state variable of the control problem, and the associated value function solves a second-order Hamilton--Jacobi--Bellman equation on the Wasserstein space. We approximate this problem by a centralized \(N\)-particle control problem with independent idiosyncratic noises and a common observation noise. The framework accommodates nonseparable rewards and controlled drifts. Since a single control is applied to the entire population, the Hamiltonian is defined by an optimization performed after integration over the population. Under smoothness of the data, uniform ellipticity, and regularity of this Hamiltonian, we establish uniform value-function error bounds of order \(N^{-1/6}\) for \(d=1\), \(N^{-1/6}(\log N)^{1/3}\) for \(d=2\), and \(N^{-1/(3d)}\) for \(d>2\). The proof combines a translation lift in the common-noise direction, Fourier--Wasserstein inf- and sup-convolutions, viscosity comparison, and particle derivative estimates uniform in \(N\).

math.OC

Long-time behavior and turnpike properties of linear-quadratic graphon mean field control problems

We investigate the asymptotic behavior and turnpike properties of graphon mean field control (GMFC) problems in the linear-quadratic setting. We consider both a finite-horizon GMFC problem and its associated ergodic counterpart, in which the controlled dynamics are governed by a graphon mean field stochastic differential equation with heterogeneous interactions. The optimal controls and state trajectories for both problems are characterized by systems of Riccati equations together with systems of generalized differential and algebraic equations on suitable Hilbert spaces. Under a stabilizability condition and appropriate positivity assumptions on the graphon-induced operators, we establish the unique solvability of the ergodic control problem and derive exponential convergence estimates for the finite-horizon system to its stationary limit. As a consequence, we establish an exponential turnpike property for the optimal pair and prove the convergence of the time-averaged value function for the finite-horizon GMFC problem.

math.OC

Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data

This paper develops a model-free framework for continuous-time mean-field control when the population evolves according to unknown controlled McKean--Vlasov dynamics and only discrete-time transition data are available. Model-based mean-field control requires the continuous-time drift and diffusion coefficients, which are not directly observed from fixed-step transitions, while a direct reduction to a discrete-time Bellman equation loses the continuous-time generator structure. To bridge these two viewpoints, we introduce a Mean-Field-PhiBE (MF-PhiBE), which incorporates discrete-time transition information into a continuous-time PDE on the Wasserstein space. The MF-PhiBE replaces the unknown infinitesimal drift and covariance in the policy-evaluation equation by one-step estimators computed from data, while preserving the generator structure of the McKean-Vlasov HJB equation. We also derive a policy-gradient theorem for entropy-regularized randomized feedback policies, expressing the actor direction through an action-wise infinitesimal advantage and the score of the policy. Combining these two ingredients yields a model-free actor-critic method. We prove a first-order consistency estimate showing that the value induced by an optimal MF-PhiBE policy approximates the optimal continuous-time value as the observation time step vanishes. For entropy-regularized LQR, we establish first-order policy convergence and second-order value convergence; under suitable conditions, the population-averaged feedback means coincide exactly. Numerical experiments on an LQR benchmark and a crowd-aversion problem illustrate the proposed framework.

math.OC

Mean-Field Control with a Common Hidden State under Decentralized Observations

We study optimal control of a system with multiple decision makers who share a common hidden state and receive fully decentralized observations through identical channels. The dynamics of the hidden state and the cost incurred by the agents depend on the agents' actions only through their empirical distribution. In the limit problem with infinitely many agents, the problem reduces to a single agent control problem where the agent affects the hidden state dynamics via the conditional law of the actions given the past values of the hidden state process. We formulate this problem as a deterministic measure valued control problem over the space of policies and provide a dynamic programming recursion. We first show that for the limiting problem randomization over the control actions is necessary for optimality. However, randomization over the selection of policies (i.e., mixture policies) is not required. We then show that the optimal symmetric policies designed for the infinite population problem are near optimal for the finite population problem. In particular, we establish convergence rates that decay with number of agents as $\frac{1}{\sqrt{N}}$, and grow exponentially with the memory length used in the policy.

math.OC

A comparison principle for Wasserstein PDEs with state- and law-dependent common noise

We prove a comparison principle for a class of second-order Hamilton--Jacobi--Bellman equations on the Wasserstein space whose second-order term is generated by a general common-noise Hessian. The main difficulty is that the relevant second-order direction is induced by a state- and measure-dependent coefficient, so the associated perturbation of the measure is no longer a translation or a fixed state-dependent transformation. We introduce a nonlinear flow of measures and use it to transform the Wasserstein-space equation into an augmented equation on $[0,T]\times \mathcal P_2(\mathbb R)\times\mathbb R$, where the general Hessian becomes an ordinary second derivative in the auxiliary variable. The construction may be viewed as a measure-dependent Lamperti transform: it removes the common-noise direction at the level of the equation, but unlike the classical one-dimensional Lamperti transform it permits degeneracy of the coefficient and dependence on the conditional law. We establish the spatial, measure-derivative, and negative-Sobolev estimates for this flow that are needed in the viscosity argument. Under structural assumptions on the transformed Hamiltonian, these estimates yield a Crandall--Ishii type comparison theorem for semicontinuous viscosity sub- and supersolutions. This gives, to the best of our knowledge, the first viscosity comparison framework of this kind for the filtering-driven equations considered here, and opens a new class of second-order PDEs on spaces of measures with state- and law-dependent common-noise directions. As an application, we identify the value function of a controlled stochastic filtering problem with state- and law-dependent common noise as the unique viscosity solution of its dynamic programming equation. We also explain how the same change-of-variable viewpoint applies to Zakai-type Kolmogorov equations on spaces of finite positive measures.

math.AP

Infinite Horizon Optimal Consumption: Intertemporal Hedging under Epstein-Zin Preferences

We study an infinite-horizon optimal consumption-investment problem for an investor with Epstein-Zin stochastic differential utility in an incomplete market with stochastic investment opportunities. Risk aversion and intertemporal substitution are separated, and we work in the regime $\theta\in(0,1)$, where there exists a unique generalised utility process for arbitrary non-negative progressively measurable consumption streams. Our main contribution is a variational characterisation of the value function. We show that the value function is the unique minimiser of a functional whose Euler-Lagrange equation coincides with the Hamilton-Jacobi-Bellman equation. Although the functional may be non-convex, the direct method yields existence, and we prove that every minimiser is a strictly positive, bounded classical solution. A verification theorem identifies any minimiser with the value function and gives feedback representations for optimal consumption and investment policies. The proof combines a change of measure to the myopic probability with uniqueness results for Epstein-Zin BSDEs and a perturbation argument for optimality. Examples with stochastic volatility, Gaussian excess returns, and fat-tailed excess returns illustrate the scope of the framework and its implications for intertemporal hedging.

q-fin.MF

Policy Gradient for Continuous-Time Mean-Field Control

This paper develops a policy gradient method for entropy-regularized mean-field control in the discounted infinite-horizon setting. We consider randomized feedback policies and a coupled representative-particle/population system, in which the representative state evolves jointly with a population law governed by a McKean--Vlasov equation. The resulting value function is therefore defined on the product space $\mathbb R^d \times \mathcal P_2(\mathbb R^d)$. A key distinction from existing policy gradient methods for mean-field control is that, after computing the value function under a fixed policy, our approach does not require solving an additional equation to obtain the policy gradient. Instead, we derive an explicit policy gradient formula directly in terms of the value function. The formulation is based on an instantaneous advantage function, which quantifies the gain of taking a given action relative to the current randomized policy. We establish a G\^ateaux policy-gradient formula, which gives the first-order variation of the objective along arbitrary policy perturbations, and then derive the corresponding ascent direction under finite-dimensional policy parametrization. The resulting formula leads to a model-based actor--critic scheme. The critic is obtained by solving the associated linear stationary Hamilton--Jacobi--Bellman equation for the value function, using cylindrical functions to represent dependence on the population law. The actor is then updated according to the derived policy-gradient formula. We further analyze the well-posedness of the PDE in a polynomial-growth function class. Finally, we illustrate the proposed method through numerical experiments on an LQR model and a crowd-motion problem.

math.OC

Analytical Approach to Continuous-Time Causal Optimal Transport

We study causal optimal transport in continuous time, with Markovian cost, between a finite-state Markov source and a diffusion target. By replacing the source with its conditional law given the observation of the target, we characterize the value of this transport problem through a fully nonlinear parabolic master equation on an enlarged state space. We further show that this value coincides with those of two equivalent stochastic control problems on the simplex: a control of the Kushner--Stratonovich filtering equation with a zero-mean condition, and a state-constrained stochastic optimal control problem. Both formulations give rise to implementable numerical schemes that approximate the value from above and below.

math.OC

Equilibrium for Time-inconsistent Mean Field Games: A Systematic Analysis by Entropy Regularization

This paper studies the existence and approximation of equilibria for general time-inconsistent mean field game (MFG) problems in continuous time. To handle the intricate nonlocal equilibrium Hamilton-Jacobi-Bellman (EHJB) system arising from initial-time dependence, such as non-exponential discounting, we develop a vanishing entropy regularization approach. Using entropy regularization, we first characterize the regularized equilibrium through a coupled exploratory equilibrium HJB (EEHJB) equation and a law-dependent stochastic differential equation. By exploiting Schauder fixed-point arguments and tailored parabolic regularity estimates in a suitable functional space involving both value functions and measure flows, we establish the global existence of regularized equilibria under mild assumptions. We then establish convergence as the entropy regularization vanishes. By employing compactness arguments, Young measure techniques, and a duality tool for divergence-form Fokker-Planck equations, we prove that the regularized equilibria converge, up to subsequences, to a mean-field equilibrium of the original MFG. Furthermore, under entropy regularization, we propose a policy iteration algorithm and establish its convergence under short-time-horizon and weak-terminal-interaction conditions.

math.OC

When Diffusion Model Can Ignore Dimension: An Entropy-Based Theory

Diffusion models perform remarkably well on high-dimensional data such as images, often using only a modest number of reverse-time steps. Despite this practical success, existing convergence theory does not fully explain why such samplers remain efficient in high dimensions. Many prior KL guarantees bound the discretization error in terms of the ambient dimension, while other improved results replace this dependence using intrinsic-dimensional or geometric structure assumptions. In this work, we develop an alternative information-theoretic perspective on diffusion sampler convergence. We prove that, for Gaussian mixture targets, the discretization error is controlled by the Shannon entropy of the latent mixture component rather than by the ambient dimension. Consequently, the leading step complexity scales linearly with latent entropy and depends only logarithmically on the second moment of the data. Our analysis also extends to discrete target distributions, where the relevant complexity is the entropy of the target rather than the dimension of the embedding space. These results suggest that diffusion sampling can remain efficient in high-dimensional spaces when the data distribution admits a compact latent representation, as is widely believed to be the case for natural images.

cs.LG

Automation, Income Incidence, and Capital Accumulation in Incomplete Markets

This paper studies how automation changes the stationary distribution of income, consumption, and wealth in an incomplete-market economy. An automating sector trades off productivity gains and labor-cost savings against adoption costs. Households differ by skill and wealth, save in a capital/equity claim, and face uninsurable skill risk. Competitive factor prices and aggregate capital clear jointly with household Hamilton--Jacobi--Bellman equations and the stationary Kolmogorov forward equation. Automation affects consumption through labor-income incidence, precautionary saving, skill mobility, ownership of automation rents, and the stationary capital stock. In an adverse-incidence scenario with high exposure, adverse reskilling, capital obsolescence, and concentrated ownership, decentralized automation lowers stationary consumption and capital relative to the no-automation allocation. With stronger productivity and complementarity, lower obsolescence, and broader ownership, automation raises output, consumption, and capital. Reversing the assumed skill-mobility response also raises consumption, output, and capital substantially at a fixed automation intensity. A proxy diagnostic combining U.S. evidence on AI adoption, investment, labor-income pass-through, equity ownership, and marginal propensities to consume places the current economy near the boundary between the two scenarios. The model provides a quantitative framework for separating automation's aggregate gains from its distributional incidence.

econ.GN

Conditional Diffusion Under Linear Constraints: Langevin Mixing and Information-Theoretic Guarantees

We study zero-shot conditional sampling with pretrained diffusion models for linear inverse problems, including inpainting and super-resolution. In these problems, the observation determines only part of the unknown signal. The remaining degrees of freedom must be sampled according to the correct conditional data distribution. Existing projection-based samplers enforce measurement consistency by correcting the observed component during reverse diffusion. However, measurement consistency alone does not determine how probability mass should be distributed along the feasible set, and this can lead to biased conditional samples. We analyze this issue through a normal--tangent decomposition of the score function. For Gaussian noising, the observed-direction score is exactly determined by the measurement; only the tangent conditional score is unknown. We prove that the error from replacing this score by the unconditional tangent score is upper bounded by a dimension-free conditional mutual information between observed and unobserved components. This gives an information-theoretic decomposition into initialization and pathwise score-mismatch errors. Motivated by the theory, we propose a projected-Langevin initialization followed by guided reverse denoising, which outperforms a strong projection-based baseline in inpainting and super-resolution experiments.

cs.LG

Continuous-time Online Learning via Mean-Field Neural Networks: Regret Analysis in Diffusion Environments

We study continuous-time online learning where data are generated by a diffusion process with unknown coefficients. The learner employs a two-layer neural network, continuously updating its parameters in a non-anticipative manner. The mean-field limit of the learning dynamics corresponds to a stochastic Wasserstein gradient flow adapted to the data filtration. We establish regret bounds for both the mean-field limit and finite-particle system. Our analysis leverages the logarithmic Sobolev inequality, Polyak-Lojasiewicz condition, Malliavin calculus, and uniform-in-time propagation of chaos. Under displacement convexity, we obtain a constant static regret bound. In the general non-convex setting, we derive explicit linear regret bounds characterizing the effects of data variation, entropic exploration, and quadratic regularization. Finally, our simulations demonstrate the outperformance of the online approach and the impact of network width and regularization parameters.

cs.LG

Tractable bank capital structure: optimal control under Basel III constraints

Banks must optimize risky investments, dividend payouts, and capital structure under tight Basel III solvency and liquidity constraints, while costly equity issuance serves as a distress-recovery tool. We formulate this as a stochastic control problem that reduces the high-dimensional balance-sheet dynamics to a tractable one-dimensional process in the asset-to-deposit ratio, with state-dependent investment limits. The resulting policy is simple and interpretable: pay dividends at an upper reflection barrier and, when needed, recapitalize only at the distress boundary, jumping to an optimal target level. We characterize these thresholds analytically and show their sensitivity to regulatory parameters. From a regulatory viewpoint, we use Monte Carlo simulation to solve an outer optimization problem and map the efficient frontier between shareholder value and survival probability, both with and without a leverage cap. In the illustrative parameter ranges studied here, tightening solvency requirements often yields the best safety--profitability trade-off.

math.OC