Search arXivSearch

arXiv subjects

Bar Light

Publications and source records attributed to Bar Light.

At least 19 recordsLinked to original sources

Equilibrium in Multi-Agent Reinforcement Learning

Standard solution concepts for stochastic games, such as Markov perfect equilibrium and Markov coarse correlated equilibrium, are computationally difficult, and thus, standard decentralized reinforcement-learning algorithms should not generally be expected to converge to them. In this paper, we study the equilibrium generated by such algorithms. In particular, we introduce a new solution concept for stochastic games, Markov Bayes coarse correlated equilibrium (MBCCE), defined as a distribution over states and stationary policy profiles such that, after observing the state but before observing her recommended action, no player can gain by choosing a different current action, with the sampled policy profile governing play thereafter. We discuss the parallels between MBCCE and coarse correlated equilibrium (CCE) in finite normal-form games and show that MBCCE retains several of its key properties. We then introduce a corresponding regret notion, adaptive Markov coarse regret (AMCR), and show that vanishing AMCR implies that every accumulation point of the empirical distribution of realized states and policy profiles is an MBCCE. Crucially, we show that achieving AMCR reduces to two standard learning tasks: minimizing external regret at each state and accurately evaluating the current joint policy. We then prove that under mild conditions these properties hold for two natural RL algorithmic designs: a decentralized asynchronous actor--critic algorithm through a new two-timescale stochastic-approximation analysis, and a standard episodic multi-agent projected policy-gradient method. Hence, both algorithms generate approximate MBCCEs, and we establish explicit finite-time convergence rates for both.

cs.GT

Tractable Relaxations of Multivariate Stochastic Dominance via Optimal Transport and CVaR

Many operational decisions involve alternatives with several uncertain attributes. When comparing such alternatives, stochastic dominance requires every decision maker in a prescribed class to prefer one alternative to the other. These orders have two known limitations: a single extreme decision maker can rule out dominance, and multivariate dominance can be difficult to verify. To address these limitations, we develop two relaxations of their standard representations: compensated stochastic dominance (CSD) relaxes Strassen's coupling representation, while reference-weighted stochastic dominance (RWSD) relaxes the integral representation. We characterize the utility class generated by CSD, recovering known almost stochastic dominance orders as special cases, and characterize it for RWSD in several cases. Importantly, we show that CSD and RWSD can each be checked by computing a single value, an optimal-transport value for CSD and a CVaR value for RWSD, and verifying that this value is nonpositive. This inequality characterization allows us to derive finite-sample guarantees for both orders, even when the outcome dimension is large. An application to Olist e-commerce data shows that, even when empirical FOSD fails, CSD and RWSD quantify how small this failure is and what is needed to overcome it.

math.PR

Conjectural Variations in Competitive Dynamic Pricing: A Learning Foundation via Experimentation Design and Feedback Structure

We study competitive dynamic pricing among multiple sellers, motivated by the rise of large-scale experimentation and algorithmic pricing in retail and online marketplaces. Sellers repeatedly set prices using simple learning rules and observe their own realized demand, while possibly observing only a subset of rivals' prices, even though demand depends on all sellers' prices and is subject to random shocks. Each seller runs local price experiments, such as switchback-style designs, and updates a focal price using a linear demand estimate fitted to its own demand data and the competitor prices it observes. Under certain conditions on demand, the resulting dynamics converge to a Conjectural Variations (CV) equilibrium, a classic static equilibrium notion in which each seller best responds under a conjecture that rivals' prices co-move systematically to changes in its own price. Unlike standard CV models that treat conjectures as behavioral primitives, we show that these conjectures arise endogenously from the interaction between the feedback structure and the correlation structure of experimentation. When a seller does not observe some rivals' prices, correlated experimentation induces an omitted-variable bias in demand estimation. We show that this bias determines the conjectures that govern the long-run equilibrium. Notably, when this learning bias vanishes, for example under full price feedback or independent experimentation of unobserved rivals, the learning dynamics converge to the standard Nash equilibrium. We provide simple sufficient conditions on demand for convergence in standard models and establish a finite-sample guarantee, showing that the mean squared price error decays at a rate of $\widetilde O (T^{-1/2})$.

cs.GT

Bounded Foresight Equilibrium in Large Dynamic Economies with Heterogeneous Agents and Aggregate Shocks

Large dynamic economies with heterogeneous agents and aggregate shocks are central to many important applications, yet their equilibrium analysis remains computationally challenging. This is because the standard solution approach, rational expectations equilibria require agents to predict the evolution of the full cross-sectional distribution of state variables, leading to an extreme curse of dimensionality. In this paper, we introduce a novel equilibrium concept, N-Bounded Foresight Equilibrium (N-BFE), and establish its existence under mild conditions. In N-BFE, agents optimize over an infinite horizon but form expectations about key economic variables only for the next N periods. Beyond this horizon, they assume that economic variables remain constant and use a predetermined continuation value. This equilibrium notion reduces computational complexity and draws a direct parallel to lookahead policies in reinforcement learning, where agents make near-term calculations while relying on approximate valuations beyond a computationally feasible horizon. At the same time, it lowers cognitive demands on agents while better aligning with the behavioral literature by incorporating time inconsistency and limited attention, all while preserving desired forward-looking behavior and ensuring that agents still respond to policy changes. Importantly, in N-BFE equilibria, forecast errors arise endogenously. We measure the foresight errors for different foresight horizons and show that foresight significantly influences the variation in endogenous equilibrium variables, distinguishing our findings from traditional risk aversion or precautionary savings channels. This variation arises from a feedback mechanism between individual decision-making and equilibrium variables, where increased foresight induces greater non-stationarity in agents' decisions and, consequently, in economic variables.

econ.GN

Computing and Learning Stationary Mean Field Equilibria with Low-Dimensional Interactions: Algorithms and Applications

Mean field equilibrium (MFE) has emerged as a computationally tractable solution concept for large dynamic games. However, computing MFE remains challenging due to nonlinearities and the absence of contraction properties, limiting its reliability for counterfactual analysis and comparative statics. This paper studies dynamic models in which agents interact through a small number of functions of the population distribution, with particular emphasis on the scalar case. Such low-dimensional interactions naturally arise in a wide range of applications in economics and operations. The main contribution of this paper is to introduce iterative algorithms that leverage this structure and have provide global convergence guarantees for computing and learning MFE under mild assumptions. Unlike existing approaches, our algorithms do not require monotonicity or contraction properties. We also provide model-free algorithms that learn an approximate MFE from simulation using reinforcement learning methods, without requiring prior knowledge of payoff or transition functions. Beyond computation, we establish existence of stationary MFE for non-compact state spaces. For scalar interactions, we also derive analytical comparative statics without requiring monotonicity of the equilibrium mapping. We apply our methods to classical models of dynamic competition, including capacity competition; to heterogeneous-agent macroeconomic models; and to models motivated by online marketplaces and learning, including inventory competition, ridesharing, and social learning. The applications illustrate how changes in market parameters affect equilibrium outcomes and how the algorithms can be used for reliable counterfactual analysis.

econ.TH

Equilibria under Dynamic Benchmark Consistency in Non-Stationary Multi-Agent Systems

We formulate and study a general time-varying multi-agent system where players repeatedly compete under incomplete information. Our work is motivated by scenarios commonly observed in online advertising and retail marketplaces, where agents and platform designers optimize algorithmic decision-making in dynamic competitive settings. In these systems, no-regret algorithms that provide guarantees relative to \emph{static} benchmarks can perform poorly and the distributions of play that emerge from their interaction do not correspond anymore to static solution concepts such as coarse correlated equilibria. Instead, we analyze the interaction of \textit{dynamic benchmark} consistent policies that have performance guarantees relative to \emph{dynamic} sequences of actions, and through a novel \textit{tracking error} notion we delineate when their empirical joint distribution of play can approximate an evolving sequence of static equilibria. In systems that change sufficiently slowly (sub-linearly in the horizon length), we show that the resulting distributions of play approximate the sequence of coarse correlated equilibria, and apply this result to establish improved welfare bounds for smooth games. On a similar vein, we formulate internal dynamic benchmark consistent policies and establish that they approximate sequences of correlated equilibria. Our findings therefore suggest that in a broad range of multi-agent systems where non-stationarity is prevalent, algorithms designed to compete with dynamic benchmarks can improve both individual and welfare guarantees, and their emerging dynamics approximate a sequence of static equilibrium outcomes.

cs.GT

A Course in Dynamic Optimization

These lecture notes are derived from a graduate-level course in dynamic optimization, offering an introduction to techniques and models extensively used in management science, economics, operations research, engineering, and computer science. The course emphasizes the theoretical underpinnings of discrete-time dynamic programming models and advanced algorithmic strategies for solving these models. Unlike typical treatments, it provides a proof for the principle of optimality for upper semi-continuous dynamic programming, a middle ground between the simpler countable state space case \cite{bertsekas2012dynamic}, and the involved universally measurable case \cite{bertsekas1996stochastic}. This approach is sufficiently rigorous to include important examples such as dynamic pricing, consumption-savings, and inventory management models. The course also delves into the properties of value and policy functions, leveraging classical results \cite{topkis1998supermodularity} and recent developments. Additionally, it offers an introduction to reinforcement learning, including a formal proof of the convergence of Q-learning algorithms. Furthermore, the notes delve into policy gradient methods for the average reward case, presenting a convergence result for the tabular case in this context. This result is simple and similar to the discounted case but appears to be new.

math.OC

A Note on the Stability of Monotone Markov Chains

This note studies monotone Markov chains, a subclass of Markov chains with extensive applications in operations research and economics. While the properties that ensure the global stability of these chains are well studied, their establishment often relies on the fulfillment of a certain splitting condition. We address the challenges of verifying the splitting condition by introducing simple, applicable conditions that ensure global stability. The simplicity of these conditions is demonstrated through various examples including autoregressive processes, portfolio allocation problems and resource allocation dynamics.

math.PR

Invariant Distributions in Nonlinear Markov Chains with Aggregators: Theory, Computation, and Applications

We study the properties of a subclass of stochastic processes called discrete time nonlinear Markov chains with an aggregator, which naturally appear in various topics such as strategic queueing systems, inventory dynamics, opinion dynamics, and wealth dynamics. In these chains, the next period's distribution depends on both the current state and a real-valued function of the current distribution. For these chains, we provide conditions for the uniqueness of an invariant distribution that do not rely on typical contraction arguments. Instead, our approach leverages flexible monotonicity properties imposed on the nonlinear Markov kernel. We demonstrate the necessity of these monotonicity conditions for proving the uniqueness of an invariant distribution through simple examples. We also provide existence results and introduce an iterative computational method that solves a simpler, tractable subproblem in each iteration and converges to the invariant distribution of the nonlinear Markov chain, even in cases where uniqueness does not hold. We leverage our findings to analyze invariant distributions in strategic queueing systems, study inventory dynamics when retailers optimize pricing and inventory decisions, establish conditions ensuring the uniqueness of solutions for a class of nonlinear equations in $\mathbb{R}^{n}$, and investigate the properties of stationary wealth distributions in large dynamic economies.

math.PR

The Principle of Optimality in Dynamic Programming: A Pedagogical Note

The principle of optimality is a fundamental aspect of dynamic programming, which states that the optimal solution to a dynamic optimization problem can be found by combining the optimal solutions to its sub-problems. While this principle is generally applicable, it is often only taught for problems with finite or countable state spaces in order to sidestep measure-theoretic complexities. Therefore, it cannot be applied to classic models such as inventory management and dynamic pricing models that have continuous state spaces, and students may not be aware of the possible challenges involved in studying dynamic programming models with general state spaces. To address this, we provide conditions and a self-contained simple proof that establish when the principle of optimality for discounted dynamic programming is valid. These conditions shed light on the difficulties that may arise in the general state space case. We provide examples from the literature that include the relatively involved case of universally measurable dynamic programming and the simple case of finite dynamic programming where our main result can be applied to show that the principle of optimality holds.

math.OC

A Characterization of the n-th Degree Bounded Stochastic Dominance

We provide a novel characterization of the $n$-th degree bounded stochastic dominance (BSD) order, linking it to the risk tolerance of decision-makers and providing a decision-theoretic foundation for these stochastic orders. Our results reveal that BSD reflects specific risk preferences through the choice of the interval $[a,b]$, by characterizing it in terms of utility functions with globally bounded Arrow--Pratt risk aversion or that satisfy an $n$-convexity condition. They also highlight limitations of BSD, including its dependence on the chosen support interval and the resulting peculiar risk aversion behavior of decision-makers included in the generator of BSD. To partially address this issue, we use our characterization to separate two roles that are combined in BSD: the largest payoff in the lotteries and the upper endpoint of the interval that determines the Arrow--Pratt lower bound. We then introduce a related lower-partial-moment order that provides a clean trade-off between expected value and downside-risk protection. Using our characterization, we present comparative statics results for decision-making under uncertainty with globally bounded risk aversion measures and savings decisions under globally bounded prudence measures, and derive inequalities for $n$-convex functions.

math.PR

Equilibria in Repeated Games under No-Regret with Dynamic Benchmarks

In repeated games, strategies are often evaluated by their ability to guarantee the performance of the single best action that is selected in hindsight, a property referred to as \emph{Hannan consistency}, or \emph{no-regret}. However, the effectiveness of the single best action as a yardstick to evaluate strategies is limited, as any static action may perform poorly in common dynamic settings. Our work therefore turns to a more ambitious notion of \emph{dynamic benchmark consistency}, which guarantees the performance of the best \emph{dynamic} sequence of actions, selected in hindsight subject to a constraint on the allowable number of action changes. Our main result establishes that for any joint empirical distribution of play that may arise when all players deploy no-regret strategies, there exist dynamic benchmark consistent strategies such that if all players deploy these strategies the same empirical distribution emerges when the horizon is large enough. This result demonstrates that although dynamic benchmark consistent strategies have a different algorithmic structure and provide significantly enhanced individual assurances, they lead to the same equilibrium set as no-regret strategies. Moreover, the proof of our main result uncovers the capacity of independent algorithms with strong individual guarantees to foster a strong form of coordination.

cs.GT

Budget Pacing in Repeated Auctions: Regret and Efficiency without Convergence

We study the aggregate welfare and individual regret guarantees of dynamic \emph{pacing algorithms} in the context of repeated auctions with budgets. Such algorithms are commonly used as bidding agents in Internet advertising platforms, adaptively learning to shade bids by a tunable linear multiplier in order to match a specified budget. We show that when agents simultaneously apply a natural form of gradient-based pacing, the liquid welfare obtained over the course of the learning dynamics is at least half the optimal expected liquid welfare obtainable by any allocation rule. Crucially, this result holds \emph{without requiring convergence of the dynamics}, allowing us to circumvent known complexity-theoretic obstacles of finding equilibria. This result is also robust to the correlation structure between agent valuations and holds for any \emph{core auction}, a broad class of auctions that includes first-price, second-price, and generalized second-price auctions as special cases. For individual guarantees, we further show such pacing algorithms enjoy \emph{dynamic regret} bounds for individual utility- and value-maximization, with respect to the sequence of budget-pacing bids, for any auction satisfying a monotone bang-for-buck property. To complement our theoretical findings, we provide semi-synthetic numerical simulations based on auction data from the Bing Advertising platform.

cs.GT

Social Learning under Platform Influence: Consensus and Persistent Disagreement

Individuals increasingly rely on social networking platforms to form opinions. However, these platforms typically aim to maximize engagement, which may not align with social good. In this paper, we introduce an opinion dynamics model where agents are connected in a social network, and update their opinions based on their neighbors' opinions and on the content shown to them by the platform. We focus on a stochastic block model with two blocks, where the initial opinions of the individuals in different blocks are different. We prove that for large and dense enough networks the trajectory of opinion dynamics in such networks can be approximated well by a simple two-agent system. The latter admits tractable analytical analysis, which we leverage to provide interesting insights into the platform's impact on the social learning outcome in our original two-block model. Specifically, by using our approximation result, we show that agents' opinions approximately converge to some limiting opinion, which is either: consensus, where all agents agree, or persistent disagreement, where agents' opinions differ. We find that when the platform is weak and there is a high number of connections between agents with different initial opinions, a consensus equilibrium is likely. In this case, even if a persistent disagreement equilibrium arises, the polarization in this equilibrium, i.e., the degree of disagreement, is low. When the platform is strong, a persistent disagreement equilibrium is likely and the equilibrium polarization is high. A moderate platform typically leads to a persistent disagreement equilibrium with moderate polarization. We analyze the effect of initial polarization on consensus and explore various extensions including a three block stochastic model and a correlation between initial opinions and agents' connection probabilities.

econ.TH

Hermite-Hadamard inequalities for (p,a,b)-convex functions

A function $f:[a,b] \rightarrow \mathbb{R}$ is called $(p,a,b)$-convex if $f$ is $p$ times continuously differentiable, $f^{(p)}$ is convex and increasing, and $f^{(k)}(a)=0$ for all $k=1,\ldots,p$ where $f^{(j)}$ is the $j$th derivative of $f$. In this note we prove Hermite-Hadamard inequalities for $(p,a,b)$-convex functions that are significantly tighter than the classical Hermite-Hadamard inequality. We also prove inequalities for fractional integrals that involve $(p,a,b)$-convex functions.

math.CA

New Jensen-type inequalities and their applications

Convex analysis is fundamental to proving inequalities that have a wide variety of applications in economics and mathematics. In this paper we provide Jensen-type inequalities for functions that are, intuitively, "very" convex. These inequalities are simple to apply and can be used to generalize and extend previous results or to derive new results. We apply our inequalities to quantify the notion "more risk averse" provided in \cite{pratt1978risk}. We also apply our results in other applications from different fields, including risk measures, Poisson approximation, moment generating functions, log-likelihood functions, and Hermite-Hadamard type inequalities.

math.OC

Concentration inequalities using higher moments information

In this paper, we generalize and improve some fundamental concentration inequalities using information on the random variables' higher moments. In particular, we improve the classical Hoeffding's and Bennett's inequalities for the case where there is some information on the random variables' first $p$ moments for every positive integer $p$. Importantly, our generalized Hoeffding's inequality is tighter than Hoeffding's inequality and is given in a simple closed-form expression for every positive integer $p$. Hence, the generalized Hoeffding's inequality is easy to use in applications. To prove our results, we derive novel upper bounds on the moment-generating function of a random variable that depend on the random variable's first $p$ moments and show that these bounds satisfy appropriate convexity properties.

math.PR

Quality Selection in Two-Sided Markets: A Constrained Price Discrimination Approach

Online platforms collect rich information about participants and then share some of this information back with them to improve market outcomes. In this paper we study the following information disclosure problem in two-sided markets: If a platform wants to maximize revenue, which sellers should the platform allow to participate, and how much of its available information about participating sellers' quality should the platform share with buyers? We study this information disclosure problem in the context of two distinct two-sided market models: one in which the platform chooses prices and the sellers choose quantities (similar to ride-sharing), and one in which the sellers choose prices (similar to e-commerce). Our main results provide conditions under which simple information structures commonly observed in practice, such as banning certain sellers from the platform while not distinguishing between participating sellers, maximize the platform's revenue. The platform's information disclosure problem naturally transforms into a constrained price discrimination problem where the constraints are determined by the equilibrium outcomes of the specific two-sided market model being studied. We analyze this constrained price discrimination problem to obtain our structural results.

econ.TH