Search arXiv⌕ Search

arXiv · 2610.09843

Minimizing Cumulative Envy in Allocating a Sequence of Items

Abstract

We study temporal fair division with indivisible goods that arrive sequentially and must be allocated irrevocably. In contrast to the usual online model, we assume that valuations and future arrivals are known in advance, and ask how unfairness evolves during the process. We introduce \emph{cumulative maximum envy}: the sum, over all rounds, of the maximum pairwise envy at that round. Equivalently, this is the area under the worst-envy curve, and it captures both the magnitude and the duration of envy. For a fixed arrival order, we show that the corresponding decision problem is strongly NP-complete and that minimizing this objective admits no constant-factor approximation unless P = NP, even under identical valuations and even under binary valuations. We complement these hardness results with a dynamic program that gives pseudopolynomial-time solvability for a constant number of agents, polynomial-time algorithms in further restricted settings, and an FPTAS for fixed $n$ under identical integer valuations. We then study a sequencing variant where the algorithm may choose the arrival order. This variant remains NP-complete even for two agents with identical valuations; however, a simple greedy algorithm achieves a $3/2$-approximation for $n=2$ agents, an $n/(n-1)$-approximation for any number of agents, and an additive guarantee depending on the maximum value of any good.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Paul W. Goldberg, Isaac Robinson, Nicholas Teh. 2026-10-07. Minimizing Cumulative Envy in Allocating a Sequence of Items. https://arxiv.org/abs/2610.09843

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Stackelberg POMDP: Learning to Lead via Reinforcement Learning

Many real-world domains--including e-commerce platform design, security planning, and multi-agent coordination--feature leader-follower problems where one decision-maker commits to a policy and others react strategically. We develop a reinforcement learning framework for such interactions in sequential environments with partial observations and multiple followers. Followers may adapt through no-regret learning or reinforcement learning, potentially departing from equilibrium behavior. The framework embeds follower adaptation into the leader's environment to construct a single-agent partially observable Markov decision process--the Stackelberg POMDP. For policy-interactive response algorithms, which access the leader's policy through queries, we prove that an optimal policy based only on the leader's game history yields an optimal commitment under the specified response procedure. We use proximal policy optimization with a centralized critic and train contextual meta-followers to respond across leader policies. In indirect mechanism design, mechanisms using buyer messages achieve higher social welfare than optimal standard sequential price mechanisms across all tested type counts, with responses certified as approximate Bayesian coarse correlated equilibria. In platform design, learned display rules increase mean consumer surplus by 8.4% over an optimized fixed price cap while accommodating hidden seller costs. In Atari bilateral trade, meta-learned follower responses support joint learning of visual gameplay and economic decisions; assigning leadership to the seller or buyer shifts transaction prices and payoffs in that agent's favor. Controlled ablations examine how response credit, policy consistency, and reward timing affect learning.

cs.GT↗

Price Competition Under Platform-Mediated Search: A Consider-Then-Choose Framework

We study the problem of predicting price equilibria on e-commerce platforms where sellers compete across multiple attributes (e.g., price, average rating, delivery speed). In these settings, a platform's design choices --- such as its display ranking, filtering tools, and promotional badges --- critically shape customer search and purchase behavior, which in turn determine sellers' equilibrium pricing strategies. Our goal is to develop a tractable framework that allows a platform to anticipate the market impact of its design interventions. We consider a behavioral model --- Consider-then-Choose with Lexicographic Choice (CLC) --- specifically tailored to platform-mediated search. We establish that any local Nash equilibrium admits a sequential-move characterization; this yields a tractable procedure for computation under an interpretable sufficient condition, which we term gradient dominance. We further prove that under gradient dominance, simple, decentralized gradient-based algorithms converge to an equilibrium, providing platforms with a method for simulating market outcomes. Finally, we use our framework to study how platform design affects market outcomes. Our framework applies to any platform in which sellers compete on multiple attributes and customer choice is guided by the platform's interface. In these environments, sellers' pricing strategies must be understood not in isolation, but as a response to the platform's design. Our work provides platform operators with a rigorous toolbox to efficiently evaluate how changes to interface design, information disclosure, and ranking policies can affect competitive outcomes.

cs.GT↗

Data-Driven Games with Coherent Risk Measures

We introduce Coherent Utility Measure Games (CUMGs) in which players' uncertainty about the distribution of payoffs is modeled using coherent utility (risk) measures. Such measures, including mean semideviation risk and conditional value-at-risk, allow for interpretable notions of players' risk aversion while retaining formal equivalence to distributionally robust games. While CUMGs, which are a subclass of distributionally robust games, are continuous games in general, they can be viewed as finite games ``lifted'' to the mixed strategy space, which illustrates computational challenges. Prior results extend to guarantee equilibrium existence in data-driven CUMGs. For CUMGs parameterized by several popular risk measures, we show that the computation of exact equilibria lies in FIXP, even for two-player games, and approximate equilibria lie in PPAD. Separately, we derive direct complementarity formulations for exact equilibrium computation for these games, which grow with $K$, the number of data samples. Unlike standard games, these programs are not linear in a two-player setting. Next, we establish the existence of approximate equilibria in finite data-driven CUMGs with small supports in the players' pure actions, yielding a quasi-polynomial time approximation scheme (QPTAS); this, together with a sparse data subsample result, guides the search for such equilibria. We also develop a stochastic first-order approach for smoothed CUMGs using data mini-batches, with bounds linking first-order error to approximate equilibrium. We include numerical experiments exploring the structure of equilibrium in CUMGs and comparing the various approaches in this work.

cs.GT↗