Search arXiv⌕ Search

arXiv · 2610.11690

Epistemic Disturbance in the Graph Model for Conflict Resolution: State-Preserving Actions, Four-Valued Assessments, and the Distinction between Capability and Intention

Abstract

In the graph model for conflict resolution (GMCR), a decision maker (DM) either moves the conflict to another state or does nothing. The basic definitions leave inaction implicit, so every action that leaves the state unchanged is treated as doing nothing. Yet announcements, exercises, leaks and selective disclosures are neither moves nor inaction: they leave the state unchanged but change what other DMs believe about which moves are available and which moves others would want to make. We introduce such state-preserving actions by augmenting states with the DMs' epistemic states: a physical move changes the physical state, a state-preserving action changes only the epistemic state, and inaction is the absence of a transition. Actions generate evidence through observer-specific interpretation maps. Building on a four-valued extension of GMCR from the author's earlier work, which separates evidence for and against, we show that evidence for a move can only enable perceived moves and evidence against can only disable them, that two of the four reduction operators ignore one kind of evidence, and that contradictory assessments are absorbing under monotone accumulation. With the monotonicity of stability in move sets, this fixes the direction in which any action moves a DM's stability judgements and characterizes when actions can enable provocation or deterrence. Capability assessments affect all sanction-based stability concepts, and on the DM's own side also Nash stability, whereas intention assessments affect only sequential stability. Hedging between two candidate types weakly expands or shrinks an observer's sequentially stable set according to how it reads contradiction. In the 1995 DVD format negotiation, general metarationality cannot distinguish its phases, since the computer industry group could always sanction; sequential stability, which asks whether it would, can.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yukiko Kato. 2026-10-08. Epistemic Disturbance in the Graph Model for Conflict Resolution: State-Preserving Actions, Four-Valued Assessments, and the Distinction between Capability and Intention. https://arxiv.org/abs/2610.11690

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Stackelberg POMDP: Learning to Lead via Reinforcement Learning

Many real-world domains--including e-commerce platform design, security planning, and multi-agent coordination--feature leader-follower problems where one decision-maker commits to a policy and others react strategically. We develop a reinforcement learning framework for such interactions in sequential environments with partial observations and multiple followers. Followers may adapt through no-regret learning or reinforcement learning, potentially departing from equilibrium behavior. The framework embeds follower adaptation into the leader's environment to construct a single-agent partially observable Markov decision process--the Stackelberg POMDP. For policy-interactive response algorithms, which access the leader's policy through queries, we prove that an optimal policy based only on the leader's game history yields an optimal commitment under the specified response procedure. We use proximal policy optimization with a centralized critic and train contextual meta-followers to respond across leader policies. In indirect mechanism design, mechanisms using buyer messages achieve higher social welfare than optimal standard sequential price mechanisms across all tested type counts, with responses certified as approximate Bayesian coarse correlated equilibria. In platform design, learned display rules increase mean consumer surplus by 8.4% over an optimized fixed price cap while accommodating hidden seller costs. In Atari bilateral trade, meta-learned follower responses support joint learning of visual gameplay and economic decisions; assigning leadership to the seller or buyer shifts transaction prices and payoffs in that agent's favor. Controlled ablations examine how response credit, policy consistency, and reward timing affect learning.

cs.GT↗

Price Competition Under Platform-Mediated Search: A Consider-Then-Choose Framework

We study the problem of predicting price equilibria on e-commerce platforms where sellers compete across multiple attributes (e.g., price, average rating, delivery speed). In these settings, a platform's design choices --- such as its display ranking, filtering tools, and promotional badges --- critically shape customer search and purchase behavior, which in turn determine sellers' equilibrium pricing strategies. Our goal is to develop a tractable framework that allows a platform to anticipate the market impact of its design interventions. We consider a behavioral model --- Consider-then-Choose with Lexicographic Choice (CLC) --- specifically tailored to platform-mediated search. We establish that any local Nash equilibrium admits a sequential-move characterization; this yields a tractable procedure for computation under an interpretable sufficient condition, which we term gradient dominance. We further prove that under gradient dominance, simple, decentralized gradient-based algorithms converge to an equilibrium, providing platforms with a method for simulating market outcomes. Finally, we use our framework to study how platform design affects market outcomes. Our framework applies to any platform in which sellers compete on multiple attributes and customer choice is guided by the platform's interface. In these environments, sellers' pricing strategies must be understood not in isolation, but as a response to the platform's design. Our work provides platform operators with a rigorous toolbox to efficiently evaluate how changes to interface design, information disclosure, and ranking policies can affect competitive outcomes.

cs.GT↗

Data-Driven Games with Coherent Risk Measures

We introduce Coherent Utility Measure Games (CUMGs) in which players' uncertainty about the distribution of payoffs is modeled using coherent utility (risk) measures. Such measures, including mean semideviation risk and conditional value-at-risk, allow for interpretable notions of players' risk aversion while retaining formal equivalence to distributionally robust games. While CUMGs, which are a subclass of distributionally robust games, are continuous games in general, they can be viewed as finite games ``lifted'' to the mixed strategy space, which illustrates computational challenges. Prior results extend to guarantee equilibrium existence in data-driven CUMGs. For CUMGs parameterized by several popular risk measures, we show that the computation of exact equilibria lies in FIXP, even for two-player games, and approximate equilibria lie in PPAD. Separately, we derive direct complementarity formulations for exact equilibrium computation for these games, which grow with $K$, the number of data samples. Unlike standard games, these programs are not linear in a two-player setting. Next, we establish the existence of approximate equilibria in finite data-driven CUMGs with small supports in the players' pure actions, yielding a quasi-polynomial time approximation scheme (QPTAS); this, together with a sparse data subsample result, guides the search for such equilibria. We also develop a stochastic first-order approach for smoothed CUMGs using data mini-batches, with bounds linking first-order error to approximate equilibrium. We include numerical experiments exploring the structure of equilibrium in CUMGs and comparing the various approaches in this work.

cs.GT↗