Search arXiv⌕ Search

arXiv · 2609.32772

Learnable Randomization as Commitment Against Adaptive Optimizers

Abstract

A pricing page can walk the posted price up to the last amount a buyer still accepts, a recommender can hold back a better item for a barely acceptable promoted one, and a classifier can shift its boundary once applicants change their features. The system predicts the response and then picks the menu that serves its own objective, so the surplus above the user's cutoff is taken. Playing the single best action publishes that cutoff, while noise on actions the user would never take throws away payoff and teaches the platform that a worse menu is still acceptable. We study unpredictable near-optimal policies (UNOP), which mix uniformly on near-best actions that remain individually rational. The mixture is a commitment about the response. On a finite price grid, when the best sure-demand price strictly out-earns the randomized band, a seller who already knows the curve posts below the band, and the purchase that occurs is deterministic. Knowing that curve is not the same as predicting the next draw. The mixture can be learned and the optimizer can match its best response, while the user's payoff stays higher because the mixture changes which action is targeted. In pricing and in policy-aware recommendation this leaves more surplus than greedy play when the platform optimizes against the curve and more than one action is acceptable. The gain goes away under quality ranking, a singleton near-optimal set, a wrong utility estimate, or a short-horizon explorer. That is also where mixing should be turned off if the other side is trying to cooperate.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zihan Deng, Chuanzhi Xu, Xiaozhen Zhong, Haoyang Li, Junjie Huang. 2026-09-26. Learnable Randomization as Commitment Against Adaptive Optimizers. https://arxiv.org/abs/2609.32772

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Honest Reporting in Scored Oversight: True-KL0 Property via the Prekopa Principle

We prove the True-KL$_0$ property for a parametric family of heterogeneous scoring rules arising in scored elicitation mechanisms (AI oversight, forecasting, expert surveys). An agent with private type $M>1$, scored through a $d$-dimensional outcome interface, reports to a principal who evaluates via a power-$p$ pseudospherical scoring rule, $p \in (d,d+1)$; $M$ captures the agent's information quality relative to a reference. Honest reporting is dominant-strategy optimal for every $d$ and every $p>1$, without a prior over the agent's type: a consequence of strict properness and identifiability, with a quadratic misreport-loss rate. True-KL$_0$, the property $R(M,p,d)<1$ for all $M>1$, $d \in \{2,3,4\}$, $p \in (d,d+1)$, is the quantitative core: $R$ is the Rayleigh quotient of the radial misreport channel of an annular oversight model, and True-KL$_0$ certifies a uniform curvature-domination margin for that channel: $1-R \ge 0.26$ ($R \le 0.7324$, semi-rigorous numerical certificate). Two structural tools drive the proof: (i) a substitution $y=(x+1)/(x-1)$ rewrites the loss integral $I_L$ as $\int_1^M F(y)(M^2-y^2)^{d/2} dy$ with $M$-independent weight $F(y)>0$; (ii) log-concavity of $I_L$ in $M$: algebraic for $d=2$ up to a small certified compact verification, via Prekopa's theorem plus semi-rigorous certificates for $d \in \{3,4\}$. True-KL$_0$ then follows from elementary tail bounds plus a certified bound on $M \in [1.001, 20]$. We also characterise the dimensional boundary: True-KL$_0$ holds for all $p \in (d,d+1)$ when $d \le 4$; $d=5$ is the unique transition, with $p_{crit}(5) \in [5.5718, 5.5750]$ (mpmath, not interval-certified); for $d=6,7$ (and conjecturally all $d \ge 6$) no threshold exists: the bound fails at every sampled $p \in (d,d+1)$.

cs.GT↗

Profit Reallocation Mechanisms in Tree-based Data Trading

Markets for data promise to unlock its economic value, yet in practice they remain far less active than expected---one reason is that those who supply data are not rewarded for the value it creates downstream. A defining feature of data is its \emph{replicability}: a buyer can refine purchased data into a new product and resell it to \emph{many} downstream buyers, so a single source seeds a branching cascade of resales that naturally forms a \emph{tree}. Because an upstream seller captures none of this downstream value, its incentive to trade is weakened. Existing works propose \emph{profit reallocation}---returning part of downstream revenue to upstream contributors---as a natural remedy. But whether profit reallocation works on the tree-structured markets that replicable data actually induces has remained open. To bridge this gap, we develop a principled framework for profit reallocation on tree-structured data markets. We introduce a sequential trading game on a tree and a general class of budget-feasible profit reallocation mechanisms (PRMs) over it. We derive efficient algorithms to compute the induced equilibria---a polynomial-time exact algorithm for discrete valuations and a fully polynomial-time approximation scheme (FPTAS) for continuous ones---via a subtree decomposition technique that tames the potential coupling across a seller's children. We then prove that, under mild assumptions, \emph{any} budget-feasible PRM weakly expands the trades that occur in equilibrium, and any budget-balanced PRM additionally weakly improves social welfare, relative to the baseline that reallocates nothing. Experiments on synthetic markets confirm that these benefits are substantial, persist even when the assumptions fail, and grow with the depth and branching of the tree.

cs.GT↗

From Reconnaissance to Response: Quantitative Risk Parameterization and Game Theoretic Containment in Modern Enterprise Attack

Modern Security Operations Centers struggle with delayed manual incident response, enabling adversaries to advance through the Cyber Kill Chain during early stage reconnaissance. While classical game theoretic defense models optimize strategic resource allocation, they rely on static utility matrices that fail to adapt to dynamic telemetry. This paper presents an integrated, metrics driven decision engine that bridges quantitative risk parameterization and continuous automated response time. Common Vulnerability Scoring Systems exploitability parameters are mapped to attacker success probabilities and evaluate defender log distributions via Factor Analysis of Information Risk Monte Carlo simulations. Real time SIEM logs streams are modeled as Poisson process arrival rates, dynamically updating defender posterior threat belief through sequential Bayesian filtering. A closed form threshold is derived by framing the interaction as a dynamic Bayesian Stackelberg game, where the expected unmitigated risk exceeds proactive containment cost. Parameterized against empirical data from the 2023 MGM Resorts and Caesars Entertainment cyber incident, simulation results demonstrate that the engine suppresses transient background noise while triggering automated SOAR network isolation within seconds of adversarial probing. Multi parameter sensitivity analysis confirms that the decision boundary dynamically adjusts to live perimeter vulnerability, offering a control theoretic foundation for sub minute automated threat containment.

cs.GT↗