Search arXiv⌕ Search

arXiv subjects

Chunyuan Zheng

Publications and source records attributed to Chunyuan Zheng.

At least 19 recordsLinked to original sources

Closing the Accuracy Gap in Stochastic First-Order Bilevel Optimization

We characterize the optimal accuracy dependence for smooth nonconvex--strongly-convex bilevel optimization with stochastic first-order oracles. For globally unbiased fresh-gradient observations with bounded variance, we prove an $Ω(ε^{-6})$ lower bound matching the accuracy exponent of existing upper bounds. We consider globally Lipschitz gradients, a bounded upper gradient in the lower variable, and Lipschitz lower Hessian blocks, with target $\mathbb{E}|\nabla F(\widehat{x})|\leε$. Let $L$ be the common first-order scale, $μ$ the lower strong-convexity modulus, $ρ$ the lower Hessian-variation budget, $κ=L/μ$, and $χ=1+ρ/μ$. For $ρ\gtrsimμ$, $χ\lesssimκ$, sufficiently high accuracy, and sufficiently large dimension, we establish $Ω\left(Lκ^2Δε^{-2}+L^3χ^2κ^8σ_g^2Δε^{-6}\right)$, where $Δ$ is the initial gap budget and $σ_g^2$ is the lower-gradient variance budget. The bound holds over full Euclidean spaces against arbitrary randomized adaptive algorithms, even when every sample is the gradient of a scalar function. In the common-scale regime $ρ=Θ(L)$, the leading stochastic term has condition-number dependence $κ^{10}$, compared with the achievable $κ^{11}$ dependence. The proof embeds a sequential hard objective into an exact lower response while preserving global regularity and finite gap. A multiscale decomposition limits the gradient signal carrying each new direction, yielding the required sample complexity. We complement this lower bound with a frozen-linear-tilt estimator whose bias is proportional to lower Hessian variation. Its analysis makes this structural dependence explicit and yields matching leading rates in the specified curvature-controlled and small-gap regimes.

math.OC↗

The Minimax Cost of Smoothness in Fully First-Order Stochastic Bilevel Optimization

We study the sample complexity of finding a point with expected hypergradient norm at most $ε>0$ in nonconvex--strongly-convex stochastic bilevel optimization using fresh, globally unbiased first-order samples with bounded variance. For lower-variable smoothness order $p\ge1$, F$^2$SA-$p$ achieves $\widetilde O(pε^{-4-2/p})$ (Chen et al., 2026), but necessity of the exponent $4+2/p$ was open. We prove matching upper and lower bounds $Θ_p(ε^{-4-2/p})$ under fixed nondegenerate regularity budgets, with constants allowed to depend on $p$. The lower bound holds for unrestricted randomized algorithms under a fixed sample budget, allows dimension to grow with accuracy, and preserves global unbiasedness and all prescribed lower-variable smoothness bounds, even with a scalar lower variable and exact upper gradients. We further characterize the minimax fixed-budget complexity under $L_j=MΛ^j(j!)^β$ for $j\ge3$: $Q_{p,β}^*(ε)=Θ\left(ε^{-4}\left[\inf_{1\le r\le p,,r\in\mathbb N} r^βε^{-1/r}\right]^2\right)$, for $p\in\mathbb N\cup{\infty}$, with constants independent of $p$. At infinite order, this gives $Θ(ε^{-4})$ for $β=0$ and $Θ(ε^{-4}\log^2(1/ε))$ for $β=1$. A gradient-difference method with weighted independent batches attains these bounds uniformly in order. Thus quantitative smoothness yields precise gains, whereas qualitative analyticity alone still permits $Θ(ε^{-6})$ complexity.

math.OC↗

Matrix-Vector Complexity of Low-Rank Approximation

We establish matching polynomial query bounds for low-rank approximation from exact matrix--vector products. Given an unknown matrix $A\in\mathbb{R}^{m\times n}$, at each step a randomized algorithm chooses either $v\in\mathbb{R}^n$ and receives $Av$, or $u\in\mathbb{R}^m$ and receives $A^\top u$. The choice may depend measurably on all previous queries and replies and on the algorithm's private randomness; each vector product costs one query. The output is a rank-$k$ right projector with Schatten-$p$ residual at most $1+\varepsilon$ times optimal. Write $N=\min\{m,n\}$ and let $Q_p^*$ denote the worst-case query budget for success probability $2/3$ on every input. For every $1\le k<N$ and sufficiently small $\varepsilon$, our lower bounds, combined with existing Krylov upper bounds, give $Q_p^*=\widetildeΘ\!\left(\min\{N,k\min\{p^{1/6}\varepsilon^{-1/3},\varepsilon^{-1/2}\}\}\right)$ $(2\le p<\infty)$, $Q_\infty^*=\widetildeΘ\!\left(\min\{N,k\varepsilon^{-1/2}\}\right)$. These bounds have universal constants and allow $p$ to vary with the problem parameters, identifying the transition at $p\varepsilon\asymp1$. A complementary result for each fixed $1\le p<2$ gives $\widetildeΘ_p(\min\{N,k\varepsilon^{-1/3}\})$, with constants and an accuracy threshold that may depend on $p$. Together, the results recover this fixed-norm rate for every fixed finite $p$, supplying the multiplicative rank dependence missing from previous lower bounds. Tildes suppress logarithmic factors. The proof extends adaptive Wishart deferred decisions to a rectangular factor with a $k$-dimensional nullspace. Posterior overlap gives a short fixed-norm argument, while persistence of small compression eigenvalues controls growing $p$ and the spectral endpoint. Exact range recovery handles target costs of order $k$; the Wishart family covers the remaining regimes.

math.OC↗

Sharp High-Probability Statistical Rates for Heavy-Tailed Nonsmooth Convex Optimization

We establish sharp high-probability statistical rates for nonsmooth stochastic convex optimization under heavy-tailed noise, closing the confidence-dependence gap in both general convex and strongly convex settings. We consider composite objectives with a Lipschitz convex loss and a known proximal regularizer, accessed through fresh unbiased stochastic subgradients. The centered noise has conditional norm and directional $p$th moments bounded by $L^p$ and $s^p$, respectively, where $1<p\le2$ and $0<s\le L$. For $T$ oracle calls and failure probability $0<δ<1/10$, set $q=1-1/p$, $\ell=\log(3/δ)$, and $A=s(L^2/s^2)^q$. With initial distance bound $D$, the minimax objective error is $Θ(D(A+s\ell^q)/T^q)$ in the nonsaturated statistical regime. For a $μ$-strongly convex regularizer, the corresponding quadratic-regime rate is $Θ((A^2+s^2\ell^{2q})/(μT^{2q}))$. These rates require domination of the deterministic and finite-variance terms; matching lower bounds require $d\ge\lceil L^2/s^2\rceil$, while the upper bounds hold in every finite dimension. The new upper bounds attain known statistical lower-bound orders, replacing the previous confidence coefficient $s^{1/p}L^{1-1/p}$ by $s$, and its square by $s^2$ under strong convexity. The key is an online inequality that separates directional fluctuations from the total second moment created by rejection sampling. Localization preserves this separation under strong convexity without additional horizon logarithms. The results characterize oracle complexity with unrestricted internal computation.

math.OC↗

Joint Lower Bounds for Zeroth-Order Nonconvex Optimization on Euclidean Balls

We prove a joint stochastic zeroth-order lower bound for Goldstein stationarity on a Euclidean query ball, even when the ball is guaranteed to contain a stationary point. In dimension $d$, let $f=\mathbb{E}[F(\cdot;ξ)]$, assume $\mathbb{E}[\operatorname{Lip}(F(\cdot;ξ))^2]\le L_0^2$, and bound the initial objective gap over the ball by $Δ$. For neighborhood radius $δ>0$ and residual tolerance $\varepsilon>0$, our smooth hard family requires $Ω(dL_0^2Δ/(δ\varepsilon^3))$ scalar evaluations for success probability $1/2$, against randomized adaptive algorithms that may retain and repeatedly query each sampled function. The result holds for $\varepsilon\le cL_0$, $Δ\ge Cδ\varepsilon$, and $d\ge C[1+\log(2+ΔL_0^2/(δ\varepsilon^3))]$, on a constructed ball of radius $Θ(Δ/\varepsilon)$. A sequence of localized regions forces repeated direction estimation, and an adaptive Gaussian posterior argument controls sample reuse. The result establishes a joint dimension and accuracy obstruction on solvable bounded-domain instances. The gap is local to the query ball; a matching minimax characterization at a common radius and the corresponding unrestricted global-gap lower bound remain open here.

math.OC↗

Sharp Minimax Rates for Highly Smooth Strongly Convex Zeroth-Order Optimization

We determine the minimax dimension dependence of noisy zeroth-order optimization for globally strongly convex functions with higher-order smoothness in multilinear operator norm. Each query returns one function value with fresh independent Gaussian noise. For dimension $d$, query budget $T$, and any fixed finite smoothness order $β>2$, the minimax expected objective error satisfies $\mathcal{E}_β(T,d)=Θ_β\left(\min\left\{1,(d^2/T)^{(β-1)/β}\right\}\right)$, with dimension-independent strong-convexity, gradient-smoothness, higher-order-smoothness, noise, and minimizer-radius bounds. Thus, for sufficiently small target error $\varepsilon$, the necessary and sufficient number of noisy values is $Θ_β(d^2\varepsilon^{-β/(β-1)})$, with no logarithmic loss. The new lower bound matches the dimension and budget dependence of the upper bound of Akhavan et al. (2024), and holds for arbitrary randomized adaptive algorithms with unrestricted query points. Its key step localizes a normalized hypercube with one radial cutoff: a query can emphasize one hidden coordinate, but the aggregate information about all coordinates remains small. A self-contained multiscale spherical estimator attains the matching upper bound by cancelling lower-order bias terms. Constants in the rates depend only on the fixed smoothness order under the stated normalization.

math.OC↗

Matching Upper and Lower Bounds for Higher-Order Nonconvex Finite-Sum Optimization

We establish tight randomized higher-order oracle complexity for finding first-order stationary points of nonconvex finite sums. Let $n$ be the number of components, $Δ>0$ the initial objective-gap bound, $L_p>0$ an individual $p$-th derivative Lipschitz bound, and $ε>0$ the target gradient norm. For every fixed integer $p\ge 2$, the minimax number of exact component queries returning the value and all derivatives through order $p$, with success probability at least $2/3$, is \[ Θ_p\!\left( n+ΔL_p^{1/p}n^{1-1/(2p)} ε^{-(p+1)/p} \right), \] where the constants depend only on $p$ and the worst case ranges over all finite dimensions. The lower bound holds for unrestricted randomized adaptive algorithms and closes the $\sqrt{n}$ gap between the previously known general-order upper and lower bounds in their dependence on $n$. We extend dense weak hiding to complete higher-order replies while keeping each component's regularity independent of the chain length. The matching upper bound retains the known finite-sum exponent, requires only mean-squared $p$-th derivative increments, and removes the fixed-confidence logarithmic loss by verifying entire recursive-estimation epochs with exact function values. The characterization includes the additive $n$ term for every positive parameter regime; it counts oracle calls with unrestricted internal computation.

math.OC↗

Second-Order Stationarity with Common Random Losses: Matching Tolerance Bounds

We establish tight polynomial tolerance bounds for stochastic second-order stationarity when each fresh oracle response is a derivative of one common random scalar loss. For a population objective $F$ with Lipschitz gradient and Hessian, the target is $\|\nabla F(x)\|\le ε$ and $λ_{\min}(\nabla^2 F(x))\ge -γ$, with independent tolerances $ε,γ>0$. Under bounded gradient variance and almost-surely bounded Hessian error, the minimax number of fresh gradient or Hessian-vector-product calls is $\widetildeΘ\!\left(ε^{-3}+γ^{-5}\right)$. The characterization fixes positive gap, smoothness, and noise parameters, suppresses logarithmic factors, and allows dimension to grow within an explicit polynomial envelope. The upper bound removes the mixed term $ε^{-2}γ^{-2}$ from the earlier fresh-HVP guarantee. Direct random-line Hessian estimates and a dyadic gradient tracker separate gradient drift from randomly signed curvature motion. The lower bound realizes the endpoint costs through globally defined smooth random losses: a smooth partition localizes scalar noise without a chain-length penalty, while exact cancellation limits the information in the entire response. Consequently, the same tolerance exponents hold even for joint value, gradient, and full-Hessian responses with bounded value variance. For a population Hessian with Hölder exponent $ν\in(0,1]$, fresh gradient/HVP complexity becomes $\widetildeΘ\!\left(ε^{-3}+γ^{-(3+2/ν)}\right)$ under the corresponding dimension envelope.

math.OC↗

Near-Optimal Higher-Order Oracle Complexity for Convex--Concave Minimax Optimization

For smooth convex--concave minimax optimization, the higher-order lower bound of Chen et al. (2026) applies to a restricted tensor-algorithm class with prescribed regularized Taylor-model updates. We establish the same bound for arbitrary adaptive deterministic and randomized algorithms, matching, up to logarithmic factors, the upper bound of Zhang et al. (2026). Fix an integer $p\ge 2$ and let $L_p>0$ bound the Lipschitz constant of the objective's $p$-th derivative on a compact convex product domain of diameter at most $D_Z>0$. Each feasible query returns the objective value and all derivatives through order $p$. For accuracy $ε>0$, set $Q_{\mathrm{tan}}=L_pD_Z^p/ε$ for tangent residual and $Q_{\mathrm{gap}}=L_pD_Z^{p+1}/ε$ for saddle gap. Let $T_E^{\mathrm{det}}(ε)$ and $T_E^{\mathrm{rand}}(ε)$ denote the high-dimensional minimax query complexities for criterion $E\in\{\mathrm{tan},\mathrm{gap}\}$, with randomized success probability at least $2/3$ on every instance. Our lower bounds and the existing upper bound give $c_pQ_E^{2/(3p-1)}\le T_E^{\mathrm{rand}}(ε)\le T_E^{\mathrm{det}}(ε)\le C_pQ_E^{2/(3p-1)}[1+\log(3+Q_E)]^{6(p-1)}$ for sufficiently large $Q_E$, where $c_p,C_p>0$ depend only on $p$. Thus the same accuracy exponent holds beyond tensor update rules, even for randomized queries and arbitrary feasible outputs. The proof constructs a scalar convex--concave chain with exactly flat gates that hide complete derivative information. Direct product-domain error witnesses and adaptive transcript arguments establish the lower bounds for both criteria.

math.OC↗

Matching Lower Bounds for Randomized First-Order Methods in Hessian-Lipschitz Nonconvex Optimization

We establish a randomized first-order lower bound for finding an $ε$-stationary point, $\|\nabla f(x)\|\le ε$, of a nonconvex function with initial gap at most $Δ$, $L_1$-Lipschitz gradient, and $L_2$-Lipschitz Hessian. Each oracle call returns the exact function value and gradient. Let $Q_{\mathrm{rand},\mathrm{FO}}^{\infty}(ε;Δ,L_1,L_2)$ denote the minimax number of calls, maximized over finite dimensions, for arbitrary adaptive randomized algorithms with per-instance success probability at least $2/3$. In the regime $ε\lesssim L_1^2/L_2$ and $ΔL_2^{1/2}ε^{-3/2}\gtrsim 1$, we prove $ Q_{\mathrm{rand},\mathrm{FO}}^{\infty}(ε;Δ,L_1,L_2) \ge c\,ΔL_1^{1/2}L_2^{1/4}ε^{-7/4}, $ where $c>0$ is an absolute constant. This matches the deterministic restarted accelerated-gradient upper bound of Li and Lin (2023) and extends the sharp deterministic lower bound of Zhou (2026) to unrestricted randomized algorithms. Thus, randomization does not improve the high-dimensional worst-case query rate. The construction keeps quadratic curvature visible while revealing the forcing directions sequentially. Large eigenspaces limit preprocessing, and a scheduled-disclosure coupling makes each fresh direction incur a separate cost.

math.OC↗

Sharp Fresh-Gradient Complexity of Nonconvex-Strongly-Concave Minimax Optimization

We characterize the fresh-gradient oracle complexity of smooth nonconvex-strongly-concave minimax optimization, with matching upper and lower bounds up to logarithmic factors. Let $Φ(x)=\max_y f(x,y)$, where $f$ is jointly $L$-smooth and $μ$-strongly concave in $y$ on unconstrained Euclidean domains, and set $κ=L/μ$. Each query returns a fresh unbiased joint gradient with conditional variance at most $σ^2$. Given initial primal gap at most $Δ$ and dual residual $\|\nabla_y f(x_0,y_0)\|\le G$, the fixed-budget complexity of finding $\|\nablaΦ(\widehat x)\|\leε$ with probability at least $2/3$ is $\widetildeΘ(\sqrtκLΔ/ε^2+κLΔσ^2/ε^4+κ^2σ^2/ε^2+\sqrtκ\log_+(G/(ε\sqrtκ)))$, where $\log_+u=\log\max\{1,u\}$. This characterization holds when $κ$ and $LΔ/ε^2$ exceed universal constants, uniformly over finite dimensions, and the suppressed logarithms are independent of $G$. It establishes the necessity of linear condition-number dependence in the global stochastic cost and identifies a separate statistical refinement cost. A proximal method separates coarse primal progress from one final refinement, while the lower bounds apply to arbitrary adaptive randomized algorithms. Dual initialization contributes only an additive logarithmic cost, yet removing its control eliminates every finite dimension-free complexity bound, even with exact gradients.

math.OC↗

Near-Optimal Deterministic Exact-Value Complexity for Smooth Convex Optimization

We study the deterministic oracle complexity of smooth convex optimization when the algorithm receives only exact function values. The objective is a globally $β$-smooth convex function, all queries and the final output are restricted to the Euclidean ball of radius $R$, and the unique minimizer lies in the ball of radius $R/2$. We establish an upper bound of $O(d\sqrt{βR^2/ε})$ using coordinate finite differences together with an error-robust accelerated projected method. Our main contribution is a matching lower bound, up to the high-accuracy saturation of the construction: any deterministic adaptive value-oracle algorithm requires $Ω\!\left(d\min\{\sqrt{βR^2/ε},(d/\log(ed))^{1/3}\}\right)$ queries. Consequently, the minimax oracle complexity is $Θ(d\sqrt{βR^2/ε})$ throughout the moderate-accuracy regime $βR^2(\log(ed)/d)^{2/3}\leqε\leq cβR^2$ for a universal constant $c>0$. The lower bound must account for the fact that a single exact real value can encode arbitrarily much information. To overcome this difficulty, we construct a single fixed smooth convex hard instance using a Moreau-smoothed biased max chain, an exact prefix-shielding mechanism, and batched delayed rotations. These techniques preserve consistency with the full adaptive transcript and establish the optimality of the square-root complexity branch for deterministic bounded-query algorithms.

math.OC↗

Near-Optimal Exact-Value Zeroth-Order Complexity for Smooth Strongly Convex Optimization

We study deterministic adaptive optimization of globally $β$-smooth, $μ$-strongly convex functions using exact scalar function values. Queries and outputs lie in $B_2^d(R)$, and the minimizer lies in $B_2^d(R/2)$. Set $κ=β/μ$, $Q=βR^2/ε$, and $D_d=(d/\log(ed))^{1/3}$. For sufficiently large $d$ and $0<ε\le c_εβR^2$, the minimax value complexity $N_ε$ satisfies \[ \begin{aligned} N_ε&\ge c d\min\{\sqrt Q,\sqrtκ,D_d\},\\ N_ε&\le C d\min\left\{ \sqrt Q,\sqrtκ[1+\log_+(Q/κ)] \right\}, \end{aligned} \] where $c,C,c_ε>0$ are universal constants and $\log_+(t)=\max\{0,\log t\}$. The lower bound uses an exactly shielded smooth chain and batched delayed rotations; the upper bound combines finite differences, acceleration, and restart. When $\min\{Q,κ\}\le D_d^2$, these bounds match up to constants in the accuracy-dominated regime $Q\leκ$ and at constant relative accuracy $ε=Θ(μR^2)$. For arbitrarily higher accuracy in the range $κ\le D_d^2$, the bounds differ by at most $1+\log(μR^2/ε)$; the optimal accuracy dependence remains unresolved in general.

math.OC↗

A Self-Triggered Agentic Push Recommendation System

Push notification is a critical recommendation scenario on large-scale platforms, allowing the system to proactively reach users outside the application to improve long-term re-engagement. However, designing an optimal push system requires handling a complex action space for the "whether and when" delivery problem under strict system resource constraints. Existing solutions typically fall into two passive paradigms: pre-planned frequency methods that allocate delivery times via offline modeling, limiting real-time adaptability; and fixed-interval triggering methods that periodically poll the system, creating a strict dilemma between excessive computational overhead and diminished optimal timing capture. Furthermore, such multi-stage frameworks severely suffer from local optima. To overcome these limitations, in this paper, we propose STEPS, a proactive, Self-Triggered End-to-end Agentic Push Recommendation System, which is already fully deployed at Douyin with over 1 billion users. STEPS reformulates push recommendation as a self-triggered agentic process in which the system decides not only whether to send a push, but also when to invoke itself again, thereby forming a closed loop that balances real-time effectiveness and efficiency. Specifically, STEPS consists of two decision transformer-based agents: a planning agent that schedules the next system invocation using a gated ordinal regression method, and an execution agent that decides whether to send a push based on trajectory rewards. Furthermore, we introduce a lightweight filtering agent to both control computational overhead and act as a crucial safeguard against unreasonable planning behaviors. Online A/B testing demonstrates that STEPS significantly increases user active days by 0.2843% and reduces the push permission disablement rate by 1.9089%, while the filtering agent reduces computational overhead by 79.42%.

cs.IR↗

Beyond Independent Manipulation: Individual Fairness-aware Strategic Classification with Peer Imitation

Strategic classification (SC) investigates scenarios where agents manipulate their features to obtain favorable decisions from predictive models. Existing fairness-aware SC approaches primarily focus on group fairness and typically assume that agents respond independently. However, when individual fairness is required, ensuring similar individuals receive similar outcomes, agents' manipulation becomes interdependent: an agent's preferred manipulation depends on the neighborhoods' outcomes. This induces a mismatch between classical SC formulations and fairness-aware decision settings, where independent models no longer accurately characterize strategic manipulations. To address this issue, we introduce individual fairness-aware strategic classification (IFSC), a framework that models peer-driven manipulation arising from individual fairness, where agents imitate nearby positively decided peers to obtain favorable outcomes. IFSC characterizes strategic manipulation as similarity-based imitation toward visible accepted peers and learns classifiers under the resulting post-manipulation distributions. To account for uncertainty in peer observability, IFSC employs a robust learning process that introduces stochastic perturbations during manipulation simulation. Experiments on synthetic and real-world datasets demonstrate that IFSC improves individual-fairness consistency and mitigates imitation-induced distortions.

cs.LG↗

When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach

Tabular foundation models based on pretrained prior-data fitted networks~(PFNs) have shown strong generalization on diverse tabular tasks, but they are typically designed for \emph{non-strategic} settings where data distributions are independent of deployed classifiers. In many real-world decision scenarios, however, individuals may strategically modify their features after deployment to obtain favorable outcomes, inducing a post-deployment distribution shift. This paper studies whether PFN-style tabular foundation models can generalize to such \emph{strategic} tabular data. We show that strategic manipulation creates a mismatch between the non-strategic prior learned during pretraining and the post-manipulation strategic prior, which leads to systematic prediction bias. To address this issue, we propose \textbf{Strategic Prior-data Fitted Network}~\textit{(SPN)}, an inference-time strategy-aware framework that adapts tabular foundation models to strategic environments without retraining. SPN constructs strategic in-context examples to approximate post-manipulation inputs and aligns PFN predictions with the induced strategic distribution. Experiments on real-world and synthetic tabular datasets show that SPN consistently improves robustness and predictive performance under strategic manipulation compared with both tabular foundation models and classical tabular methods.

cs.AI↗

Beyond Rational Illusion: Behaviorally Realistic Strategic Classification

Strategic classification(SC) studies the interaction between decision models and agents who strategically manipulate their features for favorable outcomes. Existing SC frameworks typically rely on the idealized assumption that agents are strictly rational. However, evidence from behavioral economics and psychology consistently shows that real-world decision-making is often shaped by cognitive biases, deviating from pure rationality. To formalize this limitation, we identify and define a new problem setting, termed the behaviorally realistic strategic classification problem, where agents' strategic manipulations deviate from full rationality due to psychological biases. Motivated by the identified limitation, we propose the Prospect-Guided Strategic Framework (Pro-SF) to address the problem, a principled framework grounded in prospect theory to model and learn under behaviorally realistic strategic responses. Specifically, to capture behaviorally realistic strategic manipulations, our framework reformulates the Stackelberg-style interaction between agents and the decision-maker by incorporating three key mechanisms inspired by prospect theory, including the asymmetry between benefits and costs, different subjective reference points, and non-rational probability distortion. Experiments on synthetic and real-world datasets establish Pro-SF as a behaviorally grounded approach to strategic classification, bridging machine learning and behavioral economics for more reliable deployment in the real world.

cs.AI↗

PAFO: Pareto Fairness Optimization for Personalized Reward Modeling

Large language models (LLMs) increasingly rely on reward models to align their outputs with diverse user preferences. While personalized reward models aim to capture such heterogeneity, they are often trained on imbalanced user preference data and may therefore favor users whose preferences are more common in the training population. In this paper, we identify this failure mode as personalized reward bias, where reward modeling quality varies systematically with preference support rate. We formulate its mitigation as a Pareto fairness problem over group utilities, aiming to improve under-served users without degrading other user groups. To this end, we propose PAFO, a Pareto fairness optimization framework for personalized reward modeling. PAFO first trains group-specialized reward models for majority and minority preference groups, then constructs conditional margin-level supervision to distill their heterogeneous preference boundaries into a single unified model. The resulting model uses group information only during training and requires no explicit group labels at inference time. Experiments on Personal-LLM and DSP show that PAFO improves both minority-group and majority-group accuracy while reducing user-level unfairness across multiple metrics, demonstrating its effectiveness for fairer LLM personalization.

cs.AI↗