Search arXiv⌕ Search

arXiv · 2610.03638

The Minimax Cost of Smoothness in Fully First-Order Stochastic Bilevel Optimization

Abstract

We study the sample complexity of finding a point with expected hypergradient norm at most $ε>0$ in nonconvex--strongly-convex stochastic bilevel optimization using fresh, globally unbiased first-order samples with bounded variance. For lower-variable smoothness order $p\ge1$, F$^2$SA-$p$ achieves $\widetilde O(pε^{-4-2/p})$ (Chen et al., 2026), but necessity of the exponent $4+2/p$ was open. We prove matching upper and lower bounds $Θ_p(ε^{-4-2/p})$ under fixed nondegenerate regularity budgets, with constants allowed to depend on $p$. The lower bound holds for unrestricted randomized algorithms under a fixed sample budget, allows dimension to grow with accuracy, and preserves global unbiasedness and all prescribed lower-variable smoothness bounds, even with a scalar lower variable and exact upper gradients. We further characterize the minimax fixed-budget complexity under $L_j=MΛ^j(j!)^β$ for $j\ge3$: $Q_{p,β}^*(ε)=Θ\left(ε^{-4}\left[\inf_{1\le r\le p,,r\in\mathbb N} r^βε^{-1/r}\right]^2\right)$, for $p\in\mathbb N\cup{\infty}$, with constants independent of $p$. At infinite order, this gives $Θ(ε^{-4})$ for $β=0$ and $Θ(ε^{-4}\log^2(1/ε))$ for $β=1$. A gradient-difference method with weighted independent batches attains these bounds uniformly in order. Thus quantitative smoothness yields precise gains, whereas qualitative analyticity alone still permits $Θ(ε^{-6})$ complexity.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wendao Wu, Haihan Zhang, Chenheng Zhang, Yanyi Li, Chunyuan Zheng, Cong Fang, Haoxuan Li, Zhouchen Lin. 2026-10-02. The Minimax Cost of Smoothness in Fully First-Order Stochastic Bilevel Optimization. https://arxiv.org/abs/2610.03638

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Non-convex optimal control governed by nonlinear elliptic partial differential equation: Analysis, functional-based algorithm and control landscape

Non-convex optimal control arises from various applications but may contain multiple stationary points. Classical solvers usually perform a local search and therefore rely on good initial guesses to reach appropriate local optimal controls. In this work we introduce a novel solution strategy for the non-convex optimal control of an elliptic equation, the main idea of which is to construct the control landscape based on the functional high-index saddle dynamics (FHiSD) method. This method reduces the dependence on prescribed initial guesses by using the information of high-index saddle points, and depicts the macroscopic configuration of the space of the control variable. Then various minima could be systematically computed along transition pathways and the control strategy can then be selected among them. We prove the existence of the optimal control, and then propose and analyze the FHiSD in locating saddle points. Subsequently, the applicability of FHiSD in non-convex optimal control is rigorously justified. Numerical results not only indicate the effectiveness of the proposed method, but reveal unintuitive phenomena (e.g. the non-monotonicity of the values of the cost functional with respect to the Morse indices) that support the necessity of computing multiple solutions of high indices.

math.OC↗

Mathematical model for sustainable fisheries resource management accounting for size spectrum

This paper proposes a novel modelling and control framework for growth models that incorporate a size spectrum in conjunction with numerical computation and extensive field surveys. In fisheries management, the size spectrum, characterized by individual differences in body weight and length, is a critical factor, as it influences the physiology and ecology of fish, as well as the preferences of anglers. However, a comprehensive theoretical framework for fisheries modelling and management that accounts for the size spectrum has yet to be established. We apply a growth model that considers the size spectrum to Plecoglossus altivelis altivelis (Ayu), an important inland fisheries resource in Japan. Additionally, we introduce a novel stochastic control theory for the resource management of Ayu, taking its size spectrum into account. The growth model is calibrated using data collected annually from a river system in Japan. Our control problem addresses the size spectrum of fishing benefits and terminal utility (nonlinear expectation) for sustainability, resulting in a nonstandard problem to which the dynamic programming principle does not apply. We address this difficulty using a time-inconsistent formalism, where solving the control problem is reduced to finding an appropriate solution to a system of nonlinear partial differential equations. We numerically compute the system using the finite difference method and explore the fisheries management of Ayu at the study site.

math.OC↗

The Value of Information in Resource-Constrained Pricing

Firms that price perishable resources -- airline seats, hotel rooms, seasonal inventory -- now routinely use demand predictions, but these predictions vary widely in quality. Under hard capacity constraints, acting on an inaccurate prediction can irreversibly deplete inventory needed for future periods. We study how prediction uncertainty propagates into dynamic pricing decisions with linear demand, stochastic noise, and finite capacity. A certified demand forecast with known error bound~$ε^0$ specifies where the system should operate: it shifts regret from $O(\sqrt{T})$ to $O(\log T)$ when $ε^0 \lesssim T^{-1/4}$, and we prove this threshold is tight. A misspecified surrogate model -- biased but correlated with true demand -- cannot set prices directly but reduces learning variance by a factor of $(1-ρ^2)$ through control variates. The two mechanisms compose: the forecast determines the regret regime; the surrogate tightens estimation within it. All algorithms rest on a boundary attraction mechanism that stabilizes pricing near degenerate capacity boundaries without requiring non-degeneracy assumptions. Experiments confirm the phase transition threshold, the variance reduction from surrogates, and robustness across problem instances.

math.OC↗