Search arXivSearch

arXiv · 2004.01783

Directional necessary optimality conditions for bilevel programs

Abstract

The bilevel program is an optimization problem where the constraint involves solutions to a parametric optimization problem. It is well-known that the value function reformulation provides an equivalent single-level optimization problem but it results in a nonsmooth optimization problem which never satisfies the usual constraint qualification such as the Mangasarian-Fromovitz constraint qualification (MFCQ). In this paper we show that even the first order sufficient condition for metric subregularity (which is in general weaker than MFCQ) fails at each feasible point of the bilevel program. We introduce the concept of directional calmness condition and show that under {the} directional calmness condition, the directional necessary optimality condition holds. {While the directional optimality condition is in general sharper than the non-directional one,} the directional calmness condition is in general weaker than the classical calmness condition and hence is more likely to hold. {We perform the directional sensitivity analysis of the value function and} propose the directional quasi-normality as a sufficient condition for the directional calmness. An example is given to show that the directional quasi-normality condition may hold for the bilevel program.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kuang Bai, Jane Ye. 2020-11-18. Directional necessary optimality conditions for bilevel programs. https://arxiv.org/abs/2004.01783

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Nonsmooth Newton methods with effective subspaces for polyhedral regularization

We propose several new nonsmooth Newton methods for solving convex composite optimization problems with polyhedral regularizers, while avoiding the computation of complicated second-order information on these functions. Under the tilt-stability condition at the optimal solution, these methods achieve the quadratic convergence rates expected of Newton schemes. Numerical experiments on Lasso, generalized Lasso, OSCAR-regularized least-square problems, and an image super-resolution task illustrate both the broad applicability and the accelerated convergence profile of the proposed algorithms, in comparison with first-order and several recently developed nonsmooth Newton schemes.

math.OC

Non-Asymptotic Global Convergence of PPO-Clip

Reinforcement learning has gained attention for modern Large Language Model post-training. The actor-only variants of Proximal Policy Optimization (PPO) are widely applied for their efficiency. These algorithms incorporate a clipping mechanism to improve stability. Besides, a regularization term, such as the reverse KL-divergence or a more general \(f\)-divergence, is introduced to control excessive deviation from a reference policy. Despite their empirical success, a rigorous theoretical understanding of the problem and the algorithm's properties is limited. This paper advances the theoretical foundations of the PPO-Clip algorithm by analyzing a deterministic actor-only PPO algorithm within the general RL setting with \(f\)-divergence regularization under the softmax policy parameterization. We derive a non-uniform Lipschitz smoothness condition and a Łojasiewicz inequality for the considered problem. Based on these properties, we establish non-asymptotic global linear convergence in value gap for the forward KL regularizer. For the reverse KL regularizer, we derive global linear convergence from any finite softmax initialization in both value gap and squared policy distance.

math.OC

An Occupation-Measure and Frank-Wolfe Framework for Heterogeneous Mean-Field Control

Heterogeneous mean-field control (MFC) involves multiple interacting populations with distinct dynamics, control constraints, and coupling structures. We develop a heterogeneous occupation-measure mean-field control (OM-MFC) framework that reformulates the population-distribution control problem as an optimization problem over occupation measures. In this representation, potentially nonlinear dynamics enter through linear weak Liouville constraints, while nonlocal inter-population interactions become structured bilinear functionals of the occupation-measure marginals. Exploiting this structure, we derive a Frank--Wolfe (FW) method and show that its linear minimization oracle decomposes into independent population-wise optimal control problems once the current population distributions are fixed, enabling parallel computation without an a priori discretization of the measure space. We further establish a sufficient convexity condition based on the symmetrized matrix-valued interaction kernel, under which Frank--Wolfe with an exact oracle admits the standard $O(1/k)$ convergence rate. Numerical examples on heterogeneous UAV coordination and three-dimensional search-and-rescue illustrate the framework both within the certified convex regime and for directional interactions beyond that regime.

math.OC