Search arXiv⌕ Search

arXiv · 2609.31380

Optimization with Region-Reduced ReLU Neural Networks

Abstract

Optimization of mathematical models involving integer decisions and neural networks with ReLU activation (ReLU ANNs) is a challenging task. Nevertheless, such models are an enabling technology in many application domains. A prominent example is superstructure optimization in chemical engineering, where ReLU ANNs are frequently employed as surrogate models for complex nonlinear processes. We survey recent developments in this area. We argue that in addition to network size and training options of the ANNs, the ReLU activation geometry and the number of linear regions on the domain of interest have a strong impact on computational optimization performance. While standard model compression approaches such as structured pruning reduce network size, they do not explicitly address geometric considerations. Therefore, we propose a novel \emph{region-reduced} model compression approach that combines the stabilization of unstable neurons and the merging of redundant neurons to reduce the number of linear regions while maintaining predictive accuracy through error compensation. We evaluate our method against standard compression approaches on multiple optimization use cases. First, the two-dimensional peaks function for which we can visualize the activation geometry. Second, on optimization over individual surrogate ReLU ANNs for three chemical processes, and third, on a hybrid superstructure optimization problem that involves the three ReLU ANNs, additional process submodels, and binary variables. The results for the superstructure problem demonstrate the large potential of region reduction with a decrease of 40\% to 50\% in computational time, yielding solutions closer than 1\% to the reference at negligible effort of obtaining the compressed model.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Christoph Plate, Caroline Ganzer, Mirko Hahn, Alexander Klimek, Heyuan Liu, Sebastian Sager, Kai Sundmacher, Hanna Wilhelm. 2026-09-25. Optimization with Region-Reduced ReLU Neural Networks. https://arxiv.org/abs/2609.31380

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Revisiting Inexact Fixed-Point Iterations for Min-Max Problems: Stochasticity and Structured Nonconvexity

We focus on constrained, $L$-smooth, potentially stochastic and nonconvex-nonconcave min-max problems either satisfying $ρ$-cohypomonotonicity or admitting a solution to the $ρ$-weakly Minty Variational Inequality (MVI), where larger values of the parameter $ρ>0$ correspond to a greater degree of nonconvexity. These problem classes include examples in two player reinforcement learning, interaction dominant min-max problems, and certain synthetic test problems on which classical min-max algorithms fail. It has been conjectured that first-order methods can tolerate a value of $ρ$ no larger than $\frac{1}{L}$, but existing results in the literature have stagnated at the tighter requirement $ρ< \frac{1}{2L}$. With a simple argument, we obtain optimal or best-known complexity guarantees with cohypomonotonicity or weak MVI conditions for $ρ< \frac{1}{L}$. First main insight for the improvements in the convergence analyses is to harness the recently proposed $\textit{conic nonexpansiveness}$ property of operators. Second, we provide a refined analysis for inexact Halpern iteration that relaxes the required inexactness level to improve some state-of-the-art complexity results even for constrained stochastic convex-concave min-max problems. Third, we analyze a stochastic inexact Krasnosel'ski\uı-Mann iteration with a multilevel Monte Carlo estimator when the assumptions only hold with respect to a solution.

math.OC↗

Towards Weaker Variance Assumptions for Stochastic Optimization

We revisit a classical assumption for analyzing stochastic gradient algorithms where the squared norm of the stochastic subgradient (or the variance for smooth problems) is allowed to grow as fast as the squared norm of the optimization variable. We contextualize this assumption in view of its inception in the 1960s, its seemingly independent appearance in the recent literature, its relationship to weakest-known variance assumptions for analyzing stochastic gradient algorithms, and its relevance in deterministic problems for non-Lipschitz nonsmooth convex optimization. We build on and extend a connection recently made between this assumption and the Halpern iteration. For convex nonsmooth, and potentially stochastic, optimization, we analyze horizon-free, anytime algorithms with last-iterate rates. For problems beyond simple constrained optimization, such as convex problems with functional constraints or regularized convex-concave min-max problems, we obtain rates for optimality measures that do not require boundedness of the feasible set.

math.OC↗

Spare Strategy Analysis and Design for Mega Satellite Constellations Using Markov Chain

This paper presents a Markov-chain-based method for the early-phase analysis and design of spare-management architectures for large-scale satellite constellations. To assess the long-run viability of such concepts of operations, satellite failure and replenishment processes are modeled as Markov chains and analyzed through their stationary solution. We reinvestigate an indirect spare strategy, modeled as a multi-echelon periodic-review reorder-point/order-quantity policy, in which spares are first delivered to parking orbits and then transferred to constellation planes. The stock levels in constellation and parking orbits are each modeled as independent Markov chains, and a fixed-point iteration yields a consistent joint stationary solution that describes the strategy's average behavior. This approach accurately captures the stochastic interplay within a multi-echelon model driven by orbital mechanics, avoiding the aggregation assumptions of prior works and remaining valid across a wider operating domain. Building on this fast, accurate analysis, we formulate an optimization problem and solve it via a genetic algorithm. Finally, we demonstrate the practical value of both the analysis method and the optimization framework in a real-world mega-constellation case study.

math.OC↗