Search arXiv⌕ Search

arXiv · 2610.03506

Optimality of delayed reward attainment under uncertainty

Abstract

Decision making under uncertainty requires balancing the value of additional information against the cost of delaying the decision until that information becomes available. This challenge is especially prevalent for embodied decisions such as those faced by animals moving between food sources and robots selecting navigation targets, where movement through space shapes both the quality of information and the cost of acting. We consider a version of this problem, where the decision is on choosing among rewards of uncertain value, and those rewards can only be attained through changing the decision maker's state sufficiently. Our focus is on interactions between reward learning and state evolution, restated as a deterministic optimal control problem with two terminal alternatives, one with a known reward and the other learned gradually through noisy observations. We demonstrate that delayed reward attainment is often optimal in a variety of systems, especially so when the alternatives are initially difficult to distinguish and the cost of waiting is relatively low. For a subset of problems, we also derive sufficient conditions that allow shortening the planning horizon a priori, guaranteeing that further observations would not affect the value function. Finally, when observation quality depends on the state, we show that it is often optimal for a controller to evolve the state strategically toward more informative sensing locations before making a decision, in line with the behavior broadly observed in biological and engineered systems.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chuwei Wang, Alexander Vladimirsky, Anastasia Bizyaeva. 2026-10-02. Optimality of delayed reward attainment under uncertainty. https://arxiv.org/abs/2610.03506

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Non-convex optimal control governed by nonlinear elliptic partial differential equation: Analysis, functional-based algorithm and control landscape

Non-convex optimal control arises from various applications but may contain multiple stationary points. Classical solvers usually perform a local search and therefore rely on good initial guesses to reach appropriate local optimal controls. In this work we introduce a novel solution strategy for the non-convex optimal control of an elliptic equation, the main idea of which is to construct the control landscape based on the functional high-index saddle dynamics (FHiSD) method. This method reduces the dependence on prescribed initial guesses by using the information of high-index saddle points, and depicts the macroscopic configuration of the space of the control variable. Then various minima could be systematically computed along transition pathways and the control strategy can then be selected among them. We prove the existence of the optimal control, and then propose and analyze the FHiSD in locating saddle points. Subsequently, the applicability of FHiSD in non-convex optimal control is rigorously justified. Numerical results not only indicate the effectiveness of the proposed method, but reveal unintuitive phenomena (e.g. the non-monotonicity of the values of the cost functional with respect to the Morse indices) that support the necessity of computing multiple solutions of high indices.

math.OC↗

Mathematical model for sustainable fisheries resource management accounting for size spectrum

This paper proposes a novel modelling and control framework for growth models that incorporate a size spectrum in conjunction with numerical computation and extensive field surveys. In fisheries management, the size spectrum, characterized by individual differences in body weight and length, is a critical factor, as it influences the physiology and ecology of fish, as well as the preferences of anglers. However, a comprehensive theoretical framework for fisheries modelling and management that accounts for the size spectrum has yet to be established. We apply a growth model that considers the size spectrum to Plecoglossus altivelis altivelis (Ayu), an important inland fisheries resource in Japan. Additionally, we introduce a novel stochastic control theory for the resource management of Ayu, taking its size spectrum into account. The growth model is calibrated using data collected annually from a river system in Japan. Our control problem addresses the size spectrum of fishing benefits and terminal utility (nonlinear expectation) for sustainability, resulting in a nonstandard problem to which the dynamic programming principle does not apply. We address this difficulty using a time-inconsistent formalism, where solving the control problem is reduced to finding an appropriate solution to a system of nonlinear partial differential equations. We numerically compute the system using the finite difference method and explore the fisheries management of Ayu at the study site.

math.OC↗

The Value of Information in Resource-Constrained Pricing

Firms that price perishable resources -- airline seats, hotel rooms, seasonal inventory -- now routinely use demand predictions, but these predictions vary widely in quality. Under hard capacity constraints, acting on an inaccurate prediction can irreversibly deplete inventory needed for future periods. We study how prediction uncertainty propagates into dynamic pricing decisions with linear demand, stochastic noise, and finite capacity. A certified demand forecast with known error bound~$ε^0$ specifies where the system should operate: it shifts regret from $O(\sqrt{T})$ to $O(\log T)$ when $ε^0 \lesssim T^{-1/4}$, and we prove this threshold is tight. A misspecified surrogate model -- biased but correlated with true demand -- cannot set prices directly but reduces learning variance by a factor of $(1-ρ^2)$ through control variates. The two mechanisms compose: the forecast determines the regret regime; the surrogate tightens estimation within it. All algorithms rest on a boundary attraction mechanism that stabilizes pricing near degenerate capacity boundaries without requiring non-degeneracy assumptions. Experiments confirm the phase transition threshold, the variance reduction from surrogates, and robustness across problem instances.

math.OC↗