Search arXiv⌕ Search

arXiv · 2610.04285

Amortized Score-Hamiltonian Policy Iteration: A Grid-Free Scheme for Relaxed Stochastic Control Problems

Abstract

We develop an amortized, grid-free implementation of continuous Langevin dynamics based policy-value iteration for entropy-regularized, infinite-horizon relaxed stochastic control problems. The improvement rate of the exact iteration is a discounted aggregate of relative Fisher information between the policy and the Gibbs law of its Hamiltonian. The associated score residual is the velocity with which the control's Langevin dynamics transport its law. We project this velocity onto a conditional sampler shared across states, instead of one Langevin dynamics per state, and the value dynamics onto a parametric critic, estimating both projections at sampled states to obtain coupled actor--critic flows. The score loss measures the actor's agreement with the current critic, while the policy-evaluation residual measures the critic's agreement with the actor. We also derive gradient and Hessian residuals, including a Feynman--Kac representation for the gradient equation, to control errors not detected by the projected value iteration. An exact decomposition of the HJB residual combines these errors into a policy-suboptimality bound under verification and logarithmic Sobolev assumptions. In the linear-quadratic class, both projections are exact and recover the pointwise iteration, and we provide numerical experiments on general models to demonstrate the coupled actor--critic learning in high-dimensions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Qi Feng, Gu Wang. 2026-10-03. Amortized Score-Hamiltonian Policy Iteration: A Grid-Free Scheme for Relaxed Stochastic Control Problems. https://arxiv.org/abs/2610.04285

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Non-convex optimal control governed by nonlinear elliptic partial differential equation: Analysis, functional-based algorithm and control landscape

Non-convex optimal control arises from various applications but may contain multiple stationary points. Classical solvers usually perform a local search and therefore rely on good initial guesses to reach appropriate local optimal controls. In this work we introduce a novel solution strategy for the non-convex optimal control of an elliptic equation, the main idea of which is to construct the control landscape based on the functional high-index saddle dynamics (FHiSD) method. This method reduces the dependence on prescribed initial guesses by using the information of high-index saddle points, and depicts the macroscopic configuration of the space of the control variable. Then various minima could be systematically computed along transition pathways and the control strategy can then be selected among them. We prove the existence of the optimal control, and then propose and analyze the FHiSD in locating saddle points. Subsequently, the applicability of FHiSD in non-convex optimal control is rigorously justified. Numerical results not only indicate the effectiveness of the proposed method, but reveal unintuitive phenomena (e.g. the non-monotonicity of the values of the cost functional with respect to the Morse indices) that support the necessity of computing multiple solutions of high indices.

math.OC↗

Mathematical model for sustainable fisheries resource management accounting for size spectrum

This paper proposes a novel modelling and control framework for growth models that incorporate a size spectrum in conjunction with numerical computation and extensive field surveys. In fisheries management, the size spectrum, characterized by individual differences in body weight and length, is a critical factor, as it influences the physiology and ecology of fish, as well as the preferences of anglers. However, a comprehensive theoretical framework for fisheries modelling and management that accounts for the size spectrum has yet to be established. We apply a growth model that considers the size spectrum to Plecoglossus altivelis altivelis (Ayu), an important inland fisheries resource in Japan. Additionally, we introduce a novel stochastic control theory for the resource management of Ayu, taking its size spectrum into account. The growth model is calibrated using data collected annually from a river system in Japan. Our control problem addresses the size spectrum of fishing benefits and terminal utility (nonlinear expectation) for sustainability, resulting in a nonstandard problem to which the dynamic programming principle does not apply. We address this difficulty using a time-inconsistent formalism, where solving the control problem is reduced to finding an appropriate solution to a system of nonlinear partial differential equations. We numerically compute the system using the finite difference method and explore the fisheries management of Ayu at the study site.

math.OC↗

The Value of Information in Resource-Constrained Pricing

Firms that price perishable resources -- airline seats, hotel rooms, seasonal inventory -- now routinely use demand predictions, but these predictions vary widely in quality. Under hard capacity constraints, acting on an inaccurate prediction can irreversibly deplete inventory needed for future periods. We study how prediction uncertainty propagates into dynamic pricing decisions with linear demand, stochastic noise, and finite capacity. A certified demand forecast with known error bound~$ε^0$ specifies where the system should operate: it shifts regret from $O(\sqrt{T})$ to $O(\log T)$ when $ε^0 \lesssim T^{-1/4}$, and we prove this threshold is tight. A misspecified surrogate model -- biased but correlated with true demand -- cannot set prices directly but reduces learning variance by a factor of $(1-ρ^2)$ through control variates. The two mechanisms compose: the forecast determines the regret regime; the surrogate tightens estimation within it. All algorithms rest on a boundary attraction mechanism that stabilizes pricing near degenerate capacity boundaries without requiring non-degeneracy assumptions. Experiments confirm the phase transition threshold, the variance reduction from surrogates, and robustness across problem instances.

math.OC↗