Search arXivSearch

arXiv · 2604.22188

Optimal Investment and Entropy-Regularized Learning Under Stochastic Volatility Models with Portfolio Constraints

Abstract

We study the problem of optimal portfolio selection under stochastic volatility within a continuous time reinforcement learning framework with portfolio constraints. Exploration is modeled through entropy-regularized relaxed controls, where the investor selects probability distributions over admissible portfolio allocations rather than deterministic strategies. Using dynamic programming arguments, we derive the associated entropy-regularized Hamilton-Jacobi-Bellman equation, whose Hamiltonian involves optimization over probability measures supported on a compact control set. We show that the optimal exploratory policy takes the form of a truncated Gaussian distribution characterized by spatial derivatives of the solution of the resulting nonlinear quasilinear parabolic partial differential equation. Under suitable structural conditions on the model coefficients, we prove the existence of classical solutions to this nonlinear HJB equation for the value function. We then establish a verification theorem and analyze the policy-improvement structure induced by the entropy-regularized Hamiltonian, showing how the resulting sequence of PDEs provides a continuous-time interpretation of actor-critic learning dynamics. Finally, our PDE analysis with a semi-closed form of optimal value and optimal policy enables the design of an implementable reinforcement learning algorithm by recasting the optimal problem in a martingale framework.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Thai Nguyen, Pertiny Nkuize. 2026-04-24. Optimal Investment and Entropy-Regularized Learning Under Stochastic Volatility Models with Portfolio Constraints. https://arxiv.org/abs/2604.22188

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Modeling interest rate swap volatility with GARCH processes

We examine the conditional volatility dynamics of the USD 1Yx10Y forward swap rate using GARCH(1,1), GJR-GARCH(1,1), and a two-regime Markov-switching GARCH (MSGARCH) model. The analysis uses daily data from 2007 to 2023 and incorporates market-implied measures (ATM swaption volatility and the SRVIX in- dex) together with a broad set of diagnostic tests. Standard GARCH and GJR- GARCH models show stable short-run parameters, but the intercept ω varies markedly across rolling windows, causing instability in the implied long-run vari- ance. This pattern, confirmed by the Nyblom test, motivates adopting a regime- switching specification. MSGARCH mitigates this issue by keeping regime-specific parameters stable and capturing time variation through filtered regime probabili- ties. It delivers the highest log-likelihood and lowest AIC, whereas BIC favours the more parsimonious GJR-GARCH. One-step-ahead backtesting indicates comparable short-horizon accuracy across models, but MSGARCH offers a clearer structural in- terpretation by isolating high- and low-volatility regimes aligned with major market events.

q-fin.MF

Optimal Investment and Consumption in Financial Markets with Integrated Variance Clocks

We study the infinite-horizon optimal investment and consumption problem in a general class of continuous financial markets, where uncertainty is driven by a continuous non-decreasing stochastic clock representing accumulated variance. This framework encompasses classical Markovian and non-Markovian stochastic volatility models as well as singular realized-variance models in which no spot volatility process exists. We characterize the value process and optimal investment and consumption strategies in terms of a non-linear infinite-horizon backward stochastic differential equation driven jointly by calendar time and the stochastic clock. We develop a general well-posedness theory for this new class of IVC-BSDEs based on the method of sub- and supersolutions, establishing existence, uniqueness, and stability under natural conditions that might be of independent interest beyond the financial application at hand. We are moreover able to identify the sign of the $Z$-component of the solution using Malliavin calculus. We then apply our results to Volterra Heston models with locally integrable kernels, covering both rough and hyper-rough regimes. Exploiting the affine structure of the model, we verify the optimality of the candidate strategies in incomplete markets and obtain an explicit representation of the solution in the complete market case. Owing to the generality of the framework and the weak assumptions imposed on the stochastic clock, our results unify and extend several existing results for optimal investment and consumption, including classical Markovian stochastic volatility models.

q-fin.MF

Liquidity Provision and Rebate Design in Option Markets

We provide a model for the nested optimisation problem of market making and rebate design problems in option markets and find optimal strategies. A single market maker trades multiple European call options in a local-stochastic volatility option market with both make and take strategies, modeled, respectively, as continuous and impulse controls. Her objective is to maximize, over all admissible make-take strategies, net profit of option portfolio value and cumulative rebate revenue, subject to a penalty on residual portfolio delta and vega. In addition, we demonstrate how an exchange can incentivize a market maker to improve market liquidity by setting suitable fee rebates, thereby resolving its own liquidity attraction problem. To this end, we propose a three-step rebate design scheme with flexibility to accommodate specific liquidity targets imposed by an exchange. Numerical results are provided to validate the effectiveness of the proposed scheme.

q-fin.MF