Search arXivSearch

arXiv · 2505.04553

Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions

Abstract

We propose a reinforcement learning (RL) framework under a broad class of risk objectives, characterized by convex scoring functions. This class covers many common risk measures, such as variance, Expected Shortfall, entropic Value-at-Risk, and mean-risk utility. To resolve the time-inconsistency issue, we consider an augmented state space and an auxiliary variable and recast the problem as a two-state optimization problem. We propose a customized Actor-Critic algorithm and establish some theoretical approximation guarantees. A key theoretical contribution is that our results do not require the Markov decision process to be continuous. Additionally, we propose an auxiliary variable sampling method inspired by the alternating minimization algorithm, which is convergent under certain conditions. We validate our approach in simulation experiments with a financial application in statistical arbitrage trading, demonstrating the effectiveness of the algorithm.

Explore related subjects

Keep this discovery

BibTeXRIS

Shanyu Han, Yang Liu, Xiang Yu. 2025-05-07. Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions. https://arxiv.org/abs/2505.04553

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Variance-Optimal Hedging in the Rough Hawkes--Heston Model

We study variance-optimal stock hedging and the convergence of approximate strategies in the rough Hawkes--Heston model. Starting from the model's affine conditional transform and the affine Volterra jump framework, we obtain semi-explicit hedges for European calls and a representation of the minimum quadratic error through the Galtchouk--Kunita--Watanabe projection. Our main approximation result keeps the original stock, variance driver, and information flow fixed while regularizing the kernel used to evaluate the hedge. To handle singular memory and common marked jumps, we construct the approximate holdings from histories available before trading and preserve the conditional transform's random modulus envelope. Riccati--Volterra stability and weighted truncation then yield convergence in the original stock's trading norm on compact Fourier intervals. For calls, a joint choice of kernel regularization and Fourier cutoff gives convergence of the initial capitals and strategies, uniform-in-time square-mean convergence of continuous-time gains, and convergence of the terminal mean-square error to the variance-optimal value. A numerical experiment with shifted fractional kernels illustrates the construction on common original-market paths.

q-fin.MF

Numeraire Invariance of Entropy-Projected Martingale Measures

Let \(P\) be a fixed physical law and let \(Q\) be an equivalent martingale measure selected from the martingale-measure set associated with a chosen numeraire. A change of numeraire maps \(Q\) to \(T_LQ\), where \(d(T_LQ)=L\,dQ\) and \(L\) is the terminal likelihood ratio. The forward relative-entropy projection minimizing \(D_{\mathrm{KL}}(P\Vert Q)\) commutes with this transform because its objective changes only by the constant \(-E_P\log L\). The minimal entropy martingale measure (MEMM) orientation \(D_{\mathrm{KL}}(Q\Vert P)\) does not have this property, and a trinomial counterexample shows that independently recomputed MEMMs need not be likelihood compatible. We make two economic consequences explicit. First, the two entropy orientations are precisely the \(Q\)-dependent terms in the classical convex-dual objectives for logarithmic and exponential utility, respectively. Second, likelihood compatibility is equivalent to equality of the pricing functionals obtained in the two numeraires. Hence the forward selectors value every integrable claim consistently across numeraires, whereas the two MEMMs in the counterexample assign different prices to a nonreplicable digital claim. We also prove a finite-state class-level characterization: uniform invariance over the elementary one-period likelihood-ratio families forces a smooth convex \(f\)-divergence to be logarithmic, up to scaling and affine equivalence. Finally, in finite-state markets, the forward projection exists under the usual strictly positive feasible-point condition; its density \(dP/dQ^*\) is attainable log-optimal terminal wealth, and the minimum forward entropy equals maximal expected log growth.

q-fin.MF

The Delta of a Variance Swap

We define the variance swap delta as the sensitivity of the price of variance to a change in underlying price. We use Carr-Madan spanning formulas to analyze this sensitivity when the implied volatility smile curve may depend on the underlying price. We show that the variance swap total delta is zero for the class of smile curves that are pure functions of (log) moneyness, which goes against the empirical observation that variance is up when the market is down. We propose a simple modification of the smile to correct this issue.

q-fin.MF