Search arXivSearch

arXiv · 2405.08602

Optimizing Deep Reinforcement Learning for American Put Option Hedging

Abstract

This paper contributes to the existing literature on hedging American options with Deep Reinforcement Learning (DRL). The study first investigates hyperparameter impact on hedging performance, considering learning rates, training episodes, neural network architectures, training steps, and transaction cost penalty functions. Results highlight the importance of avoiding certain combinations, such as high learning rates with a high number of training episodes or low learning rates with few training episodes and emphasize the significance of utilizing moderate values for optimal outcomes. Additionally, the paper warns against excessive training steps to prevent instability and demonstrates the superiority of a quadratic transaction cost penalty function over a linear version. This study then expands upon the work of Pickard et al. (2024), who utilize a Chebyshev interpolation option pricing method to train DRL agents with market calibrated stochastic volatility models. While the results of Pickard et al. (2024) showed that these DRL agents achieve satisfactory performance on empirical asset paths, this study introduces a novel approach where new agents at weekly intervals to newly calibrated stochastic volatility models. Results show DRL agents re-trained using weekly market data surpass the performance of those trained solely on the sale date. Furthermore, the paper demonstrates that both single-train and weekly-train DRL agents outperform the Black-Scholes Delta method at transaction costs of 1% and 3%. This practical relevance suggests that practitioners can leverage readily available market data to train DRL agents for effective hedging of options in their portfolios.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Reilly Pickard, F. Wredenhagen, Y. Lawryshyn. 2024-05-14. Optimizing Deep Reinforcement Learning for American Put Option Hedging. https://arxiv.org/abs/2405.08602

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Risk-minimizing reinsurance with adverse selection

This paper provides a comprehensive characterization of optimal reinsurance mechanisms under adverse selection when a monopolistic reinsurer faces a continuum of insurers with multi-dimensional private information on loss exposures and Value-at-Risk (VaR) preferences. Moving beyond standard single-dimensional or discrete-type models, we resolve an infinite-dimensional, non-convex screening problem by introducing a novel pre-reinsurance VaR reparameterization and an indirect utility representation. This framework converts global incentive constraints into tractable, typewise conditions. Using a Breeden-Litzenberger representation, we show that optimal indirect utility consistently adopts a hinge form across stop-loss, quota-share, and change-loss contract classes, yielding an endogenous market-exclusion threshold. Among participating types, stop-loss contracts screen via type-dependent deductibles, whereas quota-share contracts induce complete pooling at full coverage. Moreover, under mild conditions, expanding the contract space to change-loss policies yields no extra profit beyond optimal stop-loss designs. Numerical examples further illustrate the fundamental economic trade-off between extracting information rents and excluding low-exposure types. Our approach provides deeper insight into the mechanisms of discrimination and exclusion in reinsurance markets characterized by asymmetric information, enriching theoretical and practical perspectives on contract design in the presence of adverse selection.

q-fin.RM

Quantifying the 2027 Solvency II Risk Margin Reform

The 2027 Solvency II reform recalibrates the Risk Margin by reducing the prescribed cost-of-capital rate from 6% to 4.75% and introducing a time-dependent attenuation of future Solvency Capital Requirements. This paper develops an analytical and numerical framework for characterizing the effect of the final regulatory calibration. By normalizing discounted projected capital requirements into a probability distribution over run-off time, we obtain an exact representation of the ratio between the revised and previous Risk Margins. The framework yields sharp universal and horizon-specific bounds, characterizes the effect of later capital timing through stochastic dominance, and shows that mean run-off time alone does not determine the reform effect when temporal dispersion varies. Additional bounds are derived conditional on mean run-off time and horizon. Conditional on a given projected SCR path, the mechanical reduction lies between 20.83% and 60.42%. For proportional Best Estimate projections, an exact covariance decomposition identifies how departures from proportionality affect the relative reform ratio. Holding the projected SCR path fixed, the revised formula has a lower direct interest-rate semi-elasticity than the previous calibration. A reduced-form stochastic extension further quantifies convexity effects arising from uncertainty in capital persistence. Numerical applications reconstruct Risk Margin calculations from published actuarial run-off profiles and complement them with controlled long-horizon, interest-rate, and persistence experiments. The results show that the reform effect is governed by the temporal structure of future capital and provide tractable tools for assessing projected capital profiles, Risk Margin simplifications, and direct discount-rate sensitivity.

q-fin.RM

Towards foundation models for insurance risk modelling

Claim narratives, images and sensor data contain information about insured risks that is difficult to use through existing actuarial models. Foundation models learn patterns from large datasets before being adapted to particular tasks. By turning these high-dimensional sources into variables or numerical representations, they could help insurers use more of the information they already collect, potentially reducing the experience needed to develop each application. For example, a language model could identify a worsening injury in a new claim note, allowing a reserving model to recognise the change in expected cost before the payments reveal the deterioration. In this paper, we review language, vision, geospatial, time series, tabular and scientific models, explaining existing insurance applications and potential future uses. Scientific models extend this approach to future weather and climate conditions: their simulations can inform loss estimates once local hazards are linked to asset damage, repair costs and insurance coverage. We propose a process to connect these model outputs to actuarial calculations and to assess their predictive contribution, stability and compliance with rules on information use. Evaluating these applications is difficult when final claim costs become known only after long delays, large losses are rare or patterns learned elsewhere fail to transfer to the target portfolio. Richer data can reveal private information and support finer risk classification, which can change access to insurance. Reusing the same models across insurers also creates dependence on shared predictions and providers.

q-fin.RM