Search arXivSearch

arXiv · 2308.06935

Insurance pricing on price comparison websites via reinforcement learning

Abstract

The emergence of price comparison websites (PCWs) has presented insurers with unique challenges in formulating effective pricing strategies. Operating on PCWs requires insurers to strike a delicate balance between competitive premiums and profitability, amidst obstacles such as low historical conversion rates, limited visibility of competitors' actions, and a dynamic market environment. In addition to this, the capital intensive nature of the business means pricing below the risk levels of customers can result in solvency issues for the insurer. To address these challenges, this paper introduces reinforcement learning (RL) framework that learns the optimal pricing policy by integrating model-based and model-free methods. The model-based component is used to train agents in an offline setting, avoiding cold-start issues, while model-free algorithms are then employed in a contextual bandit (CB) manner to dynamically update the pricing policy to maximise the expected revenue. This facilitates quick adaptation to evolving market dynamics and enhances algorithm efficiency and decision interpretability. The paper also highlights the importance of evaluating pricing policies using an offline dataset in a consistent fashion and demonstrates the superiority of the proposed methodology over existing off-the-shelf RL/CB approaches. We validate our methodology using synthetic data, generated to reflect private commercially available data within real-world insurers, and compare against 6 other benchmark approaches. Our hybrid agent outperforms these benchmarks in terms of sample efficiency and cumulative reward with the exception of an agent that has access to perfect market information which would not be available in a real-world set-up.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tanut Treetanthiploet, Yufei Zhang, Lukasz Szpruch, Isaac Bowers-Barnard, Henrietta Ridley, James Hickey, Chris Pearce. 2023-08-14. Insurance pricing on price comparison websites via reinforcement learning. https://arxiv.org/abs/2308.06935

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Design and pricing of a transparent parametric-modeled loss CAT bond: application to German windstorm

Catastrophe (cat) bonds overcome some lack of reinsurance by sourcing capacity from the wider capital markets. We present a new type of cat bond addressing the known trade-off between moral hazard and basis risk. As our main contributions we propose a trigger mechanism which is entirely transparent and simpler to evaluate compared to indemnity modeling techniques, as well as a methodology to price this cat bond. This is relevant for insurers and public authorities in a world where natural disasters are occurring with increasing frequency and severity due to climate change, but also for players willing to enter the cat bond market for whom the lack of transparency of this asset class has been a significant obstacle. Our trigger is derived from a cost random field which separates the physical hazard, a vulnerability function and the exposure. This allows the trigger to take a flexible form between parametric and modeled loss, in case exposure is taken into account. We present a case study based on historical windstorm events impacting Germany. Using wind speed data from historical storms, we fit a max-stable random field on a resolution which is standard in the reinsurance industry. The availability of industry loss and exposure data allows us to calibrate the vulnerability component to historical observations. Besides measuring the basis risk associated with our trigger, we perform a full model assessment and discuss numerical results.

q-fin.PR

Fundamentals of Perpetual Futures

Perpetual futures are the most popular cryptocurrency derivatives. Perpetuals offer leveraged exposure to their underlying without rollover or direct ownership. Unlike fixed-maturity futures, perpetuals are not guaranteed to converge to the spot price. To minimize the gap between perpetual and spot prices, long investors periodically pay shorts a funding rate proportional to this difference. We derive no-arbitrage prices for perpetual futures in frictionless markets and bounds in markets with trading costs. Empirically, deviations from these prices in crypto are larger than in traditional currency markets, comove across currencies, and diminish over time. An implied arbitrage strategy yields high Sharpe ratios.

q-fin.PR

A deep learning approach for pricing convertible bonds with path-dependent reset and call provisions

This paper develops a deep learning framework for pricing convertible bonds with path-dependent downward reset and issuer call provisions governed by rolling-window triggers. We formulate the valuation problem as a path-dependent partial differential equation (PPDE) that captures both the historical stock-price path and the evolution of the conversion price. Model-specific PPDEs are derived under GBM, CEV, and Heston dynamics. Under suitable conditions, we establish the existence and uniqueness of a piecewise viscosity solution linked by contractual transmission conditions at monitoring dates. For computation, we construct a fixed-grid backward dynamic programming scheme and approximate its conditional expectations using neural networks, with $L^2$ convergence to the exact fixed-grid recursion as the approximation errors vanish. An application to the China CITIC Bank Convertible Bond produces stable prices across the three models and close agreement with the LSMC benchmark, but outperforms in dimensional scaling. The results show that contractual provisions have a greater valuation effect than the choice of underlying dynamics. The call provision reduces the bond value by truncating upside gains, whereas the downward reset provision increases it under the benchmark specification because improved conversion terms dominate the effect of earlier redemption. Delta and Gamma obtained by automatic differentiation of smooth local network approximations closely agree with central finite-difference estimates. The framework provides a flexible approach to pricing and sensitivity analysis for convertible bonds with complex path-dependent provisions.

q-fin.PR