Search arXivSearch

arXiv · 2604.02035

Reinforcement Learning for Speculative Trading under Exploratory Framework

Abstract

We study a speculative trading problem within the exploratory reinforcement learning (RL) framework of Wang et al. [2020]. The problem is formulated as a sequential optimal stopping problem over entry and exit times under general utility function and price process. We first consider a relaxed version of the problem in which the stopping times are modeled by the jump times of Cox processes driven by bounded, non-randomized intensity controls. Under the exploratory formulation, the agent's randomized control is characterized via the probability measure over the jump intensities, and their objective function is regularized by Shannon's differential entropy. This yields a system of the exploratory HJB equations and Gibbs distributions in closed-form as the optimal policy. Error estimates and convergence of the RL objective to the value function of the original problem are established. Finally, an RL algorithm is designed, and its implementation is showcased in a pairs-trading application.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yun Zhao, Alex S. L. Tse, Harry Zheng. 2026-04-02. Reinforcement Learning for Speculative Trading under Exploratory Framework. https://arxiv.org/abs/2604.02035

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Optimal credit portfolio and consumption with regime switching and default contagion

We study an optimal portfolio and consumption problem in a regime-switching multi-name credit market with default contagion. Default events not only generate direct portfolio losses but also alter the default intensities of surviving securities. Under a Cobb-Douglas utility, the homogeneity property reduces the associated Hamilton-Jacobi-Bellman (HJB) equation to a recursive system of ordinary differential equations indexed by the default states. Solving this system backward from the all-default state, we establish existence and uniqueness of a positive classical solution, characterize the optimal feedback controls, and provide a verification theorem. Finally, numerical experiments present sensitivity analyses and comparisons across contagion settings, regimes, utilities, as well as before and after default.

q-fin.MF

Exact calibration of structural models via time-change

In this note, we propose a general structural approach to model a default time $τ$ as the first-passage time (FPT) of a (``firm-value'') process $S$ below a (``debt'') barrier $K$ that comply with a pre-specified survival probability curve $G(t)=\Pr(τ>t)$. Following an idea of Mbaye and Vrins (Mathematical Finance, 2022) applied to reduced-form models, our approach consists in two steps: choose a latent FPT model driven by a barrier $\tilde{K}$ and process $\tilde{S}$, and time-change those using a deterministic clock $Θ$ to get $K_t=\tilde{K}_{Θ(t)}$ and $S_t:=\tilde{S}_{Θ(t)}$, leading to the final FTP model $(K,S,Θ)$. As the market curve $G$ and the latent model $(\tilde{K},\tilde{S})$ are assumed to be given, the calibration step simply consists in finding the clock $Θ$ such that the distribution of the FPT of $S$ below $K$ coincides with the survival curve $G$. We show that this is achievable for a broad class of specified curves $G$ and latent FTP models. The calibration amounts to a simple inversion of a function, which is almost immediate provided that the latent model is tractable enough. In particular, we show that the AT1P model of Brigo, Morini and Tarenghi - which is able to reproduce a broad range of CDS term-structures - can be regarded as the FPT of a time-changed drifted Brownian motion to a constant barrier: $\tilde{V}_t=μt+W_t$ and $\tilde{K}_t=k<0$. This connection offers an elegant interpretation for the instantaneous volatility function featured in AT1P and yields an immediate calibration of the latter to perfectly match a target survival curve.

q-fin.MF

Stochastic Mortality Model with Fractional Lévy Dynamics

A substantial body of empirical evidence suggests that stochastic mortality models ignoring long range dependence tend to underestimate life expectancy, which may lead to profound implications for pension schemes and funding arrangements. This paper addresses the modelling of stochastic mortality via a mixture of a fractional Lévy process and standard Brownian motion. Our stochastic mortality model exhibits nice analytical tractability in actuarial valuations and flexibility in the choice of underlying Lévy specifications. The long range dependence feature embedded in our stochastic mortality model is well reflected in our empirical studies on the mortality shocks during World War II and COVID-19. For efficient numerical pricing of longevity derivatives, we construct an effective singular value decomposition approximation scheme to overcome the computational challenges arising from the non-Markovian nature of the fractional Lévy process. Truncation errors in singular value decomposition approximation can be reduced by an effective residual correction scheme.

q-fin.MF