Search arXivSearch

arXiv · 2609.08581

AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery

Abstract

Formulaic alpha discovery is a pool-dependent symbolic search problem in which informative feedback is observed primarily when a complete expression is evaluated. This delayed feedback creates two coupled difficulties: the retained alpha pool does not preserve the full history of realized evaluation feedback, and the value of an intermediate construction action is uncertain because its consequence depends on the formula eventually completed. We introduce AlphaRJM, which addresses these difficulties through Reward-Jump Memory, an event-driven latent state that remains fixed during token construction and updates only at terminal evaluation events using the realized pool reward and evaluation outcome, and an action-conditioned SDE return critic that represents future discounted discovery returns with stochastic particles. The particles guide action selection through their mean and uncertainty and are learned using a distributional Bellman objective combining energy-distance matching, mean calibration, and jump regularization. Empirically, AlphaRJM delivers strong and stable gains across multiple equity universes, forecasting horizons, and random seeds, while ablations confirm the complementary roles of persistent evaluation history, stochastic return modeling, and distributional supervision.

Explore related subjects

Keep this discovery

BibTeXRIS

Sayan Dhan, Selvaraju Natarajan. 2026-09-08. AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery. https://arxiv.org/abs/2609.08581

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Asymptotically-informed neural networks for Black-Scholes implied volatility computation

The computation of Black-Scholes implied volatility is a fundamental task in quantitative finance, underpinning option valuation, model calibration and risk management. Although implied volatility is routinely used in practice, the inversion of the Black-Scholes pricing formula remains a challenging numerical problem, particularly in asymptotic regimes corresponding to extreme option prices, strikes or maturities, where the inverse map becomes highly sensitive to perturbations of the price. In this paper, we introduce a new family of asymptotically-informed neural-network architectures for implied-volatility computation. Exploiting the distinct behaviours of the Black-Scholes pricing function in different volatility regimes, we propose a family of architectures that learn a trainable partition of the price-log-moneyness domain through a system of gating functions and combines specialised local approximations of the implied-volatility function within each region. Extensive numerical experiments demonstrate that the proposed models consistently outperform standard feed-forward neural networks across a wide range of parameter domains, often by several orders of magnitude in relative accuracy while maintaining excellent generalisation properties. Furthermore, the neural-network outputs provide highly accurate initial guesses for a third-order Householder scheme, allowing near machine-precision implied-volatility computations after only two refinement iterations.

q-fin.CP

Machine Learning Classification and Portfolio Construction: Does the Loss Function Matter?

Classification outperforms regression across matched machine learning models in portfolio construction. A stacking ensemble of gradient boosted tree, random forest, and neural network yields a value-weighted annualized Sharpe ratio of 2.08 for classification and 1.39 for regression. This outperformance strengthens with class granularity and persists across subsamples and after transaction costs. Spanning tests show that classification retains economically large alphas after we control for regression, whereas regression alphas shrink substantially once we control for classification. These results indicate that classification extracts more return information than matched regression. Our diagnostics trace classification's advantage to more precise separation of return deciles.

q-fin.GN

Realised Volatility Forecasting: Machine Learning via Financial Word Embedding

We examine whether financial news can improve realised volatility forecasting using a parsimonious NLP-based framework that incorporates specialised financial word embeddings alongside general-purpose alternatives. News-only forecasts contain useful predictive information but generally do not outperform strong volatility-history benchmarks. Crucially, combining stock-related news forecasts with a strong volatility-history benchmark lowers forecast losses for several specifications and increases realised utility, providing evidence consistent with forecast complementarity. Performance varies across news types, embedding representations, and volatility regimes. SHAP attributions associate forecast variation with economically interpretable firm-specific and macroeconomic news themes.

q-fin.CP