Search arXivSearch

arXiv · 2601.22119

Alpha Discovery via Grammar-Guided Learning and Search

Abstract

Automatically discovering formulaic alpha factors is a central problem in quantitative finance. Existing methods often ignore syntactic and semantic constraints, relying on exhaustive search over unstructured and unbounded spaces. We present AlphaCFG, a grammar-based framework for defining and discovering alpha factors that are syntactically valid, financially interpretable, and computationally efficient. AlphaCFG uses an alpha-oriented context-free grammar to define a tree-structured, size-controlled search space, and formulates alpha discovery as a tree-structured linguistic Markov decision process, which is then solved using a grammar-aware Monte Carlo Tree Search guided by syntax-sensitive value and policy networks. Experiments on Chinese and U.S. stock market datasets show that AlphaCFG outperforms state-of-the-art baselines in both search efficiency and trading profitability. Beyond trading strategies, AlphaCFG serves as a general framework for symbolic factor discovery and refinement across quantitative finance, including asset pricing and portfolio construction.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Han Yang, Dong Hao, Zhuohan Wang, Qi Shi, Xingtong Li. 2026-01-29. Alpha Discovery via Grammar-Guided Learning and Search. https://arxiv.org/abs/2601.22119

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

CFOs Meet LLMs

Business sentiment is a closely watched economic signal, but measuring it is slow and costly: surveys typically reach only a few hundred firms, arrive periodically, and take time to compile. We show that large language models hold the potential to address these shortcomings. We prompt an LLM to role-play as the CFO of a specific company on a specific date, for every public-company CFO who responded to the Duke--Federal Reserve CFO Survey between 2002 and 2025, and answer a question about economy-wide optimism. The LLM-generated optimism score predicts the individual CFO's actual answer, even in specifications that include firm and year-quarter fixed effects as well as a control variable measuring the human CFO's lagged response. Accuracy increases with the information provided to the LLM, and the relation persists under quarterly aggregation. We find the same patterns hold for two other questions measuring CFO expectations: the respondent's optimism about their own firm and their expectation of own-firm revenues. With appropriate conditioning, LLMs may in the future be able to serve as digital twins of executives, offering scalable, high-frequency expectations data for financial research and policy.

q-fin.CP

The Physical Crash Frontier: What Finite Option Quotes Can and Cannot Reveal

Physical crash probabilities recovered from option prices depend on a pricing kernel and on a risk-neutral distribution that finitely many bid and ask quotes do not identify. For a power utility investor, we characterize the pairs of physical crash probability and expected loss below the crash threshold that the quotes admit; the boundary of this set is the physical crash frontier. Both coordinates are ratios of moments, yet when the index is bounded above the set is convex, and second-order cone programs compute it exactly at the calibrated risk aversion of two. In a decade of weekly S&P 500 cross sections, the quotes beyond the two puts nearest a 10 percent decline shrink the range of admissible crash probabilities by about 80 percent, yet its upper end remains two to three times its lower end. That lower end exists only because the index is bounded. Otherwise, for any investor more risk averse than the log investor, a vanishing probability far in the right tail inflates the denominator and drives the crash probability to zero while every quote stays inside its spread. A positive floor is therefore a joint statement about prices and a tail restriction; anything tighter than the frontier is an assumption.

q-fin.CP

OrderFusion+: Probabilistic Buy--Sell Price Trajectory Forecasting in Intraday Electricity Markets

Intraday electricity markets enable participants to adjust energy positions close to delivery, with price forecasts necessary to support trading and the scheduling of flexible electricity resources as well as increasingly responsive consumers. Forecasting model specifications have progressed from using macro-features, such as renewable generation and load, to the micro-features of continuous orderbooks. A recent advanced deep learning model, OrderFusion, explicitly models micro-level buy-sell orderbook interactions. However, despite its superior comparative forecasting performance, it considers only one delivery product at a time, ignoring neighboring-product information. Moreover, when forecasting aggregated price indices such as the liquid German ID3, ID2, and ID1 products, it omits the information contained in price trajectories. In contrast, pretrained time-series foundation models have shown success in financial, renewable-energy, and day-ahead electricity price forecasting. However, their performance on intraday orderbook data remains an open research question of considerable practical importance. In this paper, we propose OrderFusion+, an open-source deep learning model that combines historical orders from the target and neighboring delivery products to forecast probabilistic buy-sell price trajectories. We benchmark OrderFusion+ against forecasting baselines and pretrained foundation models, and investigate dynamic market conditions through the designed dynamic masking mechanism, revealing insights into market efficiency. The implementation and forecasts can be found at: https://runyao-yu.com/OrderFusion/

q-fin.CP