Search arXivSearch

arXiv · 2604.17166

The Virtue of Sparsity in Complexity

Abstract

Sparsity or complexity? In modern high-dimensional asset pricing, these are often viewed as competing principles: richer feature spaces appear to favor complexity, while economic intuition has long favored parsimony. We show that this tension is misplaced. We distinguish capacity sparsity-the dimensionality of the candidate feature space-from factor sparsity-the parsimonious structure of priced risks-and argue that the two are complements: expanding capacity enables the discovery of factor sparsity. Revisiting the benchmark empirical design of Didisheim et al. (2025) and pushing it to higher complexity regimes, we show that nonlinear feature expansions combined with basis pursuit yield portfolios whose out-of-sample performance dominates ridgeless benchmarks beyond a critical complexity threshold. The evidence shows that the gains from complexity arise not from retaining more factors, but from enlarging the space from which a sparse structure of priced risks can be identified. The virtue of complexity in asset pricing operates through factor sparsity.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nima Afsharhajari, Jonathan Yu-Meng Li. 2026-04-18. The Virtue of Sparsity in Complexity. https://arxiv.org/abs/2604.17166

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Machine Learning Classification and Portfolio Construction: Does the Loss Function Matter?

Classification outperforms regression across matched machine learning models in portfolio construction. A stacking ensemble of gradient boosted trees, random forest, and neural network yields a value-weighted annualized Sharpe ratio of 2.08 for classification and 1.39 for regression. This outperformance strengthens with class granularity and persists across subsamples and after transaction costs. Spanning tests show that classification retains economically large alphas after we control for regression, whereas regression alphas shrink substantially once we control for classification. These results indicate that classification extracts more return information than matched regression. Our diagnostics trace classification's advantage to more precise separation of return deciles.

q-fin.GN

AI for AI: Optimizing Additional Infrastructure Build-out to Power Artificial Intelligence Data Centers

The twenty-first century's transformative technology, artificial intelligence, is increasingly constrained by the twentieth century's transformative technology, the electricity grid. Rapid growth in electricity demand from data centers is leading to higher electricity prices, without a compensating supply-side response. We develop a framework linking data-center load growth, available generation capacity, and market-clearing prices to understand this phenomenon. We first analyze a deterministic model to show how differing estimates of demand and supply growth rates affect prices. We then model the expansion of new data centers and their associated electricity demand, together with build-outs of new electricity supply, as stochastic processes,resulting in probabilistic distributions of supply, demand, and prices rather than a single forecast. Finally, we formulate generation expansion as a stochastic control problem in which a revenue-maximizing investor dynamically chooses the intensity of supply-side investments. The analysis highlights a central challenge of the data-center build-out: even when rapid demand growth increases the need for new generation, the uncertainties related to load forecasts, development execution risks, and value cannibalization from overbuilding capacity may weaken incentives to invest at the pace required to keep electricity prices stable.

q-fin.GN

Measuring DeFi Risk

Decentralized finance (DeFi) lending has grown from nonexistent in 2017 to nearly 40 billion US Dollars in deposited funds in May 2022. Using cryptocurrency as collateral, the platforms match speculative margin trading with yield-seeking depositors lending coins pegged to the dollar (stable coins). Depositors receive claims guaranteed by a basket of collateral, akin to new stable coins. We develop a framework requiring only knowledge of aggregate deposits and borrowings to measure overall system risks to lenders and borrowers. Using evidence from major protocols, the measures identify an increase in system fragility beyond prudent levels around mid 2021, with a potential loss of peg for extreme variations in coin prices. Overall, the model offers an easily implementable aggregate risk metric capturing the perspectives of synthetic investors and offers early warning signals as the industry is moving from deposits guaranteed by collateral to fiat money.

q-fin.GN