Search arXivSearch

arXiv · 1904.11392

Continuous-Time Mean-Variance Portfolio Selection: A Reinforcement Learning Framework

Abstract

We approach the continuous-time mean-variance (MV) portfolio selection with reinforcement learning (RL). The problem is to achieve the best tradeoff between exploration and exploitation, and is formulated as an entropy-regularized, relaxed stochastic control problem. We prove that the optimal feedback policy for this problem must be Gaussian, with time-decaying variance. We then establish connections between the entropy-regularized MV and the classical MV, including the solvability equivalence and the convergence as exploration weighting parameter decays to zero. Finally, we prove a policy improvement theorem, based on which we devise an implementable RL algorithm. We find that our algorithm outperforms both an adaptive control based method and a deep neural networks based algorithm by a large margin in our simulations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Haoran Wang, Xun Yu Zhou. 2019-05-05. Continuous-Time Mean-Variance Portfolio Selection: A Reinforcement Learning Framework. https://arxiv.org/abs/1904.11392

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Money-Back Tontines for Retirement Decumulation: Neural-Network Optimization under Systematic Longevity Risk

Money-back guarantees (MBGs) address bequest concerns in pooled retirement income products by returning the initial purchase price through withdrawals or, after early death, through a benefit to the member's beneficiaries or estate. We study the distinct actuarial problem created by adding an MBG to an individual-account tontine with dynamic withdrawals, investment in domestic and foreign assets, and systematic longevity risk. The retiree trades expected withdrawals (EW) against the lower-tail Conditional Value-at-Risk (CVaR) of terminal wealth under a fixed-horizon plan-to-live convention. Neural networks parameterize admissible withdrawal and rebalancing controls; the MBG is then valued ex post under the learned policy through an equivalent up-front load combining the expected payout with an upper-tail CVaR prudential buffer. We also approximate the effect of contract pooling on per-contract tail risk and pricing. Using long-horizon market and mortality data calibrated for an Australian retiree, we find that expected MBG payouts are modest, while the representative-contract payout has a severe upper tail. Contract pooling substantially reduces the upper-tail exposure on a per-contract basis and lowers the approximate pooled-contract load. International diversification materially improves the EW--CVaR retirement-income trade-off, while affecting the MBG payout distribution and equivalent load only modestly. Stochastic mortality likewise has a modest effect on the efficient frontier and MBG pricing.

q-fin.PM

Large Signal Libraries: Equal-Weight Limits and the Divergent Spectra of Signals and PnL

An ensemble of roughly 3,000 signals over 20 assets was reported to have approximately 90% correlation with the leading component of the asset-space return structure. Does having about 158 signals per available linear dimension explain that alignment? The population answer depends on the research process's design distribution and its relation to returns: crowding alone imposes neither a nonzero mean nor agreement with a principal component. The population theory developed here distinguishes four objects: the equal-weight signal, signal principal components, equal-weight profit and loss (PnL), and PnL principal components. The return operator in the motivating observation is a separate object. Independent libraries converge to their design mean; exchangeable libraries can retain a random conditional mean. Cross-sectional signals over $d$ assets, once demeaned and normalized, lie on the unit sphere $S^{q-1}$ of a $q$-dimensional space, $q=d-1$. Under axial symmetry, their nonzero mean is signal-cloud PC1 exactly when its longitudinal second moment exceeds $1/q$, the isotropic energy share. Residual-and-gap bounds quantify approximate alignment. Combining design weights into signals contracts the tangent of their angle to PC1 to at most $\sqrt{λ_2/λ_1}$ times its value, where $λ_1>λ_2$ are the leading signal-kernel eigenvalues; the factor is sharp. A target-aligned frame separates transverse signal geometry, which PnL discards, from dispersion weighting and temporal centering, which also change the spectrum. A reproducible synthetic example illustrates the geometric threshold, and a proposed empirical program separates library growth from limited-history estimation. No market-data empirical results are presented.

q-fin.PM

Entropic Value-at-Risk parity for tempered stable returns

We develop Entropic Value-at-Risk (EVaR) parity for tempered stable returns. EVaR-based inverse risk parity (IRP) and equal risk contribution (ERC) portfolios are constructed using multivariate normal tempered stable models and independent component analysis with tempered stable components. We derive the corresponding asset-level EVaR and EVaR-deviation contributions and use the latter to separate the fitted location term from EVaR risk contributions. Under Gaussian returns, EVaR-deviation IRP and ERC recover conventional volatility IRP and ERC weights. We evaluate the resulting portfolios in three investment universes. Empirically, EVaR-based ERC portfolios achieve positive Sharpe differences relative to equal weight across the universes.

q-fin.PM