Search arXiv⌕ Search

arXiv · 2610.03406

PreFER: Interactive Robo-Advisor with Scoring Mechanism

Abstract

We propose an interactive robo-advising framework that learns personalized risk preferences from scores provided by clients. The resulting preference-learning problem is closely related to inverse reinforcement learning (IRL), as the robo-advisor infers the client's latent reward specification from feedback. The robo-advisor interacts with clients iteratively as follows. At each interaction time, the advisor generates investment advice based on the optimal policy distribution derived from an inferred personalized risk preference. The client scores the advice. The advisor updates its assessment of the client's risk preference based on the feedback. This learning procedure motivates us to investigate discrete-time Predictable Forward Exploratory Reward (PreFER) processes and derive an exploratory investment strategy. By interpreting the score as the acceptance probability of a piece of advice, our inverse learning procedure learns the client's exploratory investment distribution using the acceptance-rejection method pioneered by von Neumann. Under CARA preferences, we show that, even though the scores contain noise, the robo-advisor can identify the client's current risk aversion after a sufficiently large number of interactions. The PreFER process then carries the learned preference forward and generates future recommendations under updated market conditions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yuwei Wang, Hoi Ying Wong. 2026-10-02. PreFER: Interactive Robo-Advisor with Scoring Mechanism. https://arxiv.org/abs/2610.03406

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Price-Based Framework for Stochastic Portfolio Theory

We develop a price-based framework for stochastic portfolio theory in which trading strategies are generated from nominal price weights and evaluated relative to a price-weighted benchmark. Stock splits and reverse splits induce jumps in the generating weights without changing the value of existing investments. In a semimartingale market with predictable adjustment events, we incorporate the corresponding share adjustments into the self-financing condition and construct additively and multiplicatively generated strategies with explicit jump-corrected wealth decompositions. For additive generation, an entropy-based example exhibits relative wealth tending to $-\infty$ almost surely under repeated splits, demonstrating why the usual argument for long-horizon relative arbitrage does not extend directly. We also show that advance adjustment information alone provides no model-free guarantee of improved performance for functionally generated portfolios. For multiplicative generation, wealth-preserving restarts retain the classical portfolio allocation evaluated at the current price weights. We also establish a correspondence with the capitalization-based framework that preserves absolute wealth and yields a strategy-independent benchmark conversion. We illustrate the framework using daily NYSE data, comparing price- and capitalization-weighted benchmarks and the corresponding diversity-weighted portfolios across price- and capitalization-selected universes.

q-fin.MF↗

When defaults cannot be hedged: xVA calculations via local risk-minimization

We consider the pricing and hedging of counterparty credit risk and funding when there is no possibility to hedge the jump to default of either the bank or the counterparty. This represents the situation which is most often encountered in practice, due to the absence of quoted corporate bonds or CDS contracts written on the counterparty and the difficulty for the bank to buy/sell protection on her own default. We apply local risk-minimization to find the optimal strategy and compute it via a BSDE.

q-fin.MF↗

Optimal Investment and Consumption in a Stochastic Factor Model

In this article, we study optimal investment and consumption in an incomplete stochastic factor model for a power utility investor on the infinite horizon. When the state space of the stochastic factor is finite, we give a complete characterisation of the well-posedness of the problem, and provide an efficient numerical algorithm for computing the value function. When the state space is a (possibly infinite) open interval and the stochastic factor is represented by an Itô diffusion, we develop a general theory of sub- and supersolutions for second-order ordinary differential equations on open domains without boundary values to prove existence of the solution to the Hamilton-Jacobi-Bellman (HJB) equation along with explicit bounds for the solution. By characterising the asymptotic behaviour of the solution, we are also able to provide rigorous verification arguments for various models, including -- for the first time -- the Heston model. Finally, we link the discrete and continuous setting and show that that the value function in the diffusion setting can be approximated very efficiently through a fast discretisation scheme.

q-fin.MF↗