Searcharxiv⌕ Search

arXiv · 2610.03406

PreFER: Interactive Robo-Advisor with Scoring Mechanism

Abstract

We propose an interactive robo-advising framework that learns personalized risk preferences from scores provided by clients. The resulting preference-learning problem is closely related to inverse reinforcement learning (IRL), as the robo-advisor infers the client's latent reward specification from feedback. The robo-advisor interacts with clients iteratively as follows. At each interaction time, the advisor generates investment advice based on the optimal policy distribution derived from an inferred personalized risk preference. The client scores the advice. The advisor updates its assessment of the client's risk preference based on the feedback. This learning procedure motivates us to investigate discrete-time Predictable Forward Exploratory Reward (PreFER) processes and derive an exploratory investment strategy. By interpreting the score as the acceptance probability of a piece of advice, our inverse learning procedure learns the client's exploratory investment distribution using the acceptance-rejection method pioneered by von Neumann. Under CARA preferences, we show that, even though the scores contain noise, the robo-advisor can identify the client's current risk aversion after a sufficiently large number of interactions. The PreFER process then carries the learned preference forward and generates future recommendations under updated market conditions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yuwei Wang, Hoi Ying Wong. 2026-10-02. PreFER: Interactive Robo-Advisor with Scoring Mechanism. https://arxiv.org/abs/2610.03406

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

When defaults cannot be hedged: xVA calculations via local risk-minimization

We consider the pricing and hedging of counterparty credit risk and funding when there is no possibility to hedge the jump to default of either the bank or the counterparty. This represents the situation which is most often encountered in practice, due to the absence of quoted corporate bonds or CDS contracts written on the counterparty and the difficulty for the bank to buy/sell protection on her own default. We apply local risk-minimization to find the optimal strategy and compute it via a BSDE.

q-fin.MF↗

Optimal Investment and Consumption in a Stochastic Factor Model

In this article, we study optimal investment and consumption in an incomplete stochastic factor model for a power utility investor on the infinite horizon. When the state space of the stochastic factor is finite, we give a complete characterisation of the well-posedness of the problem, and provide an efficient numerical algorithm for computing the value function. When the state space is a (possibly infinite) open interval and the stochastic factor is represented by an Itô diffusion, we develop a general theory of sub- and supersolutions for second-order ordinary differential equations on open domains without boundary values to prove existence of the solution to the Hamilton-Jacobi-Bellman (HJB) equation along with explicit bounds for the solution. By characterising the asymptotic behaviour of the solution, we are also able to provide rigorous verification arguments for various models, including -- for the first time -- the Heston model. Finally, we link the discrete and continuous setting and show that that the value function in the diffusion setting can be approximated very efficiently through a fast discretisation scheme.

q-fin.MF↗

Negative Oil & Nickel Squeeze: A Feedback Model for Extreme Commodity Futures Prices

On April 20, 2020, the May front-month WTI oil futures contract, one day before its expiration date, opened near $\$17/$barrel and dropped far below zero in a single trading day, reaching an intraday low of $-\$40.32$ and settling at $-\$37.63$. Such market behavior was unforeseen at the time. This event, and the March 2022 nickel squeeze, illustrate how constraints on physical delivery create pressure to close futures positions and distort futures prices. We develop a feedback model that connects the resulting price distortion to a roll option, and to the imbalance between delivery-constrained long and short positions. Under lognormal benchmark dynamics, the roll-option value is shown to satisfy a pricing PDE which is nonlinear because its payoff depends on the observed futures price, which itself includes the feedback correction. Remarkably, in both cases, there is an explicit solution in terms of classical Black-Scholes-Margrabe exchange option formulas, up to solving a scalar equation. We show this allows that prices may go negative under distortion of a positive-price lognormal model. An analogous construction applies when the lognormal base is switched to Bachelier normal dynamics. Illustrations based on WTI and nickel show the imbalance required to reproduce the extreme event prices under the benchmark assumptions.

q-fin.MF↗