Searcharxiv⌕ Search

arXiv subjects

Hoi Ying Wong

Publications and source records attributed to Hoi Ying Wong.

18 recordsLinked to original sources

PreFER: Interactive Robo-Advisor with Scoring Mechanism

We propose an interactive robo-advising framework that learns personalized risk preferences from scores provided by clients. The resulting preference-learning problem is closely related to inverse reinforcement learning (IRL), as the robo-advisor infers the client's latent reward specification from feedback. The robo-advisor interacts with clients iteratively as follows. At each interaction time, the advisor generates investment advice based on the optimal policy distribution derived from an inferred personalized risk preference. The client scores the advice. The advisor updates its assessment of the client's risk preference based on the feedback. This learning procedure motivates us to investigate discrete-time Predictable Forward Exploratory Reward (PreFER) processes and derive an exploratory investment strategy. By interpreting the score as the acceptance probability of a piece of advice, our inverse learning procedure learns the client's exploratory investment distribution using the acceptance-rejection method pioneered by von Neumann. Under CARA preferences, we show that, even though the scores contain noise, the robo-advisor can identify the client's current risk aversion after a sufficiently large number of interactions. The PreFER process then carries the learned preference forward and generates future recommendations under updated market conditions.

q-fin.MF↗

Robust dividend policy: Equivalence of Epstein-Zin and Maenhout preferences

In a continuous-time economy, this paper formulates the Epstein-Zin preference for discounted dividends received by an investor as an Epstein-Zin singular control utility. We introduce a backward stochastic differential equation with an aggregator integrated with respect to a singular control, prove its well-posedness, and show that it coincides with the Epstein-Zin singular control utility. We then establish that this formulation is equivalent to a robust dividend policy chosen by the firm's executive under the Maenhout's ambiguity-averse preference. In particular, the robust dividend policy takes the form of a threshold strategy on the firm's surplus process, where the threshold level is characterized as the free boundary of a Hamilton-Jacobi-Bellman variational inequality. Therefore, dividend-caring investors can choose firms that match their preferences by examining stock's dividend policies and financial statements, whereas executives can make use of dividend to signal their confidence, in the form of ambiguity aversion, on realizing the earnings implied by their financial statements.

q-fin.MF↗

Robust Exploratory Stopping under Ambiguity in Reinforcement Learning

We propose and analyze a continuous-time robust reinforcement learning framework for optimal stopping under ambiguity. In this framework, an agent chooses a robust exploratory stopping time motivated by two objectives: robust decision-making under ambiguity and learning about the unknown environment. Here, ambiguity refers to considering multiple probability measures dominated by a reference measure, reflecting the agent's awareness that the reference measure representing her learned belief about the environment would be erroneous. Using the $g$-expectation framework, we reformulate the optimal stopping problem under ambiguity as a robust exploratory control problem with Bernoulli distributed controls. We then characterize the optimal Bernoulli distributed control via backward stochastic differential equations and, based on this, construct the robust exploratory stopping time that approximates the optimal stopping time under ambiguity. Last, we establish a policy iteration theorem and implement it as a reinforcement learning algorithm. Numerical experiments demonstrate the convergence, robustness, and scalability of our reinforcement learning algorithm across different levels of ambiguity and exploration.

math.OC↗

Duality and DeepMartingale for High-Dimensional Optimal Switching: Computable Upper Bounds and Approximation-Expressivity Guarantees

We study finite-horizon optimal switching with discrete intervention dates on a general filtration, allowing continuous-time observations between decision dates, and develop a deep-learning-based dual framework with computable upper bounds. We first derive a dual representation for multiple switching by introducing a family of martingale penalties. The minimal penalty is characterized by the Doob martingales of the continuation values, which yields a fully computable upper bound. We then extend DeepMartingale from optimal stopping to optimal switching and establish convergence under both the upper-bound loss and an $L^2$-surrogate loss. We also provide an expressivity analysis: under the stated structural assumptions, for any target accuracy $\varepsilon>0$, there exist neural networks of size at most $c d^{q}\varepsilon^{-r}$ whose induced dual upper bound approximates the true value within $\varepsilon$, where $c$, $q$, and $r$ are independent of $d$ and $\varepsilon$. Hence, the dual solver avoids the curse of dimensionality under the stated structural assumptions. For numerical assessment, we additionally implement a deep policy-based approach to produce feasible lower bounds and empirical upper--lower gaps. Numerical experiments on Brownian and Brownian--Poisson models demonstrate small upper--lower gaps and favorable performance in high dimensions. The learned dual martingale also yields a practical delta-hedging strategy.

math.OC↗

DeepMartingale: Duality of the Optimal Stopping Problem with Expressivity and High-Dimensional Hedging

We propose \textit{DeepMartingale}, a deep-learning framework for the dual formulation of discrete-monitoring optimal stopping problems under continuous-time models. Leveraging a martingale representation, our method implements a \emph{pure-dual} procedure that directly optimizes over a parameterized class of martingales, producing computable and tight \emph{dual upper bounds} for the value function in high-dimensional settings without requiring any primal information or Snell-envelope approximation. We prove convergence of the resulting upper bounds under mild assumptions for both first- and second-moment losses. A key contribution is an expressivity theorem showing that \textit{DeepMartingale} can approximate the true value function to any prescribed accuracy $\varepsilon$ using neural networks of size at most $\tilde{c} d^{\tilde{q}}\varepsilon^{-\tilde{r}}$, with constants independent of the dimension $d$ and accuracy $\varepsilon$, thereby avoiding the curse of dimensionality. Since expressivity in this setting translates into scalability, our theory also motivates estimating the dimension scaling law to guide architecture design and the training setup in deep learning-based numerical computation and the choice of rebalancing frequency for the related hedging strategy. The learned martingale representation further yields a practical and dimension-scalable \emph{deep delta hedging strategy}. Numerical experiments on high-dimensional Bermudan option benchmarks confirm convergence, expressivity, scalable training, and the stability of the resulting upper bounds and hedging performance.

math.OC↗

Contextual Quantile Minimization for Two-Stage Stochastic Programs

Contextual stochastic optimization is an advanced methodology to model uncertainty in the presence of contextual information during decision planning processes. Although classical methodologies focus on minimizing the expectation of a random loss, in many applications, risk-averse decision-makers may be interested in minimizing a specific quantile as a more prudent alternative. In this paper, we propose a new risk-averse contextual stochastic optimization problem with a quantile objective for general two-stage problems. Given historical data on the model's random parameters and contextual information, we model the conditional quantile by replacing the conditional expectation in its variational characterization with a generic estimator. Under two sets of mild regularity conditions, we derive the asymptotic almost-sure convergence and convergence in probability of the optimal solution and the optimal value of the associated optimization problem to their true counterparts. Optimization problems with a quantile objective is usually non-convex, which are generally regarded as challenging to solve. To address the computational difficulties, we propose a new stochastic inexact constraint generation method with convergence guarantee. Finally, through numerical experiments on a single-server appointment scheduling problem, we study the computational performance of our proposed solution method as well as operational performance of our proposed methodology. Our results demonstrate the importance of incorporating useful contextual information and decision-maker's risk attitude into the optimization model.

math.OC↗

Robust Time-inconsistent Linear-Quadratic Stochastic Controls: A Stochastic Differential Game Approach

This paper studies robust time-inconsistent (TIC) linear-quadratic stochastic control problems, formulated by stochastic differential games. By a spike variation approach, we derive sufficient conditions for achieving the Nash equilibrium, which corresponds to a time-consistent (TC) robust policy, under mild technical assumptions. To illustrate our framework, we consider two scenarios of robust mean-variance analysis, namely with state- and control-dependent ambiguity aversion. We find numerically that with time inconsistency haunting the dynamic optimal controls, the ambiguity aversion enhances the effective risk aversion faster than the linear, implying that the ambiguity in the TIC cases is more impactful than that under the TC counterparts, e.g., expected utility maximization problems.

math.OC↗

Long-range dependent mortality modeling with cointegration

Empirical studies with publicly available life tables identify long-range dependence (LRD) in national mortality data. Although the longevity market is supposed to benchmark against the national force of mortality, insurers are more concerned about the forces of mortality associated with their own portfolios than the national ones. Recent advances on mortality modeling make use of fractional Brownian motion (fBm) to capture LRD. A theoretically flexible approach even considers mixed fBm (mfBm). Using Volterra processes, we prove that the direct use of mfBm encounters the identification problem so that insurers hardly detect the LRD effect from their portfolios. Cointegration techniques can effectively bring the LRD information within the national force of mortality to the mortality models for insurers' experienced portfolios. Under the open-loop equilibrium control framework, the explicit and unique equilibrium longevity hedging strategy is derived for cointegrated forces of mortality with LRD. Using the derived hedging strategy, our numerical examples show that the accuracy of estimating cointegration is crucial for hedging against the longevity exposure of insurers with LRD national force of mortality.

q-fin.RM↗

Duality in optimal consumption--investment problems with alternative data

This study investigates an optimal consumption--investment problem in which the unobserved stock trend is modulated by a hidden Markov chain that represents different economic regimes. In the classical approach, the hidden state is estimated from historical asset prices, but recent advancements in technology enable investors to consider alternative data in their decision-making. These include social media commentary, expert opinions, COVID-19 pandemic data, and GPS data, which originate outside of the standard sources of market data but are considered useful for predicting stock trends. We develop a novel duality theory for this problem and consider a jump-diffusion process for the alternative data series. This theory helps investors in identifying ``useful'' alternative data for dynamic decision-making by offering conditions to the filter equation that permit the use of a control approach based on the dynamic programming principle. We demonstrate an application for proving a unique smooth solution for a constant relative risk-averse agent once the distributions of the signals generated from alternative data satisfy a bounded likelihood ratio condition. In doing so, we obtain an explicit consumption--investment strategy that takes advantage of different types of alternative data that have not been addressed in the literature.

q-fin.MF↗

Adaptive Robust Online Portfolio Selection

The online portfolio selection (OLPS) problem differs from classical portfolio model problems, as it involves making sequential investment decisions. Many OLPS strategies described in the literature capture market movement based on various beliefs and are shown to be profitable. In this paper, we propose a robust optimization (RO)-based strategy that takes transaction costs into account. Moreover, unlike existing studies that calibrate model parameters from benchmark data sets, we develop a novel adaptive scheme that decides the parameters sequentially. With a wide range of parameters as input, our scheme captures market uptrend and protects against market downtrend while controlling trading frequency to avoid excessive transaction costs. We numerically demonstrate the advantages of our adaptive scheme against several benchmarks under various settings. Our adaptive scheme may also be useful in general sequential decision-making problems. Finally, we compare the performance of our strategy with that of existing OLPS strategies using both benchmark and newly collected data sets. Our strategy outperforms these existing OLPS strategies in terms of cumulative returns and competitive Sharpe ratios across diversified data sets, demonstrating its adaptability-driven superiority.

q-fin.PM↗

Time-inconsistency with rough volatility

In this paper, we consider equilibrium strategies under Volterra processes and time-inconsistent preferences embracing mean-variance portfolio selection (MVP). Using a functional Itô calculus approach, we overcome the non-Markovian and non-semimartingale difficulty in Volterra processes. The equilibrium strategy is then characterized by an extended path-dependent Hamilton-Jacobi-Bellman equation system under a game-theoretic framework. A verification theorem is provided. We derive explicit solutions to three problems, including MVP with constant risk aversion, MVP for log returns, and a mean-variance objective with a linear controlled Volterra process. We also thoroughly examine the effect of volatility roughness on equilibrium strategies. Numerical experiments demonstrate that trading rules with rough volatility outperform the classic counterparts.

q-fin.MF↗

Time-consistent mean-variance reinsurance-investment problem with long-range dependent mortality rate

This paper investigates the time-consistent mean-variance reinsurance-investment (RI) problem faced by life insurers. Inspired by recent findings that mortality rates exhibit long-range dependence (LRD), we examine the effect of LRD on RI strategies. We adopt the Volterra mortality model proposed in Wang et al.(2021) to incorporate LRD into the mortality rate process and describe insurance claims using a compound Poisson process with the intensity represented by stochastic mortality rate. Under the open-loop equilibrium mean-variance criterion, we derive explicit equilibrium RI controls and study the uniqueness of these controls in cases of constant and state-dependent risk aversion. We simultaneously resolve difficulties arising from unbounded non-Markovian parameters and sudden increases in the insurer's wealth process. We also use a numerical study to reveal the influence of LRD on equilibrium strategies.

q-fin.RM↗

Optimal Expansion of Business Opportunity

Any firm whose business strategy has an exposure constraint that limits its potential gain naturally considers expansion, as this can increase its exposure. We model business expansion as an enlargement of the opportunity set for business policies. However, expansion is irreversible and has an opportunity cost attached. We use the expected optimization of utility to formulate this as a novel stochastic control problem combined with an optimal stopping time, and we derive an explicit solution for exponential utility. We apply the framework to an investment and a reinsurance scenario. In the investment problem, the cost and incentives to increase the trading exposure are analyzed, while the optimal timing for an insurer to launch its reinsurance business is investigated in the reinsurance problem. Our model predicts that the additional income gained through business expansion is the key incentive for a decision to expand. Interestingly, companies may have this incentive but are likely to wait for a period of time before expanding, although situations of zero opportunity cost or specific restrictive conditions on the model parameters are exceptions to waiting. The business policy remains on the boundary of the opportunity set before expansion during the waiting period. The length of the waiting period is related to the opportunity cost, return, and risk of the expanded business.

q-fin.RM↗

Efficient Social Distancing for COVID-19: An Integration of Economic Health and Public Health

Social distancing has been the only effective way to contain the spread of an infectious disease prior to the availability of the pharmaceutical treatment. It can lower the infection rate of the disease at the economic cost. A pandemic crisis like COVID-19, however, has posed a dilemma to the policymakers since a long-term restrictive social distancing or even lockdown will keep economic cost rising. This paper investigates an efficient social distancing policy to manage the integrated risk from economic health and public health issues for COVID-19 using a stochastic epidemic modeling with mobility controls. The social distancing is to restrict the community mobility, which was recently accessible with big data analytics. This paper takes advantage of the community mobility data to model the COVID-19 processes and infer the COVID-19 driven economic values from major market index price, which allow us to formulate the search of the efficient social distancing policy as a stochastic control problem. We propose to solve the problem with a deep-learning approach. By applying our framework to the US data, we empirically examine the efficiency of the US social distancing policy and offer recommendations generated from the algorithm.

stat.AP↗

Volterra mortality model: Actuarial valuation and risk management with long-range dependence

While abundant empirical studies support the long-range dependence (LRD) of mortality rates, the corresponding impact on mortality securities are largely unknown due to the lack of appropriate tractable models for valuation and risk management purposes. We propose a novel class of Volterra mortality models that incorporate LRD into the actuarial valuation, retain tractability, and are consistent with the existing continuous-time affine mortality models. We derive the survival probability in closed-form solution by taking into account of the historical health records. The flexibility and tractability of the models make them useful in valuing mortality-related products such as death benefits, annuities, longevity bonds, and many others, as well as offering optimal mean-variance mortality hedging rules. Numerical studies are conducted to examine the effect of incorporating LRD into mortality rates on various insurance products and hedging efficiency.

q-fin.MF↗

Mean-variance portfolio selection under Volterra Heston model

Motivated by empirical evidence for rough volatility models, this paper investigates continuous-time mean-variance (MV) portfolio selection under the Volterra Heston model. Due to the non-Markovian and non-semimartingale nature of the model, classic stochastic optimal control frameworks are not directly applicable to the associated optimization problem. By constructing an auxiliary stochastic process, we obtain the optimal investment strategy, which depends on the solution to a Riccati-Volterra equation. The MV efficient frontier is shown to maintain a quadratic curve. Numerical studies show that both roughness and volatility of volatility materially affect the optimal strategy.

q-fin.PM↗

Merton's portfolio problem under Volterra Heston model

This paper investigates Merton's portfolio problem in a rough stochastic environment described by Volterra Heston model. The model has a non-Markovian and non-semimartingale structure. By considering an auxiliary random process, we solve the portfolio optimization problem with the martingale optimality principle. Optimal strategies for power and exponential utilities are derived in semi-closed form solutions depending on the respective Riccati-Volterra equations. We numerically examine the relationship between investment demand and volatility roughness.

q-fin.PM↗

Simulation-based Value-at-Risk for Nonlinear Portfolios

Value-at-risk (VaR) has been playing the role of a standard risk measure since its introduction. In practice, the delta-normal approach is usually adopted to approximate the VaR of portfolios with option positions. Its effectiveness, however, substantially diminishes when the portfolios concerned involve a high dimension of derivative positions with nonlinear payoffs; lack of closed form pricing solution for these potentially highly correlated, American-style derivatives further complicates the problem. This paper proposes a generic simulation-based algorithm for VaR estimation that can be easily applied to any existing procedures. Our proposal leverages cross-sectional information and applies variable selection techniques to simplify the existing simulation framework. Asymptotic properties of the new approach demonstrate faster convergence due to the additional model selection component introduced. We have also performed sets of numerical results that verify the effectiveness of our approach in comparison with some existing strategies.

stat.ME↗