SearcharxivSearch

arXiv subjects

Yu-Jui Huang

Publications and source records attributed to Yu-Jui Huang.

At least 19 recordsLinked to original sources

Optimal Investment with Switching Preferences

Major life events can significantly increase individuals' risk aversion over a sustained period of time, as empirical studies reveal. How such an event-triggered shift of risk preferences impacts optimal investment is the focus of this paper. On a finite time horizon where a major life event may occur independently of the financial market, an investor aims to maximize her expected utility from terminal wealth while foreseeing a potential change in her risk aversion. We find that the associated Hamilton--Jacobi--Bellman (HJB) equation involves the post-event optimal value function (under elevated but fixed risk aversion after the event's occurrence), and the Fenchel--Legendre transform fails to linearize this HJB equation: it yields a parabolic equation with a fully nonlinear term, induced precisely by the post-event optimal value function. Through a combination of fixed-point, compactness, and verification arguments, we establish the existence of a positive convex classical solution with suitable growth to the fully-nonlinear parabolic equation. The convex conjugate of this solution is shown to satisfy the HJB equation and coincides with the pre-event optimal value function. The optimal trading strategy is obtained by concatenating the optimal pre-event and post-event strategies -- the former is expressed in terms of the solution to the HJB equation and the latter is traditional Merton's ratio.

math.OC

Policy Iteration Achieves Regularized Equilibrium under Time Inconsistency

For a general entropy-regularized time-inconsistent stochastic control problem, we propose a policy iteration algorithm (PIA) and establish its convergence to an equilibrium policy with an exponential convergence rate. The design of the PIA is based on a coupled system of non-local partial differential equations, called the exploratory equilibrium Hamilton--Jacobi--Bellman (EEHJB) equation. As opposed to the standard time-consistent case, policy improvement fails in general and the target value function (now an equilibrium value function) is not even known to exist a priori. To overcome these, we prove that the value functions generated by the PIA form a Cauchy sequence in a specialized Banach space, hence admit a limit, and the rate of convergence is exponential, on the strength of the Bismut--Elworthy--Li formula of stochastic representation. The limiting value function is shown to fulfill the EEHJB equation, which induces an equilibrium policy in a Gibbs form. Such convergence in value additionally implies uniform convergence of the generated policies to the equilibrium policy, again with an exponential rate. As a byproduct, the PIA gives a constructive proof of the global existence and uniqueness of a classical solution to our general EEHJB equation, whose well-posedness has not been explored in the literature.

math.OC

On optimal solutions of classical and sliced Wasserstein GANs with non-Gaussian data

The generative adversarial network (GAN) aims to approximate an unknown distribution via a parameterized neural network (NN). While GANs have been widely applied in reinforcement and semi-supervised learning as well as computer vision tasks, selecting their parameters often needs an exhaustive search, and only a few selection methods have been proven to be theoretically optimal. One of the most promising GAN variants is the Wasserstein GAN (WGAN). Prior work on optimal parameters for population WGAN is limited to the linear-quadratic-Gaussian (LQG) setting, where the generator NN is linear, and the data is Gaussian. In this paper, we focus on the characterization of optimal solutions of population WGAN beyond the LQG setting. As a basic result, closed-form optimal parameters for one-dimensional WGAN are derived when the NN has non-linear activation functions, and the data is non-Gaussian. For high-dimensional data, we adopt the sliced Wasserstein framework and show that the linear generator can be asymptotically optimal. Moreover, the original sliced WGAN only constrains the projected data marginal instead of the whole one in classical WGAN, and thus, we propose another new unprojected sliced WGAN and identify its asymptotic optimality. Empirical studies show that compared to the celebrated r-principal component analysis (r-PCA) solution, which has cubic complexity to the data dimension, our generator for sliced WGAN can achieve better performance with only linear complexity.

cs.LG

Mean-Variance Stackelberg Games with Asymmetric Information

This paper considers two investors who perform mean-variance portfolio selection with asymmetric information: one knows the true stock dynamics, while the other has to infer the true dynamics from observed stock evolution. Their portfolio selection is interconnected through relative performance concerns, i.e., each investor is concerned about not only her terminal wealth, but how it compares to the average terminal wealth of both investors. We model this as Stackelberg competition: the partially-informed investor (the "follower") observes the trading behavior of the fully-informed investor (the "leader") and decides her trading strategy accordingly; the leader, anticipating the follower's response, in turn selects a trading strategy that best suits her objective. To prevent information leakage, the leader adopts a randomized strategy selected under an entropy-regularized mean-variance objective, where the entropy regularizer quantifies the randomness of a chosen strategy. The follower, on the other hand, observes only the actual trading actions of the leader (sampled from the randomized strategy), but not the randomized strategy itself. Her mean-variance objective is thus a random field, in the form of an expectation conditioned on a realized path of the leader's trading actions. In the idealized case of continuous sampling of the leader's trading actions, we derive a Stackelberg equilibrium where the follower's trading strategy depends linearly on the actual trading actions of the leader and the leader samples her trading actions from Gaussian distributions. In the realistic case of discrete sampling of the leader's trading actions, the above becomes an $\epsilon$-Stackelberg equilibrium.

q-fin.MF

Mean-Field Langevin Diffusions with Density-dependent Temperature

In the context of non-convex optimization, we let the temperature of a Langevin diffusion to depend on the diffusion's own density function. The rationale is that the induced density captures to some extent the landscape imposed by the non-convex function to be minimized, such that a density-dependent temperature provides location-wise random perturbation that may better react to, for instance, the location and depth of local minimizers. As the Langevin dynamics is now self-regulated by its own density, it forms a mean-field stochastic differential equation (SDE) of the Nemytskii type, distinct from the standard McKean-Vlasov equations. Relying on Wasserstein subdifferential calculus, we first show that the corresponding (nonlinear) Fokker-Planck equation has a unique solution. Next, a weak solution to the SDE is constructed from the solution to the Fokker-Planck equation, by Trevisan's superposition principle. As time goes to infinity, we further show that the induced density converges to an invariant distribution, which admits an explicit formula in terms of the Lambert $W$ function. A numerical example suggests that the density-dependent temperature can simultaneously improve the accuracy of and rate of convergence to the estimate of global minimizers.

math.OC

Generative Modeling by Minimizing the Wasserstein-2 Loss

This paper develops a generative model by minimizing the second-order Wasserstein loss (the $W_2$ loss) through a distribution-dependent ordinary differential equation (ODE), whose dynamics involves the Kantorovich potential associated with the true data distribution and a current estimate of it. A main result shows that the time-marginal laws of the ODE form a gradient flow for the $W_2$ loss, which converges exponentially to the true data distribution. An Euler scheme for the ODE is proposed and it is shown to recover the gradient flow for the $W_2$ loss in the limit. An algorithm is designed by following the scheme and applying persistent training, which naturally fits our gradient-flow approach. In both low- and high-dimensional experiments, our algorithm outperforms Wasserstein generative adversarial networks by increasing the level of persistent training appropriately.

stat.ML

A Differential Equation Approach for Wasserstein GANs and Beyond

This paper proposes a new theoretical lens to view Wasserstein generative adversarial networks (WGANs). To minimize the Wasserstein-1 distance between the true data distribution and our estimate of it, we derive a distribution-dependent ordinary differential equation (ODE) which represents the gradient flow of the Wasserstein-1 loss, and show that a forward Euler discretization of the ODE converges. This inspires a new class of generative models that naturally integrates persistent training (which we call W1-FE). When persistent training is turned off, we prove that W1-FE reduces to WGAN. When we intensify persistent training, W1-FE is shown to outperform WGAN in training experiments from low to high dimensions, in terms of both convergence speed and training results. Intriguingly, one can reap the benefits only when persistent training is carefully integrated through our ODE perspective. As demonstrated numerically, a naive inclusion of persistent training in WGAN (without relying on our ODE framework) can significantly worsen training results.

stat.ML

Partial Information in a Mean-Variance Portfolio Selection Game

This paper considers finitely many investors who perform mean-variance portfolio selection under relative performance criteria. That is, each investor is concerned about not only her terminal wealth, but how it compares to the average terminal wealth of all investors. At the inter-personal level, each investor selects a trading strategy in response to others' strategies. This selected strategy additionally needs to yield an equilibrium intra-personally, so as to resolve time inconsistency among the investor's current and future selves (triggered by the mean-variance objective). A Nash equilibrium we look for is thus a tuple of trading strategies under which every investor achieves her intra-personal equilibrium simultaneously. We derive such a Nash equilibrium explicitly in the idealized case of full information (i.e., the dynamics of the underlying stock is perfectly known) and semi-explicitly in the realistic case of partial information (i.e., the stock evolution is observed, but the expected return of the stock is not precisely known). The formula under partial information consists of the myopic trading and intertemporal hedging terms, both of which depend on an additional state process that serves to filter the true expected return and whose influence on trading is captured by a degenerate Cauchy problem. Our results identify that relative performance criteria can induce downward self-reinforcement of investors' wealth--if every investor suffers a wealth decline simultaneously, then everyone's wealth tends to decline further. This phenomenon, as numerical examples show, is negligible under full information but pronounced under partial information.

q-fin.MF

Relaxed Equilibria for Time-Inconsistent Markov Decision Processes

This paper considers an infinite-horizon Markov decision process (MDP) that allows for general non-exponential discount functions, in both discrete and continuous time. Due to the inherent time inconsistency, we look for a randomized equilibrium policy (i.e., relaxed equilibrium) in an intra-personal game between an agent's current and future selves. When we modify the MDP by entropy regularization, a relaxed equilibrium is shown to exist by a nontrivial entropy estimate. As the degree of regularization diminishes, the entropy-regularized MDPs approximate the original MDP, which gives the general existence of a relaxed equilibrium in the limit by weak convergence arguments. As opposed to prior studies that consider only deterministic policies, our existence of an equilibrium does not require any convexity (or concavity) of the controlled transition probabilities and reward function. Interestingly, this benefit of considering randomized policies is unique to the time-inconsistent case.

math.OC

Epstein-Zin Utility Maximization on a Random Horizon

This paper solves the consumption-investment problem under Epstein-Zin preferences on a random horizon. In an incomplete market, we take the random horizon to be a stopping time adapted to the market filtration, generated by all observable, but not necessarily tradable, state processes. Contrary to prior studies, we do not impose any fixed upper bound for the random horizon, allowing for truly unbounded ones. Focusing on the empirically relevant case where the risk aversion and the elasticity of intertemporal substitution are both larger than one, we characterize the optimal consumption and investment strategies using backward stochastic differential equations with superlinear growth on unbounded random horizons. This characterization, compared with the classical fixed-horizon result, involves an additional stochastic process that serves to capture the randomness of the horizon. As demonstrated in two concrete examples, changing from a fixed horizon to a random one drastically alters the optimal strategies.

q-fin.MF

GANs as Gradient Flows that Converge

This paper approaches the unsupervised learning problem by gradient descent in the space of probability density functions. A main result shows that along the gradient flow induced by a distribution-dependent ordinary differential equation (ODE), the unknown data distribution emerges as the long-time limit. That is, one can uncover the data distribution by simulating the distribution-dependent ODE. Intriguingly, the simulation of the ODE is shown equivalent to the training of generative adversarial networks (GANs). This equivalence provides a new "cooperative" view of GANs and, more importantly, sheds new light on the divergence of GANs. In particular, it reveals that the GAN algorithm implicitly minimizes the mean squared error (MSE) between two sets of samples, and this MSE fitting alone can cause GANs to diverge. To construct a solution to the distribution-dependent ODE, we first show that the associated nonlinear Fokker-Planck equation has a unique weak solution, by the Crandall-Liggett theorem for differential equations in Banach spaces. Based on this solution to the Fokker-Planck equation, we construct a unique solution to the ODE, using Trevisan's superposition principle. The convergence of the induced gradient flow to the data distribution is obtained by analyzing the Fokker-Planck equation.

cs.LG

Convergence of Policy Iteration for Entropy-Regularized Stochastic Control Problems

For a general entropy-regularized stochastic control problem on an infinite horizon, we prove that a policy iteration algorithm (PIA) converges to an optimal relaxed control. Contrary to the standard stochastic control literature, classical H\"{o}lder estimates of value functions do not ensure the convergence of the PIA, due to the added entropy-regularizing term. To circumvent this, we carry out a delicate estimation by moving back and forth between appropriate H\"{o}lder and Sobolev spaces. This requires new Sobolev estimates designed specifically for the purpose of policy iteration and a nontrivial technique to contain the entropy growth. Ultimately, we obtain a uniform H\"{o}lder bound for the sequence of value functions generated by the PIA, thereby achieving the desired convergence result. Characterization of the optimal value function as the unique solution to an exploratory Hamilton-Jacobi-Bellman equation comes as a by-product. The PIA is numerically implemented in an example of optimal consumption.

math.OC

Minimizing the Repayment Cost of Federal Student Loans

Federal student loans are fixed-rate debt contracts with three main special features: (i) borrowers can use income-driven schemes to make payments proportional to their income above subsistence, (ii) after several years of good standing, the remaining balance is forgiven but taxed as ordinary income, and (iii) accrued interest is simple, i.e., not capitalized. For a very small loan, the cost-minimizing repayment strategy dictates maximum payments until full repayment, forgoing both income-driven schemes and forgiveness. For a very large loan, the minimal payments allowed by income-driven schemes are optimal. For intermediate balances, the optimal repayment strategy may entail an initial period of minimum payments to exploit the non-capitalization of accrued interest, but when the principal is being reimbursed maximal payments always precede minimum payments. Income-driven schemes and simple accrued interest mostly benefit borrowers with very large balances.

q-fin.MF

A Time-Inconsistent Dynkin Game: from Intra-personal to Inter-personal Equilibria

This paper studies a nonzero-sum Dynkin game in discrete time under non-exponential discounting. For both players, there are two levels of game-theoretic reasoning intertwined. First, each player looks for an intra-personal equilibrium among her current and future selves, so as to resolve time inconsistency triggered by non-exponential discounting. Next, given the other player's chosen stopping policy, each player selects a best response among her intra-personal equilibria. A resulting inter-personal equilibrium is then a Nash equilibrium between the two players, each of whom employs her best intra-personal equilibrium with respect to the other player's stopping policy. Under appropriate conditions, we show that an inter-personal equilibrium exists, based on concrete iterative procedures along with Zorn's lemma. To illustrate our theoretic results, we investigate a two-player real options valuation problem: two firms negotiate a deal of cooperation to initiate a project jointly. By deriving inter-personal equilibria explicitly, we find that coercive power in negotiation depends crucially on the impatience levels of the two firms.

math.OC

Mortality and Healthcare: a Stochastic Control Analysis under Epstein-Zin Preferences

This paper studies optimal consumption, investment, and healthcare spending under Epstein-Zin preferences. Given consumption and healthcare spending plans, Epstein-Zin utilities are defined over an agent's random lifetime, partially controllable by the agent as healthcare reduces mortality growth. To the best of our knowledge, this is the first time Epstein-Zin utilities are formulated on a controllable random horizon, via an infinite-horizon backward stochastic differential equation with superlinear growth. A new comparison result is established for the uniqueness of associated utility value processes. In a Black-Scholes market, the stochastic control problem is solved through the related Hamilton-Jacobi-Bellman (HJB) equation. The verification argument features a delicate containment of the growth of the controlled morality process, which is unique to our framework, relying on a combination of probabilistic arguments and analysis of the HJB equation. In contrast to prior work under time-separable utilities, Epstein-Zin preferences facilitate calibration. The model-generated mortality closely approximates actual mortality data in the US and UK; moreover, the efficacy of healthcare can be calibrated and compared between the two countries.

q-fin.MF

Optimal Stopping under Model Ambiguity: a Time-Consistent Equilibrium Approach

An unconventional approach for optimal stopping under model ambiguity is introduced. Besides ambiguity itself, we take into account how ambiguity-averse an agent is. This inclusion of ambiguity attitude, via an $α$-maxmin nonlinear expectation, renders the stopping problem time-inconsistent. We look for subgame perfect equilibrium stopping policies, formulated as fixed points of an operator. For a one-dimensional diffusion with drift and volatility uncertainty, we show that every equilibrium can be obtained through a fixed-point iteration. This allows us to capture much more diverse behavior, depending on an agent's ambiguity attitude, beyond the standard worst-case (or best-case) analysis. In a concrete example of real options valuation under volatility uncertainty, all equilibrium stopping policies, as well as the best one among them, are fully characterized. It demonstrates explicitly the effect of ambiguity attitude on decision making: the more ambiguity-averse, the more eager to stop -- so as to withdraw from the uncertain environment. The main result hinges on a delicate analysis of continuous sample paths in the canonical space and the capacity theory. To resolve measurability issues, a generalized measurable projection theorem, new to the literature, is also established.

q-fin.MF

Optimal Equilibria for Multi-dimensional Time-inconsistent Stopping Problems

We study an optimal stopping problem under non-exponential discounting, where the state process is a multi-dimensional continuous strong Markov process. The discount function is taken to be log sub-additive, capturing decreasing impatience in behavioral economics. On strength of probabilistic potential theory, we establish the existence of an optimal equilibrium among a sufficiently large collection of equilibria, consisting of finely closed equilibria satisfying a boundary condition. This generalizes the existence of optimal equilibria for one-dimensional stopping problems in prior literature.

q-fin.MF

Generalized Duality for Model-Free Superhedging given Marginals

In a discrete-time financial market, a generalized duality is established for model-free superhedging, given marginal distributions of the underlying asset. Contrary to prior studies, we do not require contingent claims to be upper semicontinuous, allowing for upper semi-analytic ones. The generalized duality stipulates an extended version of risk-neutral pricing. To compute the model-free superhedging price, one needs to find the supremum of expected values of a contingent claim, evaluated not directly under martingale (risk-neutral) measures, but along sequences of measures that converge, in an appropriate sense, to martingale ones. To derive the main result, we first establish a portfolio-constrained duality for upper semi-analytic contingent claims, relying on Choquet's capacitability theorem. As we gradually fade out the portfolio constraint, the generalized duality emerges through delicate probabilistic estimations.

q-fin.PR