SearcharxivSearch

arXiv subjects

Jeonggyu Huh

Publications and source records attributed to Jeonggyu Huh.

15 recordsLinked to original sources

Self-Consistent Adjoint Policy Iteration for Constrained Dynamic Portfolio Choice

We develop simulation-based policy iteration for continuous-time portfolio choice with predictable returns and convex constraints. Each outer step re-evaluates a fixed-latent open-loop backpropagation-through-time (OL-BPTT) adjoint after deployment and solves the constrained update. Shifted-adjoint cancellation controls the adjoint--HJB Hamiltonian-gradient discrepancy by the policy-improvement residual. For CRRA portfolios, exact HJB policy iteration identifies the optimal reduced value factor, while population OL-BPTT iteration converges globally when the adjoint update is directionally improving and approximate stationarity is asymptotically HJB-compatible. A theorem-matched occupation audit yields maximal $95\%$ upper endpoints of $0.066$ for the primitive directional ratio and $0.074$ for a stronger norm-relative ratio, both against the half-step threshold $0.75$. In the high-precision $50$--$50$ occupancy/broad-anchor design of a three-factor, fifty-asset benchmark, current-policy re-evaluation outperforms matched pooled refinement under the on-policy and broad evaluation laws.

math.OC

Scalable Pontryagin-Guided Adjoint-to-Control Recovery for Constrained Dynamic Portfolio Choice

We study continuous-time multi-asset portfolio choice and consumption under smooth pointwise constraints, including state-dependent feasible sets. The method separates dynamic information acquisition from local constrained recovery. A pointwise-feasible neural actor generates reference rollouts; after training, its realized latent outputs are frozen and first- and second-order adjoints are harvested from a fixed-latent open-loop backpropagation-through-time graph. Feedback therefore generates the reference trajectory without restricting the adjoint formulation to Markov controls. Conditional on the harvested adjoints, deployment solves a local generalized Pontryagin-Hamiltonian problem: quadratic-affine portfolio blocks are recovered exactly by a quadratic program, while a log barrier approximates more general regular KKT branches. We establish local chart representations, an OL-BPTT-to-adjoint correspondence retaining orthogonal martingale residuals, and an end-to-end bound from reference value loss and numerical errors to recovered-policy and local QP-gap errors. Analytical constant- and predictable-opportunity benchmarks validate the adjoints. Common-input experiments show that recovery reduces residual PMP/KKT error left by finite-budget direct policy optimization, including under a state-dependent consumption cap and with up to 100 risky assets. Scalability concerns the constrained action block rather than dimension-free state-space complexity.

q-fin.PM

From Value Bounds to Policy-Distance and Active-Face Certificates: Same-Grid Duality for Constrained Dynamic Portfolios

Neural and numerical policy solvers can produce feasible controls even when the optimal rule and its binding constraints are unavailable. A primal-dual bracket certifies value loss, but it does not locate the optimal policy or explain which constraints genuinely bind. We show that, on the same declared simulation grid, one bracket can support both conclusions. For polyhedral controls, an exact conditional budget identity rewrites the residual as a pathwise nonnegative terminal Fenchel defect plus date-by-constraint complementary-slackness terms. A canonical Doob compensation removes the budget martingale that obscures small residuals. Bellman-primitive curvature conditions then yield an occupancy-weighted policy region with the sharp O(sqrt(G)) radius, while a paired constraint relaxation lower-bounds the optimal multiplier and certifies a binding face. A finite-sample resolution theorem quantifies the path budget needed to certify a target policy tolerance or face. Locked one-asset and two-asset audits cover every external policy error and make no false face declaration. An exact-wrapper stress test remains tight through 50 assets, while a separate state-dependent pilot identifies learned-dual tightness as the high-dimensional bottleneck. Reference solutions enter only after certification.

q-fin.MF

Beyond the Bellman Recursion: A Pontryagin-Guided Framework for Non-Exponential Discounting

Most value-based and actor--critic reinforcement learning methods rely on Bellman-style recursions, yet these recursions collapse under non-exponential discounting common in human preferences and survival processes. We show the breakdown is structural: exponential discounting sits at a fragile intersection of multiplicativity and time homogeneity, and violating either property breaks standard dynamic programming. To overcome this, we propose Pontryagin-Guided Direct Policy Optimization (PG-DPO), a variational framework that abandons recursion and couples the Pontryagin Maximum Principle with Monte Carlo rollouts via an Adjoint-MC projection enforcing pointwise Hamiltonian maximization. Across multi-dimensional hyperbolic and survival-discount benchmarks, PG-DPO improves accuracy and stability where equation-driven solvers and critic-based baselines diverge.

cs.LG

Breaking the Dimensional Barrier: Dynamic Portfolio Choice with Parameter Uncertainty via Pontryagin Projection

We study continuous-time CRRA portfolio choice in diffusion markets with estimated and hence uncertain coefficients. Nature draws a latent parameter $θ\sim q$ at time $0$ and keeps it fixed; the investor never observes $θ$ and must commit to a single $θ$-blind policy maximizing an ex-ante objective, treating $q$ as a decision-time input. We propose a simulation-only two-stage solver.Stage 1 (DPO) performs BPTT-based stochastic gradient ascent through an Euler simulator while sampling $θ$ only inside the simulator. Stage 2 (Pontryagin projection) aggregates costate blocks across $θ\sim q$ and enforces the $q$-aggregated stationarity condition within the deployable class; the resulting correction can be amortized via interactive distillation. We refer to the full Stage 1 + Stage 2 pipeline as PG-DPO.We prove a uniform conditional BPTT-PMP correspondence and a residual-based policy-gap bound with explicit discretization and Monte Carlo error terms. Experiments on high-dimensional Gaussian drift-uncertainty and factor-driven benchmarks show that projection stabilizes learning and accurately recovers analytic decision-time references, while a model-free PPO baseline remains far from the targets.

q-fin.CP

MarketGANs: Multivariate financial time-series data augmentation using generative adversarial networks

This paper introduces MarketGAN, a factor-based generative framework for high-dimensional asset return generation under severe data scarcity. We embed an explicit asset-pricing factor structure as an economic inductive bias and generate returns as a single joint vector, thereby preserving cross-sectional dependence and tail co-movement alongside inter-temporal dynamics. MarketGAN employs generative adversarial learning with a temporal convolutional network (TCN) backbone, which models stochastic, time-varying factor loadings and volatilities and captures long-range temporal dependence. Using daily returns of large U.S. equities, we find that MarketGAN more closely matches empirical stylized facts of asset returns, including heavy-tailed marginal distributions, volatility clustering, leverage effects, and, most notably, high-dimensional cross-sectional correlation structures and tail co-movement across assets, than conventional factor-model-based bootstrap approaches. In portfolio applications, covariance estimates derived from MarketGAN-generated samples outperform those derived from other methods when factor information is at least weakly informative, demonstrating tangible economic value.

q-fin.ST

Breaking the Dimensional Barrier for Constrained Dynamic Portfolio Choice

We propose a scalable, policy-centric framework for continuous-time multi-asset portfolio-consumption optimization under inequality constraints. Our method integrates neural policies with Pontryagin's Maximum Principle (PMP) and enforces feasibility by maximizing a log-barrier-regularized Hamiltonian at each time-state pair, thereby satisfying KKT conditions without value-function grids. Theoretically, we show that the barrier-regularized Hamiltonian yields O($ε$) policy error and a linear Hamiltonian gap (quadratic when the KKT solution is interior), and we extend the BPTT-PMP correspondence to constrained settings with stable costate convergence. Empirically, PG-DPO and its projected variant (P-PGDPO) recover KKT-optimal policies in canonical short-sale and consumption-cap problems while maintaining strict feasibility across dimensions; unlike PDE/BSDE solvers, runtime scales linearly with the number of assets and remains practical at n=100. These results provide a rigorous and scalable foundation for high-dimensional constrained continuous-time portfolio optimization.

q-fin.PM

Breaking the Dimensional Barrier: A Pontryagin-Guided Direct Policy Optimization for Continuous-Time Multi-Asset Portfolio Choice

We introduce the Pontryagin-Guided Direct Policy Optimization (PG-DPO) framework for high-dimensional continuous-time portfolio choice. Our approach combines Pontryagin's Maximum Principle (PMP) with backpropagation through time (BPTT) to directly inform neural network-based policy learning, enabling accurate recovery of both myopic and intertemporal hedging demands--an aspect often missed by existing methods. Building on this, we develop the Projected PG-DPO (P-PGDPO) variant, which achieves nearoptimal policies with substantially improved efficiency. P-PGDPO leverages rapidly stabilizing costate estimates from BPTT and analytically projects them onto PMP's first-order conditions, reducing training overhead while improving precision. Numerical experiments show that PG-DPO matches or exceeds the accuracy of Deep BSDE, while P-PGDPO delivers significantly higher precision and scalability. By explicitly incorporating time-to-maturity, our framework naturally applies to finite-horizon problems and captures horizon-dependent effects, with the long-horizon case emerging as a stationary special case.

q-fin.PM

Pontryagin-Guided Policy Optimization for Merton's Portfolio Problem

We present a Pontryagin-Guided Direct Policy Optimization (PG-DPO) framework for Merton's portfolio problem, unifying modern neural-network-based policy parameterization with the adjoint viewpoint from Pontryagin's maximum principle (PMP). Instead of approximating the value function (as done in deep BSDE methods), we track a policy-fixed BSDE for the adjoint processes, which allows each gradient update to align with continuous-time PMP conditions. This setup yields locally optimal consumption and investment policies that are closely tied to classical stochastic control. We further incorporate an alignment penalty that nudges the learned policy toward Pontryagin-derived solutions, enhancing both convergence speed and training stability. Numerical experiments confirm that PG-DPO effectively handles both consumption and investment, achieving strong performance and interpretability without requiring large offline datasets or model-free reinforcement learning.

math.OC

Tighter 'uniform bounds for Black-Scholes implied volatility' and the applications to root-finding

Using the option delta systematically, we derive tighter lower and upper bounds of the Black-Scholes implied volatility than those in Tehranchi [SIAM J. Financ. Math. 7 (2016), 893-916]. As an application, we propose a Newton-Raphson algorithm on the log price that converges rapidly for all price ranges when using a new lower bound as an initial guess. Our new algorithm is a better alternative to the widely used naive Newton-Raphson algorithm, whose convergence is slow for extreme option prices.

q-fin.MF

Newton Raphson Emulation Network for Highly Efficient Computation of Numerous Implied Volatilities

In finance, implied volatility is an important indicator that reflects the market situation immediately. Many practitioners estimate volatility using iteration methods, such as the Newton--Raphson (NR) method. However, if numerous implied volatilities must be computed frequently, the iteration methods easily reach the processing speed limit. Therefore, we emulate the NR method as a network using PyTorch, a well-known deep learning package, and optimize the network further using TensorRT, a package for optimizing deep learning models. Comparing the optimized emulation method with the NR function in SciPy, a popular implementation of the NR method, we demonstrate that the emulation network is up to 1,000 times faster than the benchmark function.

q-fin.CP

Extensive networks would eliminate the demand for pricing formulas

In this study, we generate a large number of implied volatilities for the Stochastic Alpha Beta Rho (SABR) model using a graphics processing unit (GPU) based simulation and enable an extensive neural network to learn them. This model does not have any exact pricing formulas for vanilla options, and neural networks have an outstanding ability to approximate various functions. Surprisingly, the network reduces the simulation noises by itself, thereby achieving as much accuracy as the Monte-Carlo simulation. Extremely high accuracy cannot be attained via existing approximate formulas. Moreover, the network is as efficient as the approaches based on the formulas. When evaluating based on high accuracy and efficiency, extensive networks can eliminate the necessity of the pricing formulas for the SABR model. Another significant contribution is that a novel method is proposed to examine the errors based on nonlinear regression. This approach is easily extendable to other pricing models for which it is hard to induce analytic formulas.

q-fin.CP

Consistent and Efficient Pricing of SPX and VIX Options under Multiscale Stochastic Volatility

This study provides a consistent and efficient pricing method for both Standard & Poor's 500 Index (SPX) options and the Chicago Board Options Exchange's Volatility Index (VIX) options under a multiscale stochastic volatility model. To capture the multiscale volatility of the financial market, our model adds a fast scale factor to the well-known Heston volatility and we derive approximate analytic pricing formulas for the options under the model. The analytic tractability can greatly improve the efficiency of calibration compared to fitting procedures with the finite difference method or Monte Carlo simulation. Our experiment using options data from 2016 to 2018 shows that the model reduces the errors on the training sets of the SPX and VIX options by 9.9% and 13.2%, respectively, and decreases the errors on the test sets of the SPX and VIX options by 13.0\% and 16.5\%, respectively, compared to the single-scale model of Heston. The error reduction is possible because the additional factor reflects short-term impacts on the market, which is difficult to achieve with only one factor. It highlights the necessity of modeling multiscale volatility.

q-fin.MF

Pricing Options with Exponential Levy Neural Network

In this paper, we propose the exponential Levy neural network (ELNN) for option pricing, which is a new non-parametric exponential Levy model using artificial neural networks (ANN). The ELNN fully integrates the ANNs with the exponential Levy model, a conventional pricing model. So, the ELNN can improve ANN-based models to avoid several essential issues such as unacceptable outcomes and inconsistent pricing of over-the-counter products. Moreover, the ELNN is the first applicable non-parametric exponential Levy model by virtue of outstanding researches on optimization in the field of ANN. The existing non-parametric models are too vulnerable to be employed in practice. The empirical tests with S\&P 500 option prices show that the ELNN outperforms two parametric models, the Merton and Kou models, in terms of fitting performance and stability of estimates.

q-fin.PR

Measuring Systematic Risk with Neural Network Factor Model

In this paper, we measure systematic risk with a new nonparametric factor model, the neural network factor model. The suitable factors for systematic risk can be naturally found by inserting daily returns on a wide range of assets into the bottleneck network. The network-based model does not stick to a probabilistic structure unlike parametric factor models, and it does not need feature engineering because it selects notable features by itself. In addition, we compare performance between our model and the existing models using 20-year data of S&P 100 components. Although the new model can not outperform the best ones among the parametric factor models due to limitations of the variational inference, the estimation method used for this study, it is still noteworthy in that it achieves the performance as best the comparable models could without any prior knowledge.

q-fin.CP