Searcharxiv⌕ Search

arXiv subjects

Kyunghyun Park

Publications and source records attributed to Kyunghyun Park.

12 recordsLinked to original sources

Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise

In this article, we present a robust $Q$-learning algorithm for discrete-time mean-field control problems under Wasserstein uncertainty in the common noise law. The algorithm combines a quantization-and-projection scheme with a Wasserstein dual reformulation on the common-noise space. We establish its convergence together with finite-time iteration bounds for both synchronous and asynchronous learning schemes. Numerical experiments on systemic risk and epidemic models compare the asynchronous implementation with an idealized Bellman iteration, illustrate the robustness-performance tradeoff under common-noise misspecification, and report the observed convergence behavior of the asynchronous $Q$-learning algorithm.

math.OC↗

Scaling limits of multi-period distributionally robust optimization problems

We examine the scaling limit of multi-period distributionally robust optimization (DRO) problems via a semigroup approach. Each period involves a worst-case maximization over distributions in a Wasserstein ball around the transition probability of a reference process with radius proportional to the length of the period, and the multi-period DRO problem arises through its sequential composition. We show that the scaling limit of the multi-period DRO, as the length of each period tends to zero, is a strongly continuous monotone semigroup on $\mathrm{C_b}$. Furthermore, we show that its infinitesimal generator is equal to the generator associated with the non-robust scaling limit plus an additional perturbation term induced by the Wasserstein uncertainty. As an application, we show that when the reference process follows an Itô process, the viscosity solution of the associated nonlinear PDE coincides with the value of continuous-time robust optimization problems under parametric uncertainty.

math.OC↗

IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents

Large language model (LLM)-based agents solve complex tasks by leveraging multi-step reasoning with iterative tool calls and environment interactions, which incur idle time while waiting for observations. Despite the prevalence of idle time in most agentic scenarios, existing works treat it as an unavoidable overhead or propose restricted solutions that overlook varying computational budgets across different tool calls and future observation uncertainty, thereby leading to suboptimal utilization of idle time. In this paper, we introduce IdleSpec, a scalable and generic inference approach that leverages idle-time computation to improve agent performance while minimizing latency overhead. Specifically, IdleSpec iteratively generates plan candidates during idle periods and, once observations become available, aggregates them to guide the next reasoning step. For effective plan generation under observation uncertainty, IdleSpec samples between complementary drafting strategies (i.e., progressive and recovery) from a learned distribution that is updated via posterior feedback. Our experiments demonstrate that IdleSpec significantly improves agent performance in various agentic scenarios by effectively utilizing idle time. In particular, on the GAIA and FRAMES, IdleSpec achieves 55.6% average accuracy with Gemini-2.5-Flash, surpassing the vanilla baseline without idle-time usage by 5.1%. Furthermore, for MLE-Bench, which involves substantial delay from code executions, IdleSpec achieves performance gains of up to 9.1% on the Any Medal rate, highlighting its generalizability to long-horizon tasks.

cs.AI↗

Robust dividend policy: Equivalence of Epstein-Zin and Maenhout preferences

In a continuous-time economy, this paper formulates the Epstein-Zin preference for discounted dividends received by an investor as an Epstein-Zin singular control utility. We introduce a backward stochastic differential equation with an aggregator integrated with respect to a singular control, prove its well-posedness, and show that it coincides with the Epstein-Zin singular control utility. We then establish that this formulation is equivalent to a robust dividend policy chosen by the firm's executive under the Maenhout's ambiguity-averse preference. In particular, the robust dividend policy takes the form of a threshold strategy on the firm's surplus process, where the threshold level is characterized as the free boundary of a Hamilton-Jacobi-Bellman variational inequality. Therefore, dividend-caring investors can choose firms that match their preferences by examining stock's dividend policies and financial statements, whereas executives can make use of dividend to signal their confidence, in the form of ambiguity aversion, on realizing the earnings implied by their financial statements.

q-fin.MF↗

Robust Exploratory Stopping under Ambiguity in Reinforcement Learning

We propose and analyze a continuous-time robust reinforcement learning framework for optimal stopping under ambiguity. In this framework, an agent chooses a robust exploratory stopping time motivated by two objectives: robust decision-making under ambiguity and learning about the unknown environment. Here, ambiguity refers to considering multiple probability measures dominated by a reference measure, reflecting the agent's awareness that the reference measure representing her learned belief about the environment would be erroneous. Using the $g$-expectation framework, we reformulate the optimal stopping problem under ambiguity as a robust exploratory control problem with Bernoulli distributed controls. We then characterize the optimal Bernoulli distributed control via backward stochastic differential equations and, based on this, construct the robust exploratory stopping time that approximates the optimal stopping time under ambiguity. Last, we establish a policy iteration theorem and implement it as a reinforcement learning algorithm. Numerical experiments demonstrate the convergence, robustness, and scalability of our reinforcement learning algorithm across different levels of ambiguity and exploration.

math.OC↗

Numerical method for nonlinear Kolmogorov PDEs via sensitivity analysis

We examine nonlinear Kolmogorov partial differential equations (PDEs). Here the nonlinear part of the PDE comes from its Hamiltonian where one maximizes over all possible drift and diffusion coefficients which fall within a $\varepsilon$-neighborhood of pre-specified baseline coefficients. Our goal is to quantify and compute how sensitive those PDEs are to such a small nonlinearity, and then use the results to develop an efficient numerical method for their approximation. We show that as $\varepsilon\downarrow 0$, the nonlinear Kolmogorov PDE equals the linear Kolmogorov PDE defined with respect to the corresponding baseline coefficients plus $\varepsilon$ times a correction term which can be also characterized by the solution of another linear Kolmogorov PDE involving the baseline coefficients. As these linear Kolmogorov PDEs can be efficiently solved in high-dimensions by exploiting their Feynman-Kac representation, our derived sensitivity analysis then provides a Monte Carlo based numerical method which can efficiently solve these nonlinear Kolmogorov equations. We establish an error and complexity analysis for our numerical method. Moreover, we provide numerical examples in up to 100 dimensions to empirically demonstrate the applicability of our numerical method.

math.NA↗

Robust mean-field control under common noise uncertainty

We propose and analyze a framework for discrete-time robust mean-field control problems under common noise uncertainty. In this framework, the mean-field interaction describes the collective behavior of infinitely many cooperative agents' state and action, while the common noise -- a random disturbance affecting all agents' state dynamics -- is uncertain. A social planner optimizes over open-loop controls on an infinite horizon to maximize the representative agent's worst-case expected reward, where worst-case corresponds to the most adverse probability measure among all candidates inducing the unknown true law of the common noise process. We refer to this optimization as a robust mean-field control problem under common noise uncertainty. We first show that this problem arises as the asymptotic limit of a cooperative $N$-agent robust optimization problem, commonly known as propagation of chaos. We then prove the existence of an optimal open-loop control by linking the robust mean field control problem to a lifted robust Markov decision problem on the space of probability measures and by establishing the dynamic programming principle and Bellman--Isaac fixed point theorem for the lifted robust Markov decision problem. Finally, we complement our theoretical results with numerical experiments motivated by distribution planning and systemic risk in finance, highlighting the advantages of accounting for common noise uncertainty.

math.OC↗

Sensitivity of robust optimization problems under drift and volatility uncertainty

We examine optimization problems in which an investor has the opportunity to trade in $d$ stocks with the goal of maximizing her worst-case cost of cumulative gains and losses. Here, worst-case refers to taking into account all possible drift and volatility processes for the stocks that fall within a $\varepsilon$-neighborhood of predefined fixed baseline processes. Although solving the worst-case problem for a fixed $\varepsilon>0$ is known to be very challenging in general, we show that it can be approximated as $\varepsilon\to 0$ by the baseline problem (computed using the baseline processes) in the following sense: Firstly, the value of the worst-case problem is equal to the value of the baseline problem plus $\varepsilon$ times a correction term. This correction term can be computed explicitly and quantifies how sensitive a given optimization problem is to model uncertainty. Moreover, approximately optimal trading strategies for the worst-case problem can be obtained using optimal strategies from the corresponding baseline problem.

math.OC↗

Markov-Nash equilibria in mean-field games under model uncertainty

We propose and analyze a framework for mean-field Markov games under model uncertainty. In this framework, a state-measure flow describing the collective behavior of a population affects the given reward function as well as the unknown transition kernel of the representative agent. The agent's objective is to choose an optimal Markov policy in order to maximize her worst-case expected reward, where worst-case refers to the most adverse scenario among all transition kernels considered to be feasible to describe the unknown true law of the environment. We prove the existence of a mean-field equilibrium under model uncertainty, where the agent chooses the optimal policy that maximizes the worst-case expected reward, and the state-measure flow aligns with the agent's state distribution under the optimal policy and the worst-case transition kernel. Moreover, we prove that for suitable multi-agent Markov games under model uncertainty the optimal policy from the mean-field equilibrium forms an approximate Markov-Nash equilibrium whenever the number of agents is large enough.

math.OC↗

Feature-aligned N-BEATS with Sinkhorn divergence

We propose Feature-aligned N-BEATS as a domain-generalized time series forecasting model. It is a nontrivial extension of N-BEATS with doubly residual stacking principle (Oreshkin et al. [45]) into a representation learning framework. In particular, it revolves around marginal feature probability measures induced by the intricate composition of residual and feature extracting operators of N-BEATS in each stack and aligns them stack-wise via an approximate of an optimal transport distance referred to as the Sinkhorn divergence. The training loss consists of an empirical risk minimization from multiple source domains, i.e., forecasting loss, and an alignment loss calculated with the Sinkhorn divergence, which allows the model to learn invariant features stack-wise across multiple source data sequences while retaining N-BEATS's interpretable design and forecasting power. Comprehensive experimental evaluations with ablation studies are provided and the corresponding results demonstrate the proposed model's forecasting and generalization capabilities.

cs.LG↗

Extensive networks would eliminate the demand for pricing formulas

In this study, we generate a large number of implied volatilities for the Stochastic Alpha Beta Rho (SABR) model using a graphics processing unit (GPU) based simulation and enable an extensive neural network to learn them. This model does not have any exact pricing formulas for vanilla options, and neural networks have an outstanding ability to approximate various functions. Surprisingly, the network reduces the simulation noises by itself, thereby achieving as much accuracy as the Monte-Carlo simulation. Extremely high accuracy cannot be attained via existing approximate formulas. Moreover, the network is as efficient as the approaches based on the formulas. When evaluating based on high accuracy and efficiency, extensive networks can eliminate the necessity of the pricing formulas for the SABR model. Another significant contribution is that a novel method is proposed to examine the errors based on nonlinear regression. This approach is easily extendable to other pricing models for which it is hard to induce analytic formulas.

q-fin.CP↗

Optimal Insurance with Limited Commitment in a Finite Horizon

We study a finite horizon optimal contracting problem of a risk-neutral principal and a risk-averse agent who receives a stochastic income stream when the agent is unable to make commitments. The problem involves an infinite number of constraints at each time and each state of the world. Miao and Zhang (2015) have developed a dual approach to the problem by considering a Lagrangian and derived a Hamilton-Jacobi-Bellman equation in an infinite horizon. We consider a similar Lagrangian in a finite horizon, but transform the dual problem into an infinite series of optimal stopping problems. For each optimal stopping problem we provide an analytic solution by providing an integral equation representation for the free boundary. We provide a verification theorem that the value function of the original principal's problem is the Legender-Fenchel transform of the integral of the value functions of the optimal stopping problems. We also provide some numerical simulation results of optimal contracting strategies

econ.TH↗