SearcharxivSearch

arXiv subjects

Chang-Han Rhee

Publications and source records attributed to Chang-Han Rhee.

17 recordsLinked to original sources

First-Exit Time Analysis for Truncated Heavy-Tailed Dynamical Systems

In this paper, we study the first-exit time of stochastic difference equation $X^\eta_{j+1}(x) = X^\eta_{j}(x) + \eta a\big( X^\eta_{j}(x)\big) + \eta \sigma\big( X^\eta_{j}(x)\big)Z_{j+1}$ and its truncated variant $X^{\eta|b}_{j+1}(x) = X^{\eta|b}_{j}( x) + \varphi_b\big(\eta a\big( X^{\eta|b}_{j}( x)\big) + \eta \sigma\big( X^{\eta|b}_{j}( x)\big) Z_{j+1}\big)$, where $\varphi_b(x) = (x/|x|)\min\{|x|, b\}$ and the law of the noise $Z_t$ is multivariate regularly varying. The truncation operator $\varphi_b(\cdot)$ is often introduced as a modulation mechanism in heavy-tailed systems, such as stochastic gradient descent algorithms in deep learning. By developing a framework that connects large deviations with metastability, we leverage the locally uniform sample-path large deviations for both processes in Wang and Rhee (2024) to obtain precise characterizations of the joint distributions of the first exit times and exit locations. The resulting limit theorem unveils a discrete hierarchy of phase transitions (i.e., exit times) as the truncation threshold $b$ varies, and manifests the catastrophe principle, whereby key events or metastable behaviors in heavy-tailed systems are driven by catastrophic behavior in a few components while the rest of the system behaves nominally. These developments lead to a comprehensive heavy-tailed counterpart of the classical Freidlin-Wentzell theory.

math.PR

Global Dynamics of Heavy-Tailed SGDs in Nonconvex Loss Landscape: Characterization and Control

Stochastic gradient descent (SGD) and its variants enable modern artificial intelligence. However, theoretical understanding lags far behind their empirical success. It is widely believed that SGD has a curious ability to avoid sharp local minima in the loss landscape, which are associated with poor generalization. To unravel this mystery and further enhance such capability of SGDs, it is imperative to go beyond the traditional local convergence analysis and obtain a comprehensive understanding of SGDs' global dynamics. In this paper, we develop a set of technical machinery based on the recent large deviations and metastability analysis in Wang and Rhee (2023) and obtain sharp characterization of the global dynamics of heavy-tailed SGDs. In particular, we reveal a fascinating phenomenon in deep learning: by injecting and then truncating heavy-tailed noises during the training phase, SGD can almost completely avoid sharp minima and achieve better generalization performance for the test data. Simulation and deep learning experiments confirm our theoretical prediction that heavy-tailed SGD with gradient clipping finds local minima with a more flat geometry and achieves better generalization performance.

cs.LG

Exit Time Analysis For Kesten's Stochastic Recurrence Equations

Kesten's stochastic recurrent equation is a classical subject of research in probability theory and its applications. Recently, it has garnered attention as a model for stochastic gradient descent with a quadratic objective function and the emergence of heavy-tailed dynamics in machine learning. This context calls for analysis of its asymptotic behavior under both negative and positive Lyapunov exponents. This paper studies the exit times of the Kesten's stochastic recurrence equation in both cases. Depending on the sign of Lyapunov exponent, the exit time scales either polynomially or logarithmically as the radius of the exit boundary increases.

math.PR

Sample-Path Large Deviations for L\'evy Processes and Random Walks with Lognormal Increments

The large deviations theory for heavy-tailed processes has seen significant advances in the recent past. In particular, Rhee et al. (2019) and Bazhba et al. (2020) established large deviation asymptotics at the sample-path level for L\'evy processes and random walks with regularly varying and (heavy-tailed) Weibull-type increments. This leaves the lognormal case -- one of the three most prominent classes of heavy-tailed distributions, alongside regular variation and Weibull -- open. This article establishes the \emph{extended large deviation principle} (extended LDP) at the sample-path level for one-dimensional L\'evy processes and random walks with lognormal-type increments. Building on these results, we also establish the extended LDPs for multi-dimensional processes with independent coordinates. We demonstrate the sharpness of these results by constructing counterexamples, thereby proving that our results cannot be strengthened to a standard LDP under $J_1$ topology and $M_1'$ topology.

math.PR

Strongly Efficient Rare-Event Simulation for Regularly Varying Lévy Processes with Infinite Activities

In this paper, we address rare-event simulation for heavy-tailed Lévy processes with infinite activities. The presence of infinite activities poses a critical challenge, making it impractical to simulate or store the precise sample path of the Lévy process. We present a rare-event simulation algorithm that incorporates an importance sampling strategy based on heavy-tailed large deviations, the stick-breaking approximation for the extrema of Lévy processes, the Asmussen-Rosiński approximation, and the randomized debiasing technique. By establishing a novel characterization for the Lipschitz continuity of the law of Lévy processes, we show that the proposed algorithm is unbiased and strongly efficient under mild conditions, and hence applicable to a broad class of Lévy processes. In numerical experiments, our algorithm demonstrates significant improvements in efficiency compared to the crude Monte-Carlo approach.

math.PR

Sample-path large deviations for a class of heavy-tailed Markov additive processes

For a class of additive processes driven by the affine recursion $X_{n+1} = A_n X_n + B_n$, we develop a sample-path large deviations principle in the $M_1'$ topology on $D [0,1]$. We allow $B_n$ to have both signs and focus on the case where Kesten's condition holds on $A_1$, leading to heavy-tailed distributions. The most likely paths in our large deviations results are step functions with both positive and negative jumps.

math.PR

Sample-path large deviations for unbounded additive functionals of the reflected random walk

We prove a sample path large deviation principle (LDP) with sub-linear speed for unbounded functionals of certain Markov chains induced by the Lindley recursion. The LDP holds in the Skorokhod space $\mathbb{D}[0,T]$ equipped with the $M_1'$ topology. Our technique hinges on a suitable decomposition of the Markov chain in terms of regeneration cycles. Each regeneration cycle denotes the area accumulated during the busy period of the reflected random walk. We prove a large deviation principle for the area under the busy period of the MRW, and we show that it exhibits a heavy-tailed behavior.

math.PR

Large Deviations and Metastability Analysis for Heavy-Tailed Dynamical Systems

This paper introduces novel frameworks for large deviations and metastability analysis in heavy-tailed stochastic dynamical systems. We develop and apply these frameworks within the context of stochastic difference equation $X^\eta_{j+1}(x) = X^\eta_{j}(x) + \eta a\big( X^\eta_{j}(x)\big) + \eta \sigma\big( X^\eta_{j}(x)\big)Z_{j+1}$ and its variation with truncated dynamics $X^{\eta|b}_{j+1}(x) = X^{\eta|b}_{j}( x) + \varphi_b\big(\eta a\big( X^{\eta|b}_{j}( x)\big) + \eta \sigma\big( X^{\eta|b}_{j}( x)\big) Z_{j+1}\big)$, where $\varphi_b(x) = (x/\|x\|)\max\{\|x\|, b\}$. The truncation operator $\varphi_b(\cdot)$ is often introduced as a modulation mechanism in heavy-tailed systems, such as stochastic gradient descent algorithms in deep learning. Thus, it is crucial to successfully analyze both $X^{\eta}_{j}(x)$ and $X^{\eta|b}_{j}(x)$. We establish locally uniform sample-path large deviations for both processes and translate these asymptotics into precise characterizations of the joint distributions of the first exit times and exit locations. Our large deviations asymptotics are sharp enough to rigorously characterize \emph{the catastrophe principle} by establishing the distributional limit of the sample paths conditional on the rare events of interest, thereby revealing the most likely paths through which rare events arise in heavy-tailed dynamical systems. Moreover the resulting limit theorem unveils a discrete hierarchy of phase transitions (i.e., exit times) as the truncation threshold $b$ varies. Together, these developments serve as a heavy-tailed counterpart of the classical Freidlin-Wentzell theory. We also present the corresponding results for continuous-time processes in the appendix.

math.PR

Large deviations for stochastic fluid networks with Weibullian tails

We consider a stochastic fluid network where the external input processes are compound Poisson with heavy-tailed Weibullian jumps. Our results comprise of large deviations estimates for the buffer content process in the vector-valued Skorokhod space which is endowed with the product $J_1$ topology. To illustrate our framework, we provide explicit results for a tandem queue. At the heart of our proof is a recent sample-path large deviations result, and a novel continuity result for the Skorokhod reflection map in the product $J_1$ topology.

math.PR

Eliminating Sharp Minima from SGD with Truncated Heavy-tailed Noise

The empirical success of deep learning is often attributed to SGD's mysterious ability to avoid sharp local minima in the loss landscape, as sharp minima are known to lead to poor generalization. Recently, empirical evidence of heavy-tailed gradient noise was reported in many deep learning tasks, and it was shown in Şimşekli (2019a,b) that SGD can escape sharp local minima under the presence of such heavy-tailed gradient noise, providing a partial solution to the mystery. In this work, we analyze a popular variant of SGD where gradients are truncated above a fixed threshold. We show that it achieves a stronger notion of avoiding sharp minima: it can effectively eliminate sharp local minima entirely from its training trajectory. We characterize the dynamics of truncated SGD driven by heavy-tailed noises. First, we show that the truncation threshold and width of the attraction field dictate the order of the first exit time from the associated local minimum. Moreover, when the objective function satisfies appropriate structural conditions, we prove that as the learning rate decreases, the dynamics of heavy-tailed truncated SGD closely resemble those of a continuous-time Markov chain that never visits any sharp minima. Real data experiments on deep learning confirm our theoretical prediction that heavy-tailed SGD with gradient clipping finds a "flatter" local minima and achieves better generalization.

cs.LG

Efficient Rare-Event Simulation for Multiple Jump Events in Regularly Varying Lévy Processes with Infinite Activities

In this paper we address the problem of rare-event simulation for heavy-tailed Lévy processes with infinite activities. We propose a strongly efficient importance sampling algorithm that builds upon the sample path large deviations for heavy-tailed Lévy processes, stick-breaking approximation of extrema of Lévy processes, and the randomized debiasing Monte Carlo scheme. The proposed importance sampling algorithm can be applied to a broad class of Lévy processes and exhibits significant improvements in efficiency when compared to crude Monte-Carlo method in our numerical experiments.

math.PR

Sample-path large deviations for Lévy processes and random walks with Weibull increments

We study sample-path large deviations for Lévy processes and random walks with heavy-tailed jump-size distributions that are of Weibull type. Our main results include an extended form of an LDP (large deviations principle) in the $J_1$ topology, and a full LDP in the $M_1'$ topology. The rate function can be represented as the solution to a quasi-variational problem. The sharpness and applicability of these results are illustrated by a counterexample proving the nonexistence of a full LDP in the $J_1$ topology, and by an application to a first passage problem.

math.PR

Space-filling design for nonlinear models

Performing a computer experiment can be viewed as observing a mapping between the model parameters and the corresponding model outputs predicted by the computer model. In view of this, experimental design for computer experiments can be thought of as devising a reliable procedure for finding configurations of design points in the parameter space so that their images represent the manifold parametrized by such a mapping (i.e., computer experiments). Traditional space-filling designs aim to achieve this goal by filling the parameter space with design points that are as "uniform" as possible in the parameter space. However, the resulting design points may be non-uniform in the model output space and hence fail to provide a reliable representation of the manifold, becoming highly inefficient or even misleading in case the computer experiments are non-linear. In this paper, we propose an iterative algorithm that fills in the model output manifold uniformly---rather than the parameter space uniformly---so that one could obtain a reliable understanding of the model behaviors with the minimal number of design points.

stat.CO

Sample Path Large Deviations for Lévy Processes and Random Walks with Regularly Varying Increments

Let $X$ be a Lévy process with regularly varying Lévy measure $ν$. We obtain sample-path large deviations for scaled processes $\bar X_n(t) \triangleq X(nt)/n$ and obtain a similar result for random walks. Our results yield detailed asymptotic estimates in scenarios where multiple big jumps in the increment are required to make a rare event happen; we illustrate this through detailed conditional limit theorems. In addition, we investigate connections with the classical large deviations framework. In that setting, we show that a weak large deviation principle (with logarithmic speed) holds, but a full large deviation principle does not hold.

math.PR

Lyapunov Conditions for Differentiability of Markov Chain Expectations: the Absolutely Continuous Case

We consider a family of Markov chains whose transition dynamics are affected by model parameters. Understanding the parametric dependence of (complex) performance measures of such Markov chains is often of significant interest. The derivatives of the performance measures w.r.t. the parameters play important roles, for example, in numerical optimization of the performance measures, and quantification of the uncertainties in the performance measures when there are uncertainties in the parameters from the statistical estimation procedures. In this paper, we establish conditions that guarantee the differentiability of various types of intractable performance measures---such as the stationary and random horizon discounted performance measures---of general state space Markov chains and provide probabilistic representations for the derivatives.

math.PR

Efficient Rare-Event Simulation for Multiple Jump Events in Regularly Varying Random Walks and Compound Poisson Processes

We propose a class of strongly efficient rare event simulation estimators for random walks and compound Poisson processes with a regularly varying increment/jump-size distribution in a general large deviations regime. Our estimator is based on an importance sampling strategy that hinges on the heavy-tailed sample path large deviations result recently established in Rhee, Blanchet, and Zwart (2016). The new estimators are straightforward to implement and can be used to systematically evaluate the probability of a wide range of rare events with bounded relative error. They are "universal" in the sense that a single importance sampling scheme applies to a very general class of rare events that arise in heavy-tailed systems. In particular, our estimators can deal with rare events that are caused by multiple big jumps (therefore, beyond the usual principle of a single big jump) as well as multidimensional processes such as the buffer content process of a queueing network. We illustrate the versatility of our approach with several applications that arise in the context of mathematical finance, actuarial science, and queueing theory.

math.PR

Importance sampling of heavy-tailed iterated random functions

We consider a stochastic recurrence equation of the form $Z_{n+1} = A_{n+1} Z_n+B_{n+1}$, where $\mathbb{E}[\log A_1]<0$, $\mathbb{E}[\log^+ B_1]<\infty$ and $\{(A_n,B_n)\}_{n\in\mathbb{N}}$ is an i.i.d. sequence of positive random vectors. The stationary distribution of this Markov chain can be represented as the distribution of the random variable $Z \triangleq \sum_{n=0}^\infty B_{n+1}\prod_{k=1}^nA_k$. Such random variables can be found in the analysis of probabilistic algorithms or financial mathematics, where $Z$ would be called a stochastic perpetuity. If one interprets $-\log A_n$ as the interest rate at time $n$, then $Z$ is the present value of a bond that generates $B_n$ unit of money at each time point $n$. We are interested in estimating the probability of the rare event $\{Z>x\}$, when $x$ is large; we provide a consistent simulation estimator using state-dependent importance sampling for the case, where $\log A_1$ is heavy-tailed and the so-called Cramér condition is not satisfied. Our algorithm leads to an estimator for $P(Z>x)$. We show that under natural conditions, our estimator is strongly efficient. Furthermore, we extend our method to the case, where $\{Z_n\}_{n\in\mathbb{N}}$ is defined via the recursive formula $Z_{n+1}=Ψ_{n+1}(Z_n)$ and $\{Ψ_n\}_{n\in\mathbb{N}}$ is a sequence of i.i.d. random Lipschitz functions.

math.PR