SearcharxivSearch

arXiv subjects

Hailin Sang

Publications and source records attributed to Hailin Sang.

At least 19 recordsLinked to original sources

The Double Descent Behavior in Two Layer Neural Network for Binary Classification

Recent studies observed a surprising concept on model test error called the double descent phenomenon, where the increasing model complexity decreases the test error first and then the error increases and decreases again. To observe this, we work on a two layer neural network model with a ReLU activation function designed for binary classification under supervised learning. Our aim is to observe and investigate the mathematical theory behind the double descent behavior of model test error for varying model sizes. We quantify the model size by the ratio of number of training samples to the dimension of the model. Due to the complexity of the empirical risk minimization procedure, we use the Convex Gaussian Min Max Theorem to find a suitable candidate for the global training loss.

stat.ML

Self-Normalized Moderate Deviations for Degenerate U-Statistics

In this paper, we study self-normalized moderate deviations for degenerate { $U$}-statistics of order $2$. Let $\{X_i, i \geq 1\}$ be i.i.d. random variables and consider symmetric and degenerate kernel functions in the form $h(x,y)=\sum_{l=1}^{\infty} λ_l g_l (x) g_l(y)$, where $λ_l > 0$, $E g_l(X_1)=0$, and $g_l (X_1)$ is in the domain of attraction of a normal law for all $l \geq 1$. Under the condition $\sum_{l=1}^{\infty}λ_l<\infty$ and some truncated conditions for $\{g_l(X_1): l \geq 1\}$, we show that $ \text{log} P({\frac{\sum_{1 \leq i \neq j \leq n}h(X_{i}, X_{j})} {\max_{1\le l<\infty}λ_l V^2_{n,l} }} \geq x_n^2) \sim - { \frac {x_n^2}{ 2}}$ for $x_n \to \infty$ and $x_n =o(\sqrt{n})$, where $V^2_{n,l}=\sum_{i=1}^n g_l^2(X_i)$. As application, a law of the iterated logarithm is also obtained.

math.PR

Error analysis of generative adversarial network

The generative adversarial network (GAN) is an important model developed for high-dimensional distribution learning in recent years. However, there is a pressing need for a comprehensive method to understand its error convergence rate. In this research, we focus on studying the error convergence rate of the GAN model that is based on a class of functions encompassing the discriminator and generator neural networks. These functions are VC type with bounded envelope function under our assumptions, enabling the application of the Talagrand inequality. By employing the Talagrand inequality and Borel-Cantelli lemma, we establish a tight convergence rate for the error of GAN. This method can also be applied on existing error estimations of GAN and yields improved convergence rates. In particular, the error defined with the neural network distance is a special case error in our definition.

stat.ML

Least absolute deviation estimation for AR(1) processes with roots close to unity

We establish the asymptotic theory of least absolute deviation estimators for AR(1) processes with autoregressive parameter satisfying $n(ρ_n-1)\toγ$ for some fixed $γ$ as $n\to\infty$, which is parallel to the results of ordinary least squares estimators developed by Andrews and Guggenberger (2008) in the case $γ=0$ or Chan and Wei (1987) and Phillips (1987) in the case $γ\ne 0$. Simulation experiments are conducted to confirm the theoretical results and to demonstrate the robustness of the least absolute deviation estimation.

math.ST

On the integrated mean squared error of wavelet density estimation for linear processes

Let $\{X_n: n\in \N\}$ be a linear process with density function $f(x)\in L^2(\R)$. We study wavelet density estimation of $f(x)$. Under some regular conditions on the characteristic function of innovations, we achieve, based on the number of nonzero coefficients in the linear process, the minimax optimal convergence rate of the integrated mean squared error of density estimation. Considered wavelets have compact support and are twice continuously differentiable. The number of vanishing moments of mother wavelet is proportional to the number of nonzero coefficients in the linear process and to the rate of decay of characteristic function of innovations. Theoretical results are illustrated by simulation studies with innovations following Gaussian, Cauchy and chi-squared distributions.

math.ST

Nonparametric regression with modified ReLU networks

We consider regression estimation with modified ReLU neural networks in which network weight matrices are first modified by a function $α$ before being multiplied by input vectors. We give an example of continuous, piecewise linear function $α$ for which the empirical risk minimizers over the classes of modified ReLU networks with $l_1$ and squared $l_2$ penalties attain, up to a logarithmic factor, the minimax rate of prediction of unknown $β$-smooth function.

stat.ML

Limit theorems for linear random fields with innovations in the domain of attraction of a stable law

In this paper we study the convergence in distribution and the local limit theorem for the partial sums of linear random fields with i.i.d. innovations that have infinite second moment and belong to the domain of attraction of a stable law with index $0<α\leq2$ under the condition that the innovations are centered if $1<α\leq2$ and are symmetric if $α=1$. We establish these two types of limit theorems as long as the linear random fields are well-defined, the coefficients are either absolutely summable or not absolutely summable.

math.PR

A Statistician Teaches Deep Learning

Deep learning (DL) has gained much attention and become increasingly popular in modern data science. Computer scientists led the way in developing deep learning techniques, so the ideas and perspectives can seem alien to statisticians. Nonetheless, it is important that statisticians become involved -- many of our students need this expertise for their careers. In this paper, developed as part of a program on DL held at the Statistical and Applied Mathematical Sciences Institute, we address this culture gap and provide tips on how to teach deep learning to statistics graduate students. After some background, we list ways in which DL and statistical perspectives differ, provide a recommended syllabus that evolved from teaching two iterations of a DL graduate course, offer examples of suggested homework assignments, give an annotated list of teaching resources, and discuss DL in the context of two research areas.

stat.ML

Variable bandwidth kernel regression estimation

In this paper we propose a variable bandwidth kernel regression estimator for $i.i.d.$ observations in $\mathbb{R}^2$ to improve the classical Nadaraya-Watson estimator. The bias is improved to the order of $O(h_n^4)$ under the condition that the fifth order derivative of the density function and the sixth order derivative of the regression function are bounded and continuous. We also establish the central limit theorems for the proposed ideal and true variable kernel regression estimators. The simulation study confirms our results and demonstrates the advantage of the variable bandwidth kernel method over the classical kernel method.

math.ST

Shannon entropy estimation for linear processes

In this paper, we estimate the Shannon entropy $S(f) = -\E[ \log (f(x))]$ of a one-sided linear process with probability density function $f(x)$. We employ the integral estimator $S_n(f)$, which utilizes the standard kernel density estimator $f_n(x)$ of $f(x)$. We show that $S_n (f)$ converges to $S(f)$ almost surely and in $Ł^2$ under reasonable conditions.

math.ST

A Local Limit Theorem for Linear Random Fields

In this paper, we establish a local limit theorem for linear fields of random variables constructed from independent and identically distributed innovations each with finite second moment. When the coefficients are absolutely summable we do not restrict the region of summation. However, when the coefficients are only square-summable we add the variables on unions of rectangle and we impose regularity conditions on the coefficients depending on the number of rectangles considered. Our results are new also for the dimension 1, i.e. for linear sequences of random variables. The examples include the fractionally integrated processes for which the results of a simulation study is also included.

math.PR

A Berry-Esseen bound of order $ 1/\sqrt{n} $ for martingales

Renz (Ann. Probab. 1996) has established a rate of convergence $1/\sqrt{n}$ in the central limit theorem for martingales with some restrictive conditions. In the present paper a modification of the methods, developed by Bolthausen (Ann. Probab. 1982) and Grama and Haeusler (Stochastic Process. Appl. 2000), is applied for obtaining the same convergence rate for a class of more general martingales. An application to linear processes is discussed.

math.PR

Cramér Type Moderate Deviations for Random Fields

We study the Cramér type moderate deviation for partial sums of random fields by applying the conjugate method. The results are applicable to the partial sums of linear random fields with short or long memory and to nonparametric regression with random field errors.

math.ST

On mutual information estimation for mixed-pair random variables

We study the mutual information estimation for mixed-pair random variables. One random variable is discrete and the other one is continuous. We develop a kernel method to estimate the mutual information between the two random variables. The estimates enjoy a central limit theorem under some regular conditions on the distributions. The theoretical results are demonstrated by simulation study.

math.ST

A rank-based Cramér-von-Mises-type test for two samples

We study a rank based univariate two-sample distribution-free test. The test statistic is the difference between the average of between-group rank distances and the average of within-group rank distances. This test statistic is closely related to the two-sample Cramér-von Mises criterion. They are different empirical versions of a same quantity for testing the equality of two population distributions. Although they may be different for finite samples, they share the same expected value, variance and asymptotic properties. The advantage of the new rank based test over the classical one is its ease to generalize to the multivariate case. Rather than using the empirical process approach, we provide a different easier proof, bringing in a different perspective and insight. In particular, we apply the Hájek projection and orthogonal decomposition technique in deriving the asymptotics of the proposed rank based statistic. A numerical study compares power performance of the rank formulation test with other commonly-used nonparametric tests and recommendations on those tests are provided. Lastly, we propose a multivariate extension of the test based on the spatial rank.

stat.ME

Central limit theorem for the variable bandwidth kernel density estimators

In this paper we study the ideal variable bandwidth kernel density estimator introduced by McKay (1993) and Jones, McKay and Hu (1994) and the plug-in practical version of the variable bandwidth kernel estimator with two sequences of bandwidths as in Giné and Sang (2013). Based on the bias and variance analysis of the ideal and true variable bandwidth kernel density estimators, we study the central limit theorems for each of them.

math.ST

Exact Moderate and Large Deviations for Linear Random Fields

By extending the methods in Peligrad et al. (2014a, b), we establish exact moderate and large deviation asymptotics for linear random fields with independent innovations. These results are useful for studying nonparametric regression with random field errors and strong limit theorems.

math.PR

Further Refinement of Self-normalized Cram\a'{e}r-type moderate Deviations

In this paper, we study the self-normalized Cram\a'{e}r-type moderate deviations for centered independent random variables $X_1, X_2,...$ with $0<E |X_i|^3 <\infty$. The main results refine Theorems 1.1 and 1.2 of Wang (2011), the Berry-Esseen bound (2.11) and Corollaries 2.2 and 2.3 of Jing, Shao and Wang (2003) under stronger moment conditions.

math.PR