SearcharxivSearch

arXiv subjects

Yijun Zuo

Publications and source records attributed to Yijun Zuo.

At least 19 recordsLinked to original sources

On the super-efficiency and robustness of the least squares of depth-trimmed regression estimator

The least squares of depth-trimmed (LST) residuals regression, proposed and studied in Zuo and Zuo (2023), serves as a robust alternative to the classic least squares (LS) regression as well as a strong competitor to the renowned robust least trimmed squares (LTS) regression of Rousseeuw (1984). The aim of this article is three-fold. (i) to reveal the super-efficiency of the LST and demonstrate it can be as efficient as (or even more efficient than) the LS in the scenarios with errors uncorrelated and mean zero and homoscedastic with finite variance and to explain this anti-Gaussian-Markov-Theorem phenomenon; (ii) to demonstrate that the LST can outperform the LTS, the benchmark of robust regression estimator, on robustness, and the MM of Yohai (1987), the benchmark of efficient and robust estimator, on both efficiency and robustness, consequently, could serve as an alternative to both; (iii) to promote the implementation and computation of the LST regression for a broad group of statisticians in statistical practice and to demonstrate that it can be computed as fast as (or even faster than) the LTS based on a newly improved algorithm.

stat.AP

Non-asymptotic analysis of the performance of the penalized least trimmed squares in sparse models

The least trimmed squares (LTS) estimator is a renowned robust alternative to the classic least squares estimator and is popular in location, regression, machine learning, and AI literature. Many studies exist on LTS, including its robustness, computation algorithms, extension to non-linear cases, asymptotics, etc. The LTS has been applied in the penalized regression in a high-dimensional real-data sparse-model setting where dimension $p$ (in thousands) is much larger than sample size $n$ (in tens, or hundreds). In such a practical setting, the sample size $n$ often is the count of sub-population that has a special attribute (e.g. the count of patients of Alzheimer's, Parkinson's, Leukemia, or ALS, etc.) among a population with a finite fixed size N. Asymptotic analysis assuming that $n$ tends to infinity is not practically convincing and legitimate in such a scenario. A non-asymptotic or finite sample analysis will be more desirable and feasible. This article establishes some finite sample (non-asymptotic) error bounds for estimating and predicting based on LTS with high probability for the first time.

stat.ML

Computation of least squares trimmed regression--an alternative to least trimmed squares regression

The least squares of depth trimmed (LST) residuals regression, proposed in Zuo and Zuo (2023) \cite{ZZ23}, serves as a robust alternative to the classic least squares (LS) regression as well as a strong competitor to the famous least trimmed squares (LTS) regression of Rousseeuw (1984) \cite{R84}. Theoretical properties of the LST were thoroughly studied in \cite{ZZ23}. The article aims to promote the implementation and computation of the LST residuals regression for a broad group of statisticians in statistical practice and demonstrates that (i) the LST is as robust as the benchmark of robust regression, the LTS regression, and much more efficient than the latter. (ii) It can be as efficient as (or even more efficient than) the LS in the scenario with errors uncorrelated with mean zero and homoscedastic with finite variance. (iii) It can be computed as fast as (or even faster than) the LTS based on a newly proposed algorithm.

stat.ME

Weighted least squares regression with the best robustness and high computability

A novel regression method is introduced and studied. The procedure weights squared residuals based on their magnitude. Unlike the classic least squares which treats every squared residual equally important, the new procedure exponentially down-weights squared-residuals that lie far away from the cloud of all residuals and assigns a constant weight (one) to squared-residuals that lie close to the center of the squared-residual cloud. The new procedure can keep a good balance between robustness and efficiency, it possesses the highest breakdown point robustness for any regression equivariant procedure, much more robust than the classic least squares, yet much more efficient than the benchmark of robust method, the least trimmed squares (LTS) of Rousseeuw (1984). With a smooth weight function, the new procedure could be computed very fast by the first-order (first-derivative) method and the second-order (second-derivative) method. Assertions and other theoretical findings are verified in simulated and real data examples.

stat.ME

Robust penalized least squares of depth trimmed residuals regression for high-dimensional data

Challenges with data in the big-data era include (i) the dimension $p$ is often larger than the sample size $n$ (ii) outliers or contaminated points are frequently hidden and more difficult to detect. Challenge (i) renders most conventional methods inapplicable. Thus, it attracts tremendous attention from statistics, computer science, and bio-medical communities. Numerous penalized regression methods have been introduced as modern methods for analyzing high-dimensional data. Disproportionate attention has been paid to the challenge (ii) though. Penalized regression methods can do their job very well and are expected to handle the challenge (ii) simultaneously. Most of them, however, can break down by a single outlier (or single adversary contaminated point) as revealed in this article. The latter systematically examines leading penalized regression methods in the literature in terms of their robustness, provides quantitative assessment, and reveals that most of them can break down by a single outlier. Consequently, a novel robust penalized regression method based on the least sum of squares of depth trimmed residuals is proposed and studied carefully. Experiments with simulated and real data reveal that the newly proposed method can outperform some leading competitors in estimation and prediction accuracy in the cases considered.

stat.ML

Non-asymptotic robustness analysis of regression depth median

The maximum depth estimator (aka depth median) ($\bsβ^*_{RD}$) induced from regression depth (RD) of Rousseeuw and Hubert (1999) (RH99) is one of the most prevailing estimators in regression. It possesses outstanding robustness similar to the univariate location counterpart. Indeed, $\bsβ^*_{RD}$ can, asymptotically, resist up to $33\%$ contamination without breakdown, in contrast to the $0\%$ for the traditional (least squares and least absolute deviations) estimators (see Van Aelst and Rousseeuw, 2000) (VAR00)). The results from VAR00 are pioneering, yet they are limited to regression-symmetric populations (with a strictly positive density) and the $ε$-contamination and maximum-bias model. With a fixed finite-sample size practice, the most prevailing measure of robustness for estimators is the finite-sample breakdown point (FSBP) (Donoho and Huber (1983)). Despite many attempts made in the literature, only sporadic partial results on FSBP for $\bsβ^*_{RD}$ were obtained whereas an exact FSBP for $\bsβ^*_{RD}$ remained open in the last twenty-plus years. Furthermore, is the asymptotic breakdown value $1/3$ (the limit of an increasing sequence of finite-sample breakdown values) relevant in the finite-sample practice? (Or what is the difference between the finite-sample and the limit breakdown values?). Such discussions are yet to be given in the literature. This article addresses the above issues, revealing an intrinsic connection between the regression depth of $\bsβ^*_{RD}$ and the newly obtained exact FSBP. It justifies the employment of $\bsβ^*_{RD}$ as a robust alternative to the traditional estimators and demonstrates the necessity and the merit of using the FSBP in finite-sample real practice.

math.ST

Least sum of squares of trimmed residuals regression

In the famous least sum of trimmed squares (LTS) of residuals estimator (Rousseeuw (1984)), residuals are first squared and then trimmed. In this article, we first trim residuals - using a depth trimming scheme - and then square the rest of residuals. The estimator that can minimize the sum of squares of the trimmed residuals, is called an LST estimator. It turns out that LST is a robust alternative to the classic least sum of squares (LS) estimator. Indeed, it has a very high finite sample breakdown point, and can resist, asymptotically, up to $50\%$ contamination without breakdown - in sharp contrast to the $0\%$ of the LS estimator. The population version of LST is Fisher consistent, and the sample version is strong and root-$n$ consistent and asymptotically normal. Approximate algorithms for computing LST are proposed and tested in synthetic and real data examples. These experiments indicate that one of the algorithms can compute the LST estimator very fast and with relatively smaller variances than the famous LTS estimator. All the evidence suggests that LST deserves to be a robust alternative to the LS estimator and is feasible in practice for high dimensional data sets (with possible contamination and outliers).

stat.ME

New algorithms for computing the least trimmed squares estimator

Instead of minimizing the sum of all $n$ squared residuals as the classical least squares (LS) does, Rousseeuw (1984) proposed to minimize the sum of $h$ ($n/2 \leq h < n$) smallest squared residuals, the resulting estimator is called least trimmed squares (LTS). The idea of the LTS is simple but its computation is challenging since no LS-type analytical computation formula exists anymore. Attempts had been made since its presence, the feasible solution algorithm (Hawkins (1994)), fastlts.f (Rousseeuw and Van Driessen (1999)), and FAST-LTS (Rousseeuw and Van Driessen (2006)), among others, are promising approximate algorithms. The latter two have been incorporated into R function ltsReg by Valentin Todorov. These algorithms utilize combinatorial- or subsampling- approaches. With the great software accessibility and fast speed, the LTS, enjoying many desired properties, has become one of the most popular robust regression estimators across multiple disciplines. This article proposes analytic approaches -- employing first-order derivative (gradient) and second-order derivative (Hessian matrix) of the objective function. Our approximate algorithms for the LTS are vetted in synthetic and real data examples. Compared with ltsReg -- the benchmark in robust regression and well-known for its speed, our algorithms are comparable (and sometimes even favorable) with respect to both speed and accuracy criteria. Other major contributions include (i) originating the uniqueness and the strong and Fisher consistency at empirical and population settings respectively; (ii) deriving the influence function in a general setting; (iii) re-establishing the asymptotic normality (consequently root-n consistency) of the estimator with a neat and general approach.

stat.CO

Large sample behavior of the least trimmed squares estimator

The least trimmed squares (LTS) estimator is popular in location, regression, machine learning, and AI literature. Despite the empirical version of least trimmed squares (LTS) being repeatedly studied in the literature, the population version of the LTS has never been introduced and studied. The lack of the population version hinders the study of the large sample properties of the LTS utilizing the empirical process theory. Novel properties of the objective function in both empirical and population settings of the LTS and other properties are established for the first time in this article. The primary properties of the objective function facilitate the establishment of other original results, including the influence function and Fisher consistency. The strong consistency is established with the help of a generalized Glivenko-Cantelli Theorem over a class of functions for the first time. Differentiability and stochastic equicontinuity promote the establishment of asymptotic normality with a concise and novel approach.

math.ST

Asymptotic normality of the least sum of squares of trimmed residuals estimator

To enhance the robustness of the classic least sum of squares (LS) of the residuals estimator, Zuo (2022) introduced the least sum of squares of trimmed (LST) residuals estimator. The LST enjoys many desired properties and serves well as a robust alternative to the LS. Its asymptotic properties, including strong and root-n consistency, have been established whereas the asymptotic normality is left unaddressed. This article solves this remained problem.

math.ST

Non-asymptotic analysis and inference for an outlyingness induced winsorized mean

Robust estimation of a mean vector, a topic regarded as obsolete in the traditional robust statistics community, has recently surged in machine learning literature in the last decade. The latest focus is on the sub-Gaussian performance and computability of the estimators in a non-asymptotic setting. Numerous traditional robust estimators are computationally intractable, which partly contributes to the renewal of the interest in the robust mean estimation. Robust centrality estimators, however, include the trimmed mean and the sample median. The latter has the best robustness but suffers a low-efficiency drawback. Trimmed mean and median of means, %as robust alternatives to the sample mean, and achieving sub-Gaussian performance have been proposed and studied in the literature. This article investigates the robustness of leading sub-Gaussian estimators of mean and reveals that none of them can resist greater than $25\%$ contamination in data and consequently introduces an outlyingness induced winsorized mean which has the best possible robustness (can resist up to $50\%$ contamination without breakdown) meanwhile achieving high efficiency. Furthermore, it has a sub-Gaussian performance for uncontaminated samples and a bounded estimation error for contaminated samples at a given confidence level in a finite sample setting. It can be computed in linear time.

stat.ML

Computation of projection regression depth and its induced median

Notions of depth in regression have been introduced and studied in the literature. The most famous example is Regression Depth (RD), which is a direct extension of location depth to regression. The projection regression depth (PRD) is the extension of another prevailing location depth, the projection depth, to regression. The computation issues of the RD have been discussed in the literature. The computation issues of the PRD have never been dealt with before. The computation issues of the PRD and its induced median (maximum depth estimator) in a regression setting are addressed now. For a given $\bsβ\in\R^p$ exact algorithms for the PRD with cost $O(n^2\log n)$ ($p=2$) and $O(N(n, p)(p^{3}+n\log n+np^{1.5}+npN_{Iter}))$ ($p>2$) and approximate algorithms for the PRD and its induced median with cost respectively $O(N_{\mb{v}}np)$ and $O(Rp N_{\bsβ}(p^2+nN_{\mb{v}}N_{Iter}))$ are proposed. Here $N(n, p)$ is a number defined based on the total number of $(p-1)$ dimensional hyperplanes formed by points induced from sample points and the $\bsβ$; $N_{\mb{v}}$ is the total number of unit directions $\mb{v}$ utilized; $N_{\bsβ}$ is the total number of candidate regression parameters $\bsβ$ employed; $N_{Iter}$ is the total number of iterations carried out in an optimization algorithm; $R$ is the total number of replications. Furthermore, as the second major contribution, three PRD induced estimators, which can be computed up to 30 times faster than that of the PRD induced median while maintaining a similar level of accuracy are introduced. Examples and simulation studies reveal that the depth median induced from the PRD is favorable in terms of robustness and efficiency, compared to the maximum depth estimator induced from the RD, which is the current leading regression median.

stat.CO

Large sample properties of the regression depth induced median

Notions of depth in regression have been introduced and studied in the literature. Regression depth (RD) of Rousseeuw and Hubert (1999), the most famous one, is a direct extension of Tukey location depth (Tukey (1975)) to regression. Like its location counterpart, the most remarkable advantage of the notion of depth in regression is to directly introduce the maximum (or deepest) regression depth estimator (aka depth induced median) for regression parameters in a multi-dimensional setting. Classical questions for the regression depth induced median include (i) is it a consistent estimator (or rather under what sufficient conditions, it is consistent)? and (ii) is there any limiting distribution? Bai and He (1999) (BH99) pioneered an attempt to answer these questions. Under some stringent conditions on (i) the design points, (ii) the conditional distributions of $y$ given $\bs{x}_i$, and (iii) the error distributions, BH99 proved the strong consistency of the depth induced median. Under another set of conditions, BH99 showed the existence of the limiting distribution of the estimator. This article establishes the strong consistency of the depth induced median without any of the stringent conditions in BH99, and proves the existence of the limiting distribution of the estimator by sufficient conditions and an approach different from BH99.

math.ST

Exact computation of projection regression depth and fast computation of its induced median and other estimators

Zuo (2019) (Z19) addressed the computation of the projection regression depth (PRD) and its induced median (the maximum depth estimator). Z19 achieved the exact computation of PRD via a modified version of regular univariate sample median, which resulted in the loss of invariance of PRD and the equivariance of depth induced median. This article achieves the exact computation without scarifying the invariance of PRD and the equivariance of the regression median. Z19 also addressed the approximate computation of PRD induced median, the naive algorithm in Z19 is very slow. This article modifies the approximation in Z19 and adopts Rcpp package and consequently obtains a much (could be $100$ times) faster algorithm with an even better level of accuracy meanwhile. Furthermore, as the third major contribution, this article introduces three new depth induced estimators which can run $300$ times faster than that of Z19 meanwhile maintaining the same (or a bit better) level of accuracy. Real as well as simulated data examples are presented to illustrate the difference between the algorithms of Z19 and the ones proposed in this article. Findings support the statements above and manifest the major contributions of the article.

stat.ME

Depth induced regression medians and uniqueness

Notion of median in one dimension is a foundational element in nonparametric statistics. It has been extended to multi-dimensional cases both in location and in regression via notions of data depth. Regression depth (RD) and projection regression depth (PRD) represent the two most promising notions in regression. Carrizosa depth $D_C$ is another depth notion in regression.Depth induced regression medians (maximum depth estimators) serve as robust alternatives to the classical least squares estimator. The uniqueness of regression medians is indispensable in the discussion of their properties and the asymptotics (consistency and limiting distribution) of sample regression medians. Are the regression medians induced from RD, PRD, and $D_C$ unique? Answering this question is the main goal of this article. It is found that only the regression median induced from PRD possesses the desired uniqueness property. The conventional remedy measure for non-uniqueness, taking average of all medians, might yield an estimator that no longer possesses the maximum depth in both RD and $D_C$ cases. These and other findings indicate that the PRD and its induced median are highly favorable among their leading competitors.

math.ST

On general notions of depth for regression

Depth notions in location have fascinated tremendous attention in the literature. In fact data depth and its applications remain one of the most active research topics in statistics in the last two decades. Most favored notions of depth in location include Tukey (1975) halfspace depth (HD), Liu (1990) simplicial depth, and projection depth (Stahel (1981) and Donoho (1982), Liu (1992), Zuo and Serfling (2000) (ZS00) and Zuo (2003)), among others. Depth notions in regression have also been proposed, sporadically nevertheless. Regression depth (RD) of Rousseeuw and Hubert (1999) (RH99) is the most famous one which is a direct extension of Tukey HD to regression. Others include Carrizosa (1996) and the ones induced from Marrona and Yohai (1993) (MY93) proposed in this article. Is there any relationship between Carrizosa depth and RD of RH99? Do these depth notions possess desirable properties? What are the desirable properties? Can existing notions really serve as depth functions in regression? These questions remain open. Revealing the equivalence between Carrizosa depth and RD of RH99; expanding location depth evaluating criteria in ZS00 for regression depth notions; examining the existing regression notions with respect to the gauges; and proposing the regression counterpart of the eminent projection depth in location are the four major objectives of the article.

stat.ME

Robustness of deepest projection regression functional

Depth notions in regression have been systematically proposed and examined in Zuo (2018). One of the prominent advantages of notion of depth is that it can be directly utilized to introduce median-type deepest estimating functionals (or estimators in empirical distribution case) for location or regression parameters in a multi-dimensional setting. Regression depth shares the advantage. Depth induced deepest estimating functionals are expected to inherit desirable and inherent robustness properties ( e.g. bounded maximum bias and influence function and high breakdown point) as their univariate location counterpart does. Investigating and verifying the robustness of the deepest projection estimating functional (in terms of maximum bias, asymptotic and finite sample breakdown point, and influence function) is the major goal of this article. It turns out that the deepest projection estimating functional possesses a bounded influence function and the best possible asymptotic breakdown point as well as the best finite sample breakdown point with robust choice of its univariate regression and scale component.

math.ST

The limit of finite sample breakdown point of Tukey's halfspace median for general data

Under special conditions on data set and underlying distribution, the limit of finite sample breakdown point of Tukey's halfspace median ($\frac{1} {3}$) has been obtained in literature. In this paper, we establish the result under \emph{weaker assumption} imposed on underlying distribution (halfspace symmetry) and on data set (not necessary in general position). The representation of Tukey's sample depth regions for data set \emph{not necessary in general position} is also obtained, as a by-product of our derivation.

math.ST