SearcharxivSearch

arXiv subjects

Stanislav Nagy

Publications and source records attributed to Stanislav Nagy.

At least 19 recordsLinked to original sources

Resampling simplicial depth

The simplicial depth (SD) is a commonly used indicator of the centrality of points $x\in\mathbb{R}^d$ with respect to distributions $P$ on $\mathbb{R}^d$. Asymptotic theory for the sample SD is based on its representation as a $U$-statistic, which can be either non-degenerate or degenerate. For $d=2$, we prove under mild conditions that this $U$-statistic is degenerate with rate $n$ if and only if $x$ is a center of symmetry of $P$. Otherwise, the asymptotic distribution of SD is non-degenerate with rate $\sqrt{n}$. Because the location of the center of symmetry of $P$ is usually unknown, these two modes of behavior complicate the estimation of the sample distribution of SD at $x$. We propose a two-step adaptive subsampling procedure for estimating that distribution. First, an estimator $\widehat \gamma$ of a parameter $\gamma\in\{1/2, 1\}$ characterizing the correct rate of convergence $n^\gamma$ of SD is constructed based on subsampling. Our estimator uses a bias correction suitable for $U$-statistics. Second, $\widehat \gamma$ is employed for approximating the distribution of the sample SD. We prove the consistency of this subsampling approach and illustrate its usefulness (i) in the construction of confidence intervals for SD, and (ii) in an SD-based supervised classification task

stat.ME

Projection depth for functional data: Practical issues, computation and applications

Statistical analysis of functional data is challenging due to their complex patterns, for which functional depth provides an effective means of reflecting their ordering structure. In this work, we investigate practical aspects of the recently proposed regularized projection depth (RPD), which induces a meaningful ordering of functional data while appropriately accommodating their complex shape features. Specifically, we examine the impact and choice of its tuning parameter, which regulates the degree of effective dimension reduction applied to the data, and propose a random projection-based approach for its efficient computation, supported by theoretical justification. Through comprehensive numerical studies, we explore a wide range of statistical applications of the RPD and demonstrate its particular usefulness in uncovering shape features in functional data analysis. This ability allows us to propose a novel RPD-based outlier detection method, and to demonstrate that the RPD outperforms competing depth-based approaches in tasks such as functional classification and two-sample hypothesis testing.

stat.ME

Estimating axial symmetry using random projections

This paper studies the problem of identifying directions of axial symmetry in multivariate distributions. Theoretical results are derived on how the measure or cardinality of the set of symmetry directions relates to spherical symmetry. The problem is framed using random projections, leading to a proof that in \(\RR^2\), agreement on two random projections is enough to identify the true axes of symmetry. A corresponding result for higher dimensions is conjectured. An estimator for the symmetry directions is proposed and proved to be consistent in the plane.

math.ST

Projection depth for functional data: Theoretical properties

We introduce a novel projection depth for data lying in a general Hilbert space, called the regularized projection depth, with a focus on functional data. By regularizing projection directions, the proposed depth does not suffer from the degeneracy issue that may arise when the classical projection depth is naively defined on an infinite-dimensional space. Compared to existing functional depth notions, the regularized projection depth has several advantages: (i) it requires no moment assumptions on the underlying distribution, (ii) it satisfies many desirable depth properties including invariance, monotonicity, and vanishing at infinity, (iii) its sample version uniformly converges under mild conditions, and (iv) it generates a highly robust median. Furthermore, the proposed depth is statistically useful as it (v) does not produce ties in the induced ranks and (vi) effectively detects shape outlying functions. This paper focuses mainly on the theoretical properties of the regularized projection depth.

stat.ME

Location and scatter halfspace median under {\alpha}-symmetric distributions

In a landmark result, Chen et al. (2018) showed that multivariate medians induced by halfspace depth attain the minimax optimal convergence rate under Huber contamination and elliptical symmetry, for both location and scatter estimation. We extend some of these findings to the broader family of {\alpha}-symmetric distributions, which includes both elliptically symmetric and multivariate heavy-tailed distributions. For location estimation, we establish an upper bound on the estimation error of the location halfspace median under the Huber contamination model. An analogous result for the standard scatter halfspace median matrix is feasible only under the assumption of elliptical symmetry, as ellipticity is deeply embedded in the definition of scatter halfspace depth. To address this limitation, we propose a modified scatter halfspace depth that better accommodates {\alpha}-symmetric distributions, and derive an upper bound for the corresponding {\alpha}-scatter median matrix. Additionally, we identify several key properties of scatter halfspace depth for {\alpha}-symmetric distributions.

math.ST

Which depth to use to construct functional boxplots?

This paper answers the question of which functional depth to use to construct a boxplot for functional data. It shows that integrated depths, e.g., the popular modified band depth, do not result in well-defined boxplots. Instead, we argue that infimal depths are the only functional depths that provide a valid construction of a functional boxplot. We also show that the properties of the boxplot are completely determined by properties of the one-dimensional depth function used in defining the infimal depth for functional data. Our claims are supported by (i) a motivating example, (ii) theoretical results concerning the properties of the boxplot, and (iii) a simulation study.

stat.ME

Theoretical properties of angular halfspace depth

The angular halfspace depth (ahD) is a natural modification of the celebrated halfspace (or Tukey) depth to the setup of directional data. It allows us to define elements of nonparametric inference, such as the median, the inter-quantile regions, or the rank statistics, for datasets supported in the unit sphere. Despite being introduced in 1987, ahD has never received ample recognition in the literature, mainly due to the lack of efficient algorithms for its computation. With the recent progress on the computational front, ahD however exhibits the potential for developing viable nonparametric statistics techniques for directional datasets. In this paper, we thoroughly treat the theoretical properties of ahD. We show that similarly to the classical halfspace depth for multivariate data, also ahD satisfies many desirable properties of a statistical depth function. Further, we derive uniform continuity/consistency results for the associated set of directional medians, and the central regions of ahD, the latter representing a depth-based analogue of the quantiles for directional data.

math.ST

Robust Functional Regression with Discretely Sampled Predictors

The functional linear model is an important extension of the classical regression model allowing for scalar responses to be modeled as functions of stochastic processes. Yet, despite the usefulness and popularity of the functional linear model in recent years, most treatments, theoretical and practical alike, suffer either from (i) lack of resistance towards the many types of anomalies one may encounter with functional data or (ii) biases resulting from the use of discretely sampled functional data instead of completely observed data. To address these deficiencies, this paper introduces and studies the first class of robust functional regression estimators for partially observed functional data. The proposed broad class of estimators is based on thin-plate splines with a novel computationally efficient quadratic penalty, is easily implementable and enjoys good theoretical properties under weak assumptions. We show that, in the incomplete data setting, both the sample size and discretization error of the processes determine the asymptotic rate of convergence of functional regression estimators and the latter cannot be ignored. These theoretical properties remain valid even with multi-dimensional random fields acting as predictors and random smoothing parameters. The effectiveness of the proposed class of estimators in practice is demonstrated by means of a simulation study and a real-data example.

stat.ME

Another look at halfspace depth: Flag halfspaces with applications

The halfspace depth is a well studied tool of nonparametric statistics in multivariate spaces, naturally inducing a multivariate generalisation of quantiles. The halfspace depth of a point with respect to a measure is defined as the infimum mass of closed halfspaces that contain the given point. In general, a closed halfspace that attains that infimum does not have to exist. We introduce a flag halfspace - an intermediary between a closed halfspace and its interior. We demonstrate that the halfspace depth can be equivalently formulated also in terms of flag halfspaces, and that there always exists a flag halfspace whose boundary passes through any given point $x$, and has mass exactly equal to the halfspace depth of $x$. Flag halfspaces allow us to derive theoretical results regarding the halfspace depth without the need to differentiate absolutely continuous measures from measures containing atoms, as was frequently done previously. The notion of flag halfspaces is used to state results on the dimensionality of the halfspace median set for random samples. We prove that under mild conditions, the dimension of the sample halfspace median set of $d$-variate data cannot be $d-1$, and that for $d=2$ the sample halfspace median set must be either a two-dimensional convex polygon, or a data point. The latter result guarantees that the computational algorithm for the sample halfspace median form the R package TukeyRegion is exact also in the case when the median set is less-than-full-dimensional in dimension $d=2$.

stat.ME

Exact and approximate computation of the scatter halfspace depth

The scatter halfspace depth (sHD) is an extension of the location halfspace (also called Tukey) depth that is applicable in the nonparametric analysis of scatter. Using sHD, it is possible to define minimax optimal robust scatter estimators for multivariate data. The problem of exact computation of sHD for data of dimension $d \geq 2$ has, however, not been addressed in the literature. We develop an exact algorithm for the computation of sHD in any dimension $d$ and implement it efficiently using C++ for $d \leq 5$, and in R for any dimension $d \geq 1$. Since the exact computation of sHD is slow especially for higher dimensions, we also propose two fast approximate algorithms. All our programs are freely available in the R package scatterdepth.

stat.CO

On exact computation of Tukey depth central regions

The Tukey (or halfspace) depth extends nonparametric methods toward multivariate data. The multivariate analogues of the quantiles are the central regions of the Tukey depth, defined as sets of points in the $d$-dimensional space whose Tukey depth exceeds given thresholds $k$. We address the problem of fast and exact computation of those central regions. First, we analyse an efficient Algorithm A from Liu et al. (2019), and prove that it yields exact results in dimension $d=2$, or for a low threshold $k$ in arbitrary dimension. We provide examples where Algorithm A fails to recover the exact Tukey depth region for $d>2$, and propose a modification that is guaranteed to be exact. We express the problem of computing the exact central region in its dual formulation, and use that viewpoint to demonstrate that further substantial improvements to our algorithm are unlikely. An efficient C++ implementation of our exact algorithm is freely available in the R package TukeyRegion.

stat.CO

Partial reconstruction of measures from halfspace depth

The halfspace depth of a $d$-dimensional point $x$ with respect to a finite (or probability) Borel measure $\mu$ in $\mathbb{R}^d$ is defined as the infimum of the $\mu$-masses of all closed halfspaces containing $x$. A natural question is whether the halfspace depth, as a function of $x \in \mathbb{R}^d$, determines the measure $\mu$ completely. In general, it turns out that this is not the case, and it is possible for two different measures to have the same halfspace depth function everywhere in $\mathbb{R}^d$. In this paper we show that despite this negative result, one can still obtain a substantial amount of information on the support and the location of the mass of $\mu$ from its halfspace depth. We illustrate our partial reconstruction procedure in an example of a non-trivial bivariate probability distribution whose atomic part is determined successfully from its halfspace depth.

math.ST

Integrated shape-sensitive functional metrics

This paper develops a new integrated ball (pseudo)metric which provides an intermediary between a chosen starting (pseudo)metric d and the L_p distance in general function spaces. Selecting d as the Hausdorff or Fr\'echet distances, we introduce integrated shape-sensitive versions of these supremum-based metrics. The new metrics allow for finer analyses in functional settings, not attainable applying the non-integrated versions directly. Moreover, convergent discrete approximations make computations feasible in practice.

math.MG

Halfspace depth for general measures: The ray basis theorem and its consequences

The halfspace depth is a prominent tool of nonparametric multivariate analysis. The upper level sets of the depth, termed the trimmed regions of a measure, serve as a natural generalization of the quantiles and inter-quantile regions to higher-dimensional spaces. The smallest non-empty trimmed region, coined the halfspace median of a measure, generalizes the median. We focus on the (inverse) ray basis theorem for the halfspace depth, a crucial theoretical result that characterizes the halfspace median by a covering property. First, a novel elementary proof of that statement is provided, under minimal assumptions on the underlying measure. The proof applies not only to the median, but also to other trimmed regions. Motivated by the technical development of the amended ray basis theorem, we specify connections between the trimmed regions, floating bodies, and additional equi-affine convex sets related to the depth. As a consequence, minimal conditions for the strict monotonicity of the depth are obtained. Applications to the computation of the depth and robust estimation are outlined.

math.ST

Statistical Depth Meets Machine Learning: Kernel Mean Embeddings and Depth in Functional Data Analysis

Statistical depth is the act of gauging how representative a point is compared to a reference probability measure. The depth allows introducing rankings and orderings to data living in multivariate, or function spaces. Though widely applied and with much experimental success, little theoretical progress has been made in analysing functional depths. This article highlights how the common $h$-depth and related statistical depths for functional data can be viewed as a kernel mean embedding, a technique used widely in statistical machine learning. This connection facilitates answers to open questions regarding statistical properties of functional depths, as well as it provides a link between the depth and empirical characteristic function based procedures for functional data.

math.ST

Approximate computation of projection depths

Data depth is a concept in multivariate statistics that measures the centrality of a point in a given data cloud in $\IR^d$. If the depth of a point can be represented as the minimum of the depths with respect to all one-dimensional projections of the data, then the depth satisfies the so-called projection property. Such depths form an important class that includes many of the depths that have been proposed in literature. For depths that satisfy the projection property an approximate algorithm can easily be constructed since taking the minimum of the depths with respect to only a finite number of one-dimensional projections yields an upper bound for the depth with respect to the multivariate data. Such an algorithm is particularly useful if no exact algorithm exists or if the exact algorithm has a high computational complexity, as is the case with the halfspace depth or the projection depth. To compute these depths in high dimensions, the use of an approximate algorithm with better complexity is surely preferable. Instead of focusing on a single method we provide a comprehensive and fair comparison of several methods, both already described in the literature and original.

stat.CO

Uniform convergence rates for the approximated halfspace and projection depth

The computational complexity of some depths that satisfy the projection property, such as the halfspace depth or the projection depth, is known to be high, especially for data of higher dimensionality. In such scenarios, the exact depth is frequently approximated using a randomized approach: The data are projected into a finite number of directions uniformly distributed on the unit sphere, and the minimal depth of these univariate projections is used to approximate the true depth. We provide a theoretical background for this approximation procedure. Several uniform consistency results are established, and the corresponding uniform convergence rates are provided. For elliptically symmetric distributions and the halfspace depth it is shown that the obtained uniform convergence rates are sharp. In particular, guidelines for the choice of the number of random projections in order to achieve a given precision of the depths are stated.

math.ST

Scalar-on-function local linear regression and beyond

Regressing a scalar response on a random function is nowadays a common situation. In the nonparametric setting, this paper paves the way for making the local linear regression based on a projection approach a prominent method for solving this regression problem. Our asymptotic results demonstrate that the functional local linear regression outperforms its functional local constant counterpart. Beyond the estimation of the regression operator itself, the local linear regression is also a useful tool for predicting the functional derivative of the regression operator, a promising mathematical object on its own. The local linear estimator of the functional derivative is shown to be consistent. On simulated datasets we illustrate good finite sample properties of both proposed methods. On a real data example of a single-functional index model we indicate how the functional derivative of the regression operator provides an original and fast, widely applicable estimating method.

stat.ME