SearcharxivSearch

arXiv subjects

Jana Gauss

Publications and source records attributed to Jana Gauss.

3 recordsLinked to original sources

Asymptotics for estimating a diverging number of parameters -- with and without sparsity

We develop a general asymptotic theory for estimating equations whose dimension diverges with the sample size. For both unpenalized and sparse penalized problems, we establish population-level conditions for existence, consistency, uniqueness, and asymptotic normality; in the penalized setting, we also establish selection consistency under a generalized version of the mutual incoherence condition. Our results cover stepwise procedures with a diverging number of steps, independent but non-identically distributed observations, and dependent data. We allow penalties that are simultaneously nonconvex, non-coordinate-separable, and group-structured and may involve heterogeneous tuning parameters. Our population-level conditions imply a weak form of restricted strong convexity, and we provide an explicit example where the commonly used stronger form fails. The results are illustrated by several applications including Group SCAD-penalized estimation in generalized linear models, distributed inference under network dependence, and penalized stepwise estimation in causal inference.

math.ST

Properties of stepwise parameter estimation in high-dimensional vine copulas

The increasing use of vine copulas in high-dimensional settings, where the number of parameters is often of the same order as the sample size, calls for asymptotic theory beyond the traditional fixed-$p$, large-$n$ framework. We establish consistency and asymptotic normality of the stepwise maximum likelihood estimator for vine copulas when the number of parameters diverges as $n \to \infty$. Our theoretical results cover both parametric and nonparametric estimation of the marginal distributions, as well as truncated vines, and are also applicable to general estimation problems, particularly other sequential procedures. Numerical experiments suggest that the derived assumptions are satisfied if the pair copulas in higher trees converge to independence copulas sufficiently fast. A simulation study substantiates these findings and identifies settings in which estimation becomes challenging. In particular, the vine structure strongly affects estimation accuracy, with D-vines being more difficult to estimate than C-vines, and estimates in Gumbel vines exhibit substantially larger biases than those in Gaussian vines.

math.ST

DCSI -- An improved measure of cluster separability based on separation and connectedness

Whether class labels in a given data set correspond to meaningful clusters is crucial for the evaluation of clustering algorithms using real-world data sets. This property can be quantified by separability measures. The central aspects of separability for density-based clustering are between-class separation and within-class connectedness, and neither classification-based complexity measures nor cluster validity indices (CVIs) adequately incorporate them. A newly developed measure (density cluster separability index, DCSI) aims to quantify these two characteristics and can also be used as a CVI. Extensive experiments on synthetic data indicate that DCSI correlates strongly with the performance of DBSCAN measured via the adjusted Rand index (ARI) but lacks robustness when it comes to multi-class data sets with overlapping classes that are ill-suited for density-based hard clustering. Detailed evaluation on frequently used real-world data sets shows that DCSI can correctly identify touching or overlapping classes that do not correspond to meaningful density-based clusters.

stat.ML