Searcharxiv⌕ Search

arXiv subjects

Sungkyu Jung

Publications and source records attributed to Sungkyu Jung.

At least 19 recordsLinked to original sources

An association measure for mixed-type variables

Quantifying the association between a real-valued variable and a categorical variable is a fundamental task in data analysis. Existing methods often rely on parametric assumptions or arbitrary integer encoding, which may lead to unstable results. We propose a label-invariant population measure of association, $ξ'$, specifically designed for the mixed real-valued-categorical setting. The proposed measure is normalized between 0 and 1; it equals 0 if and only if the variables are independent and 1 if and only if the categorical variable is a measurable function of the real-valued one. We also introduce a corresponding sample estimator, $ξ_n'$, computable in $O(n \log n)$ time. These measures are invariant to permutations of category labels and strictly monotone transformations of the real-valued variable. We establish the strong consistency and asymptotic normality of the estimator $ξ_n'$, enabling a computationally efficient, permutation-free Wald test for independence, and an asymptotic confidence interval for the population measure $ξ'$. Extensive simulations and an application to The Cancer Genome Atlas (TCGA) data demonstrate that the proposed method provides coding stability, competitive power, and substantial computational advantages in nominal mixed-type settings.

stat.ME↗

Joint estimation of high-dimensional spiked covariance matrices via a partially shared subspace

Statistical analysis of high-dimensional data is often hampered by limited sample sizes, yet auxiliary datasets from related sources are often readily available. When two such datasets share part of their covariance structure, but not all of it, exploiting the shared part can substantially improve estimation. We propose a spiked covariance model that explicitly captures this partial sharing: two datasets share a subspace of unknown rank and arbitrary position in the spectrum, while each retains its own distinct spiked directions. The model treats the two datasets symmetrically and strictly generalizes existing models for shared covariance structure. We develop a complete estimation procedure that includes joint estimation of the shared subspace and its rank, a closed-form pooling weight for combining the two datasets, and asymptotic guarantees derived from random matrix theory in the proportional-growth regime. The framework also resolves a gap in contrastive dimension reduction by providing a principled estimator for high-dimensional settings. We illustrate the methodology on portfolio construction during the early COVID-19 pandemic and on contrastive analysis of brain tumor gene expression.

stat.ME↗

General M-estimators of location on Riemannian manifolds: existence and uniqueness

We study general M-estimators of location on Riemannian manifolds, extending classical notions such as the Frechet mean by replacing the squared loss with a broad class of loss functions. Under minimal regularity conditions on the loss function and the underlying probability distribution, we establish theoretical guarantees for the existence and uniqueness of such estimators. In particular, we provide sufficient conditions under which the population and sample M-estimators exist and are uniquely defined. Our results offer a general framework for robust location estimation in non-Euclidean geometric spaces and unify prior uniqueness results under a broad class of convex losses.

math.ST↗

Enhancing Differentially Private Mechanisms via Empirical Bayes

Differential privacy (DP) has become the gold standard for ensuring the privacy protection of machine learning and statistical algorithms in recent decades. A plethora of algorithms and methods have been developed to enhance the utility of DP algorithms while maintaining the same level of DP. However, these are often overly complex or computationally ineffective. We propose a novel approach focusing on denoising the output of the simple additive Gaussian mechanism by adopting the idea of \textit{empirical Bayes estimation}. We highlight that the empirical Bayes approach can reduce the mean-squared error solely by taking the output of the Gaussian mechanism as input. Our numerical studies show that this simple yet powerful approach can be applied to improve upon various statistical problems, including histogram release, principal component analysis, and linear regression, often outperforming existing private algorithms.

cs.LG↗

Robust and Differentially Private Principal Component Analysis

Recent advances have sparked significant interest in the development of privacy-preserving Principal Component Analysis (PCA). However, many existing approaches rely on restrictive assumptions, such as assuming sub-Gaussian data or being vulnerable to data contamination. Additionally, some methods are computationally expensive or depend on unknown model parameters that must be estimated, limiting their accessibility for data analysts seeking privacy-preserving PCA. In this paper, we propose a differentially private PCA method applicable to heavy-tailed and potentially contaminated data. Our approach leverages the property that the covariance matrix of properly rescaled data preserves eigenvectors and their order under elliptical distributions, which include Gaussian and heavy-tailed distributions. By applying a bounded transformation, we enable straightforward computation of principal components in a differentially private manner. Additionally, boundedness guarantees robustness against data contamination. We conduct both theoretical analysis and empirical evaluations of the proposed method, focusing on its ability to recover the subspace spanned by the leading principal components. Extensive numerical experiments demonstrate that our method consistently outperforms existing approaches in terms of statistical utility, particularly in non-Gaussian or contaminated data settings.

stat.ME↗

Predicting Current Outcomes From Historical Survey Data With Weighted Conformal Prediction

In large-scale complex surveys such as the National Health and Nutrition Examination Survey (NHANES), some outcomes are measured only in selected years, leaving incomplete records across survey waves. We develop a weighted conformal prediction framework that enables valid population-level prediction of unobserved outcomes using information from earlier surveys. The method accommodates covariate shift, where both continuous and categorical covariate distributions evolve over time while survey design affects representativeness. It integrates subgroup-specific density ratio and subgroup-proportion estimation to approximate likelihood ratios between the historical and target covariate distributions, and we establish coverage guarantees for the resulting prediction sets. Simulation studies and an application predicting low-density lipoprotein cholesterol (LDL-C) for the current U.S. population show that the proposed approach achieves coverage close to the nominal level and improved efficiency over existing methods, particularly when covariate distributions are complex or unknown.

stat.ME↗

iLBA: An R package for confidentially disseminating aggregated frequency tables

Statistical agencies frequently release frequency tables derived from microdata, but small frequency cells may lead to disclosure risks. We present \texttt{iLBA}, an open-source \textsf{R} package for confidential dissemination of aggregated frequency tables. The package implements the Information-Loss-Bounded Aggregation (iLBA) algorithm, which combines Small Cell Adjustment (SCA) at the finest level table with an aggregation procedure that introduces controlled ambiguity while bounding information loss. The software enables users to construct masked finest level tables, generate confidential aggregated tables for selected variables, and obtain masked frequencies for single-cell queries. By providing an accessible implementation of the iLBA method, the package facilitates reproducible and efficient disclosure control for tabular data derived from microdata.

stat.CO↗

Generalized Fréchet means with random minimizing domains and its strong consistency

This paper introduces a novel extension of Fréchet means, called \textit{generalized Fréchet means} as a comprehensive framework for characterizing features in probability distributions in general topological spaces. The generalized Fréchet means are defined as minimizers of a suitably defined cost function. The framework encompasses various extensions of Fréchet means in the literature. The most distinctive difference of the new framework from the previous works is that we allow the domain of minimization of the empirical means be random and different from that of the population means. This expands the applicability of the Fréchet mean framework to diverse statistical scenarios, including dimension reduction for manifold-valued data.

math.ST↗

Huber means on Riemannian manifolds

This article introduces Huber means on Riemannian manifolds, providing a robust alternative to the Frechet mean by integrating elements of both square and absolute loss functions. The Huber means are designed to be highly resistant to outliers while maintaining efficiency, making it a valuable generalization of Huber's M-estimator for manifold-valued data. We comprehensively investigate the statistical and computational aspects of Huber means, demonstrating their utility in manifold-valued data analysis. Specifically, we establish nearly minimal conditions for ensuring the existence and uniqueness of the Huber mean and discuss regularity conditions for unbiasedness. The Huber means are consistent and enjoy the central limit theorem. Additionally, we propose a novel moment-based estimator for the limiting covariance matrix, which is used to construct a robust one-sample location test procedure and an approximate confidence region for location parameters. The Huber mean is shown to be highly robust and efficient in the presence of outliers or under heavy-tailed distributions. Specifically, it achieves a breakdown point of at least 0.5, the highest among all isometric equivariant estimators, and is more efficient than the Frechet mean under heavy-tailed distributions. Numerical examples on spheres and the space of symmetric positive-definite matrices further illustrate the efficiency and reliability of the proposed Huber means on Riemannian manifolds.

math.ST↗

Optimal Test-Data Piling in HDLSS Classification with Covariance Heterogeneity

This work addresses a longstanding question in high-dimensional linear classification: Is perfect classification achievable in heterogeneous covariance structures? We focus on the phenomenon of data piling, where projected data points collapse onto discrete values. We provide a comprehensive characterization of two distinct types of data piling. The first type of data piling refers to the phenomenon where projecting the training data onto a certain direction yields exactly two distinct values-one for each class. This occurs universally when the data dimension $p$ exceeds the sample size $n$. The second type concerns independent test data and arises asymptotically as $p \to \infty$ with fixed $n$. While previous work established the existence of such double data piling under homogeneously spiked covariance structures using negatively ridged classifiers, our analysis extends to the more general and realistic case of heterogeneous covariance. We identify an optimal direction among all piling directions that maximizes the separation between test data piles, which is called the Second Maximal Data Piling direction. An algorithm based on data splitting is proposed to compute this direction using only training data. Our analysis reveals a key insight: the main obstacle to discovering this direction is the imbalance of the tail eigenvalues, rather than differences in spike count, spike magnitude, or the alignment of leading eigenspaces. Extensive simulations confirm our theoretical results and demonstrate the effectiveness of the proposed classifier across a wide range of high-dimensional scenarios.

math.ST↗

Adaptive Reference-Guided Estimation of Principal Component Subspace in High Dimensions

We propose a novel estimator for the principal component (PC) subspace tailored to the high-dimension, low-sample size (HDLSS) context. The method, termed Adaptive Reference-Guided (ARG) estimator, is designed for data exhibiting spiked covariance structures and seeks to improve upon the conventional sample PC subspace by leveraging auxiliary information from reference vectors, presumed to carry prior knowledge about the true PC subspace. The estimator is constructed by first identifying vectors asymptotically orthogonal to the true PC subspace within a signal subspace, the subspace spanned by the leading sample PC directions and the references, and then taking the orthogonal complement. The estimator is adaptive, as it automatically selects the subspace asymptotically closest to the true PC subspace inside the signal subspace, without requiring parameter tuning. We show that when the reference vectors carry nontrivial information, the proposed estimator asymptotically reduces all principal angles between the estimated and true PC subspaces compared to the naive sample-based estimator. Interestingly, despite being derived from a completely different rationale, the ARG estimator is theoretically equivalent to an estimator based on James-Stein shrinkage. Our results thus establish a theoretical foundation that unifies these two distinct approaches.

math.ST↗

Subspace Recovery in Winsorized PCA: Insights into Accuracy and Robustness

In this paper, we explore the theoretical properties of subspace recovery using Winsorized Principal Component Analysis (WPCA), utilizing a common data transformation technique that caps extreme values to mitigate the impact of outliers. Despite the widespread use of winsorization in various tasks of multivariate analysis, its theoretical properties, particularly for subspace recovery, have received limited attention. We provide a detailed analysis of the accuracy of WPCA, showing that increasing the number of samples while decreasing the proportion of outliers guarantees the consistency of the sample subspaces from WPCA with respect to the true population subspace. Furthermore, we establish perturbation bounds that ensure the WPCA subspace obtained from contaminated data remains close to the subspace recovered from pure data. Additionally, we extend the classical notion of breakdown points to subspace-valued statistics and derive lower bounds for the breakdown points of WPCA. Our analysis demonstrates that WPCA exhibits strong robustness to outliers while maintaining consistency under mild assumptions. A toy example is provided to numerically illustrate the behavior of the upper bounds for perturbation bounds and breakdown points, emphasizing winsorization's utility in subspace recovery.

stat.ML↗

Variable selection and basis learning for ordinal classification

We propose a method for variable selection and basis learning for high-dimensional classification with ordinal responses. The proposed method extends sparse multiclass linear discriminant analysis, with the aim of identifying not only the variables relevant to discrimination but also the variables that are order-concordant with the responses. For this purpose, we compute for each variable an ordinal weight, where larger weights are given to variables with ordered group-means, and penalize the variables with smaller weights more severely. A two-step construction for ordinal weights is developed, and we show that the ordinal weights correctly separate ordinal variables from non-ordinal variables with high probability. The resulting sparse ordinal basis learning method is shown to consistently select either the discriminant variables or the ordinal and discriminant variables, depending on the choice of a tunable parameter. Such asymptotic guarantees are given under a high-dimensional asymptotic regime where the dimension grows much faster than the sample size. We also discuss a two-step procedure of post-screening ordinal variables among the selected discriminant variables. Simulated and real data analyses confirm that the proposed basis learning provides sparse and interpretable basis, as it mostly consists of ordinal variables.

stat.ME↗

Robust SVD Made Easy: A fast and reliable algorithm for large-scale data analysis

The singular value decomposition (SVD) is a crucial tool in machine learning and statistical data analysis. However, it is highly susceptible to outliers in the data matrix. Existing robust SVD algorithms often sacrifice speed for robustness or fail in the presence of only a few outliers. This study introduces an efficient algorithm, called Spherically Normalized SVD, for robust SVD approximation that is highly insensitive to outliers, computationally scalable, and provides accurate approximations of singular vectors. The proposed algorithm achieves remarkable speed by utilizing only two applications of a standard reduced-rank SVD algorithm to appropriately scaled data, significantly outperforming competing algorithms in computation times. To assess the robustness of the approximated singular vectors and their subspaces against data contamination, we introduce new notions of breakdown points for matrix-valued input, including row-wise, column-wise, and block-wise breakdown points. Theoretical and empirical analyses demonstrate that our algorithm exhibits higher breakdown points compared to standard SVD and its modifications. We empirically validate the effectiveness of our approach in applications such as robust low-rank approximation and robust principal component analysis of high-dimensional microarray datasets. Overall, our study presents a highly efficient and robust solution for SVD approximation that overcomes the limitations of existing algorithms in the presence of outliers.

stat.ML↗

A genericity property of Fréchet sample means on Riemannian manifolds

Let $(M,g)$ be a Riemannian manifold. If $μ$ is a probability measure on $M$ given by a continuous density function, one would expect the Fréchet means of data-samples $Q=(q_1,q_2,\dots, q_N)\in M^N$, with respect to $μ$, to behave ``generically''; e.g. the probability that the Fréchet mean set $\mbox{FM}(Q)$ has any elements that lie in a given, positive-codimension submanifold, should be zero for any $N\geq 1$. Even this simplest instance of genericity does not seem to have been proven in the literature, except in special cases. The main result of this paper is a general, and stronger, genericity property: given i.i.d. absolutely continuous $M$-valued random variables $X_1,\dots, X_N$, and a subset $A\subset M$ of volume-measure zero, $\mbox{Pr}\left\{\mbox{FM}(\{X_1,\dots,X_N\})\subset M\backslash A\right\}=1.$ We also establish a companion theorem for equivariant Fréchet means, defined when $(M,g)$ arises as the quotient of a Riemannian manifold $(\widetilde{M},\tilde{g})$ by a free, isometric action of a finite group. The equivariant Fréchet means lie in $\widetilde{M}$, but, as we show, project down to the ordinary Fréchet sample means, and enjoy a similar genericity property. Both these theorems are proven as consequences of a purely geometric (and quite general) result that constitutes the core mathematics in this paper: If $A\subset M$ has volume zero in $M$ , then the set $\{Q\in M^N : \mbox{FM}(Q) \cap A\neq\emptyset\}$ has volume zero in $M^N$. We conclude the paper with an application to partial scaling-rotation means, a type of mean for symmetric positive-definite matrices.

math.PR↗

Averaging symmetric positive-definite matrices on the space of eigen-decompositions

We study extensions of Fréchet means for random objects in the space ${\rm Sym}^+(p)$ of $p \times p$ symmetric positive-definite matrices using the scaling-rotation geometric framework introduced by Jung et al. [\textit{SIAM J. Matrix. Anal. Appl.} \textbf{36} (2015) 1180-1201]. The scaling-rotation framework is designed to enjoy a clearer interpretation of the changes in random ellipsoids in terms of scaling and rotation. In this work, we formally define the \emph{scaling-rotation (SR) mean set} to be the set of Fréchet means in ${\rm Sym}^+(p)$ with respect to the scaling-rotation distance. Since computing such means requires a difficult optimization, we also define the \emph{partial scaling-rotation (PSR) mean set} lying on the space of eigen-decompositions as a proxy for the SR mean set. The PSR mean set is easier to compute and its projection to ${\rm Sym}^+(p)$ often coincides with SR mean set. Minimal conditions are required to ensure that the mean sets are non-empty. Because eigen-decompositions are never unique, neither are PSR means, but we give sufficient conditions for the sample PSR mean to be unique up to the action of a certain finite group. We also establish strong consistency of the sample PSR means as estimators of the population PSR mean set, and a central limit theorem. In an application to multivariate tensor-based morphometry, we demonstrate that a two-group test using the proposed PSR means can have greater power than the two-group test using the usual affine-invariant geometric framework for symmetric positive-definite matrices.

stat.ME↗

Differentially Private Multivariate Statistics with an Application to Contingency Table Analysis

Differential privacy (DP) has become a rigorous central concept for privacy protection in the past decade. We use Gaussian differential privacy (GDP) in gauging the level of privacy protection for releasing statistical summaries from data. The GDP is a natural and easy-to-interpret differential privacy criterion based on the statistical hypothesis testing framework. The Gaussian mechanism is a natural and fundamental mechanism that can be used to perturb multivariate statistics to satisfy a $μ$-GDP criterion, where $μ>0$ stands for the level of privacy protection. Requiring a certain level of differential privacy inevitably leads to a loss of statistical utility. We improve ordinary Gaussian mechanisms by developing rank-deficient James-Stein Gaussian mechanisms for releasing private multivariate statistics, and show that the proposed mechanisms have higher statistical utilities. Laplace mechanisms, the most commonly used mechanisms in the pure DP framework, are also investigated under the GDP criterion. We show that optimal calibration of multivariate Laplace mechanisms requires more information on the statistic than just the global sensitivity, and derive the minimal amount of Laplace perturbation for releasing $μ$-GDP contingency tables. Gaussian mechanisms are shown to have higher statistical utilities than Laplace mechanisms, except for very low levels of privacy. The utility of proposed multivariate mechanisms is further demonstrated using differentially private hypotheses tests on contingency tables. Bootstrap-based goodness-of-fit and homogeneity tests, utilizing the proposed rank-deficient James--Stein mechanisms, exhibit higher powers than natural competitors.

stat.ME↗