SearcharxivSearch

arXiv subjects

Kwangok Seo

Publications and source records attributed to Kwangok Seo.

7 recordsLinked to original sources

Block-Independent Likelihood Ratio Testing for High-Dimensional Mean Vectors with Applications to Matrix-Variate Data

Testing the equality of two high-dimensional mean vectors is a fundamental problem in multivariate analysis. While the classical Hotelling's $T^2$ test is optimal in low-dimensional settings, it fails when the dimension $p$ is comparable to or exceeds the sample size $n$. Several extensions, including the Diagonal Likelihood Ratio Test (DLRT), have been proposed under the working independence assumption among variables. However, such an assumption can lead to a substantial loss of power when correlations are present. In this paper, we propose a new test, the Block Independent Likelihood Ratio Test (BILT), which generalizes DLRT by relaxing the working independence assumption to a block independence assumption. We establish its asymptotic normality of the null distribution of the BILT statistic for 'increasing $p$ with small $n$' under mild regularity conditions. We further analyze the asymptotic power of BILT under a local alternatives. Extensive simulation studies show that BILT maintains Type I error control and achieves substantially higher power than DLRT across a wide range of covariance structures. An application to the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset further demonstrates the application of BILT to testing mean differences between two matrix-variate populations.

stat.ME

Uncertainty-Aware Ideal Point Estimation via Variational EM

Roll-call data analysis aims to estimate legislators' ideal points and quantify the associated uncertainty. Existing approaches either rely on Bayesian methods implemented via Markov chain Monte Carlo sampling or focus primarily on point estimation, with uncertainty typically assessed through resampling procedures such as the bootstrap. Consequently, the computational burden of these approaches can become substantial when applied to large roll-call datasets. To address this challenge, we propose a computationally efficient likelihood method for estimating ideal points and their standard errors. Leveraging the P\'{o}lya--Gamma identity, we develop a variational expectation--maximization algorithm for estimating ideal points and introduce a variational Louis' method to approximate the observed Fisher information for standard error estimation. Numerical studies and applications to U.S. congressional roll-call data demonstrate that the proposed method produces accurate ideal point estimates and reliable standard errors while being substantially more computationally efficient than existing approaches.

stat.ME

Conformalized Method for Empirical Bayes Normal Mean Inference Problem with Heteroscedastic Variance

We study the normal mean inference problem, which involves simultaneous testing of the means of many normal distributions. This problem has been extensively studied within the empirical Bayes (EB) framework. However, the reliability of most EB methods heavily depends on two key conditions: (i) the prior distribution is correctly specified, and (ii) it can be accurately estimated. In practice, both conditions are difficult to satisfy, and it is often unclear whether they hold in a given application. To overcome these limitations, we propose a new algorithm, called COIN (COnformal Inference for Normal mean inference problem). Unlike traditional empirical Bayes approaches, COIN produces decision rules whose validity does not depend on the correct specification or accurate estimation of the prior. We theoretically prove that COIN asymptotically controls the false discovery rate at the nominal level, even in the presence of prior misspecification or estimation errors. Since the COIN algorithm requires an external training dataset to estimate the prior distribution and conformity score function, we introduce two data-splitting strategies -- sample-splitting and feature-splitting -- for the case where such external data are unavailable. We provide theoretical guarantees for the data-splitting strategies and demonstrate their effectiveness through extensive numerical studies and three real data examples.

stat.ME

On parameter estimation for the truncated skew-normal distribution

Parameter estimation for the truncated skew-normal distribution is challenging, as truncation introduces additional nonlinearity into the likelihood function and often leads to numerical instability in existing estimation procedures. In this paper, we propose a grid-based estimation method, referred to as GRID-MOM, for parameter estimation in the truncated skew-normal distribution. The proposed approach fixes the shape parameter on a pre-specified grid and, for each grid point, estimates the location and scale parameters using the method of moments. The optimal value of the shape parameter is then selected via likelihood-based comparison, yielding the final parameter estimates. By decoupling the estimation of the shape parameter from that of the location and scale parameters, the proposed method reduces the complexity of the optimization problem and improves numerical stability. We evaluate the finite-sample performance of the proposed estimator through an extensive numerical study, comparing it with existing methods under a variety of scenarios. The results demonstrate that the proposed method provides stable and accurate estimation, particularly for the shape parameter, suggesting that the proposed method offers a practical alternative for inference in truncated skew-normal models. We further demonstrate the practical applicability of the proposed method using phosphoproteomics data and hospital admission data.

stat.ME

Multiple Testing of One-Sided Hypotheses with Conservative $p$-values

We study a large-scale one-sided multiple testing problem in which test statistics follow normal distributions with unit variance, and the goal is to identify signals with positive mean effects. A conventional approach is to compute $p$-values under the assumption that all null means are exactly zero and then apply standard multiple testing procedures such as the Benjamini-Hochberg (BH) or Storey-BH method. However, because the null hypothesis is composite, some null means may be strictly negative. In this case, the resulting $p$-values are conservative, leading to a substantial loss of power. Existing methods address this issue by modifying the multiple testing procedure itself, for example through conditioning strategies or discarding rules. In contrast, we focus on correcting the $p$-values so that they are exact under the null. Specifically, we estimate the marginal null distribution of the test statistics within an empirical Bayes framework and construct refined $p$-values based on this estimated distribution. These refined $p$-values can then be directly used in standard multiple testing procedures without modification. Extensive simulation studies show that the proposed method substantially improves power when conventional $p$-values are conservative, while achieving comparable performance to existing methods when conventional $p$-values are exact. An application to phosphorylation data further demonstrates the practical effectiveness of our approach.

stat.ME

Empirical Bayes Method for Large Scale Multiple Testing with Heteroscedastic Errors

In this paper, we address the normal mean inference problem, which involves testing multiple means of normal random variables with heteroscedastic variances. Most existing empirical Bayes methods for this setting are developed under restrictive assumptions, such as the scaled inverse-chi-squared prior for variances and unimodality for the non-null mean distribution. However, when either of these assumptions is violated, these methods often fail to control the false discovery rate (FDR) at the target level or suffer from a substantial loss of power. To overcome these limitations, we propose a new empirical Bayes method, gg-Mix, which assumes only independence between the normal means and variances, without imposing any structural restrictions on their distributions. We thoroughly evaluate the FDR control and power of gg-Mix through extensive numerical studies and demonstrate its superior performance compared to existing methods. Finally, we apply gg-Mix to three real data examples to further illustrate the practical advantages of our approach.

stat.ME

$\ell_0$-Regularized Item Response Theory Model for Robust Ideal Point Estimation

Ideal point estimation methods face a significant challenge when legislators engage in protest voting -- strategically voting against their party to express dissatisfaction. Such votes introduce attenuation bias, making ideologically extreme legislators appear artificially moderate. We propose a novel statistical framework that extends the fast EM-based estimation approach of \cite{Imai2016} using $\ell_0$ regularization method to handle protest votes. Through simulation studies, we demonstrate that our proposed method maintains estimation accuracy even with high proportions of protest votes, while being substantially faster than MCMC-based methods. Applying our method to the 116th and 117th U.S. House of Representatives, we successfully recover the extreme liberal positions of ``the Squad'', whose protest votes had caused conventional methods to misclassify them as moderates. While conventional methods rank Ocasio-Cortez as more conservative than 69\% of Democrats, our method places her firmly in the progressive wing, aligning with her documented policy positions. This approach provides both robust ideal point estimates and systematic identification of protest votes, facilitating deeper analysis of strategic voting behavior in legislatures.

stat.AP