SearcharxivSearch

arXiv subjects

Sangkon Oh

Publications and source records attributed to Sangkon Oh.

4 recordsLinked to original sources

Semiparametric robust mixture of experts based on nonparametric maximum likelihood

The mixture of experts (MoE) model provides a flexible approach for modeling heterogeneous regression relationships by allowing covariate-dependent mixing through a gating network, but most existing MoE models rely on parametric assumptions for expert error distributions, typically Gaussian, which can lead to inefficiency and sensitivity to outliers or heavy-tailed behavior when misspecified. We propose a semiparametric MoE model in which each expert error distribution is represented as a nonparametric Gaussian scale mixture estimated via nonparametric maximum likelihood, relaxing parametric assumptions within the Gaussian scale-mixture class while preserving the interpretability and structure of the MoE framework. The resulting model adapts to complex error structures, improves robustness under contamination and heavy tails, and remains competitive under well-specified Gaussian settings, providing a practical and theoretically grounded alternative to parametric MoE formulations.

stat.ME

Differential gene expression analysis via two-component mixture models with a semiparametric skew-normal scale mixture alternative

Two-component mixture models are particularly useful for identifying differentially expressed genes, but their performance can deteriorate markedly when the alternative distribution departs from parametric assumptions or symmetry. We propose a semiparametric mixture model in which the null component is standard normal and the alternative follows a skew-normal scale mixture with an unspecified scale mixing distribution. This formulation accommodates skewness and heavy tails, providing a flexible and computationally tractable tool for differential gene-expression analysis without restrictive distributional assumptions. We establish identifiability and consistency of the model and develop an efficient estimation algorithm that incorporates nonparametric maximum likelihood estimation of the scale distribution. Numerical studies show notable improvements over existing parametric and nonparametric approaches for modeling the alternative distribution, and applications to colon cancer and leukemia datasets demonstrate reduced false discovery and false negative rates.

stat.ME

Mixture of partially linear experts

In the mixture of experts model, a common assumption is the linearity between a response variable and covariates. While this assumption has theoretical and computational benefits, it may lead to suboptimal estimates by overlooking potential nonlinear relationships among the variables. To address this limitation, we propose a partially linear structure that incorporates unspecified functions to capture nonlinear relationships. We establish the identifiability of the proposed model under mild conditions and introduce a practical estimation algorithm. We present the performance of our approach through numerical studies, including simulations and real data analysis.

stat.ME

Adaptive Accelerated Failure Time modeling with a Semiparametric Skewed Error Distribution

The accelerated failure time (AFT) model is widely used to analyze relationships between variables in the presence of censored observations. However, this model relies on some assumptions such as the error distribution, which can lead to biased or inefficient estimates if these assumptions are violated. In order to overcome this challenge, we propose a novel approach that incorporates a semiparametric skew-normal scale mixture distribution for the error term in the AFT model. By allowing for more flexibility and robustness, this approach reduces the risk of misspecification and improves the accuracy of parameter estimation. We investigate the identifiability and consistency of the proposed model and develop a practical estimation algorithm. To evaluate the performance of our approach, we conduct extensive simulation studies and real data analyses. The results demonstrate the effectiveness of our method in providing robust and accurate estimates in various scenarios.

stat.ME