SearcharxivSearch

arXiv subjects

Hien D Nguyen

Publications and source records attributed to Hien D Nguyen.

17 recordsLinked to original sources

Non-asymptotic oracle inequalities for the Lasso in high-dimensional mixture of experts

We investigate the estimation properties of the mixture of experts (MoE) model in a high-dimensional setting, where the number of predictors is much larger than the sample size, and for which the literature is particularly lacking in theoretical results. We consider the class of softmax-gated Gaussian MoE (SGMoE) models, defined as MoE models with softmax gating functions and Gaussian experts, and focus on the theoretical properties of their $l_1$-regularized estimation via the Lasso. To the best of our knowledge, we are the first to investigate the $l_1$-regularization properties of SGMoE models from a non-asymptotic perspective, under the mildest assumptions, namely the boundedness of the parameter space. We provide a lower bound on the regularization parameter of the Lasso penalty that ensures non-asymptotic theoretical control of the Kullback--Leibler loss of the Lasso estimator for SGMoE models. Finally, we carry out a simulation study to empirically validate our theoretical findings.

math.ST

Finite sample inference for empirical Bayesian methods

In recent years, empirical Bayesian (EB) inference has become an attractive approach for estimation in parametric models arising in a variety of real-life problems, especially in complex and high-dimensional scientific applications. However, compared to the relative abundance of available general methods for computing point estimators in the EB framework, the construction of confidence sets and hypothesis tests with good theoretical properties remains difficult and problem specific. Motivated by the universal inference framework of Wasserman et al. (2020), we propose a general and universal method, based on holdout likelihood ratios, and utilizing the hierarchical structure of the specified Bayesian model for constructing confidence sets and hypothesis tests that are finite sample valid. We illustrate our method through a range of numerical studies and real data applications, which demonstrate that the approach is able to generate useful and meaningful inferential statements in the relevant contexts.

stat.ME

Order selection with confidence for finite mixture models

The determination of the number of mixture components (the order) of a finite mixture model has been an enduring problem in statistical inference. We prove that the closed testing principle leads to a sequential testing procedure (STP) that allows for confidence statements to be made regarding the order of a finite mixture model. We construct finite sample tests, via data splitting and data swapping, for use in the STP, and we prove that such tests are consistent against fixed alternatives. Simulation studies and real data examples are used to demonstrate the performance of the finite sample tests-based STP, yielding practical recommendations of their use as confidence estimators in combination with point estimates such as the Akaike information or Bayesian information criteria. In addition, we demonstrate that a modification of the STP yields a method that consistently selects the order of a finite mixture model, in the asymptotic sense. Our STP is not only applicable for order selection of finite mixture models, but is also useful for making confidence statements regarding any sequence of nested models.

stat.ME

Approximation of probability density functions via location-scale finite mixtures in Lebesgue spaces

The class of location-scale finite mixtures is of enduring interest both from applied and theoretical perspectives of probability and statistics. We prove the following results: to an arbitrary degree of accuracy, (a) location-scale mixtures of a continuous probability density function (PDF) can approximate any continuous PDF, uniformly, on a compact set; and (b) for any finite $p\ge1$, location-scale mixtures of an essentially bounded PDF can approximate any PDF in $\mathcal{L}_{p}$, in the $\mathcal{L}_{p}$ norm.

math.ST

Universal inference with composite likelihoods

Maximum composite likelihood estimation is a useful alternative to maximum likelihood estimation when data arise from data generating processes (DGPs) that do not admit tractable joint specification. We demonstrate that generic composite likelihoods consisting of marginal and conditional specifications permit the simple construction of composite likelihood ratio-like statistics from which finite-sample valid confidence sets and hypothesis tests can be constructed. These statistics are universal in the sense that they can be constructed from any estimator for the parameter of the underlying DGP. We demonstrate our methodology via a simulation study using a pair of conditionally specified bivariate models.

stat.ME

A binary-response regression model based on support vector machines

The soft-margin support vector machine (SVM) is a ubiquitous tool for prediction of binary-response data. However, the SVM is characterized entirely via a numerical optimization problem, rather than a probability model, and thus does not directly generate probabilistic inferential statements as outputs. We consider a probabilistic regression model for binary-response data that is based on the optimization problem that characterizes the SVM. Under weak regularity assumptions, we prove that the maximum likelihood estimate (MLE) of our model exists, and that it is consistent and asymptotically normal. We further assess the performance of our model via simulation studies, and demonstrate its use in real data applications regarding spam detection and well water access.

stat.ME

Approximation by finite mixtures of continuous density functions that vanish at infinity

Given sufficiently many components, it is often cited that finite mixture models can approximate any other probability density function (pdf) to an arbitrary degree of accuracy. Unfortunately, the nature of this approximation result is often left unclear. We prove that finite mixture models constructed from pdfs in $\mathcal{C}_{0}$ can be used to conduct approximation of various classes of approximands in a number of different modes. That is, we prove approximands in $\mathcal{C}_{0}$ can be uniformly approximated, approximands in $\mathcal{C}_{b}$ can be uniformly approximated on compact sets, and approximands in $\mathcal{L}_{p}$ can be approximated with respect to the $\mathcal{L}_{p}$, for $p\in\left[1,\infty\right)$. Furthermore, we also prove that measurable functions can be approximated, almost everywhere.

math.ST

Asymptotic normality of the time-domain generalized least squares estimator for linear regression models

In linear models, the generalized least squares (GLS) estimator is applicable when the structure of the error dependence is known. When it is unknown, such structure must be approximated and estimated in a manner that may lead to misspecification. The large-sample analysis of incorrectly-specified GLS (IGLS) estimators requires careful asymptotic manipulations. When performing estimation in the frequency domain, the asymptotic normality of the IGLS estimator, under the so-called Grenander assumptions, has been proved for a broad class of error dependence models. Under the same assumptions, asymptotic normality results for the time-domain IGLS estimator are only available for a limited class of error structures. We prove that the time-domain IGLS estimator is asymptotically normal for a general class of dependence models.

math.ST

A Simple Online Parameter Estimation Technique with Asymptotic Guarantees

In many modern settings, data are acquired iteratively over time, rather than all at once. Such settings are known as online, as opposed to offline or batch. We introduce a simple technique for online parameter estimation, which can operate in low memory settings, settings where data are correlated, and only requires a single inspection of the available data at each time period. We show that the estimators---constructed via the technique---are asymptotically normal under generous assumptions, and present a technique for the online computation of the covariance matrices for such estimators. A set of numerical studies demonstrates that our estimators can be as efficient as their offline counterparts, and that our technique generates estimates and confidence intervals that match their offline counterparts in various parameter estimation settings.

stat.CO

A Note on the Convergence of the Gaussian Mean Shift Algorithm

Mean shift (MS) algorithms are popular methods for mode finding in pattern analysis. Each MS algorithm can be phrased as a fixed-point iteration scheme, which operates on a kernel density estimate (KDE) based on some data. The ability of an MS algorithm to obtain the modes of its KDE depends on whether or not the fixed-point scheme converges. The convergence of MS algorithms have recently been proved under some general conditions via first principle arguments. We complement the recent proofs by demonstrating that the MS algorithm operating on a Gaussian KDE can be viewed as an MM (minorization-maximization) algorithm, and thus permits the application of convergence techniques for such constructions. For the Gaussian case, we extend upon the previously results by showing that the fixed-points of the MS algorithm are all stationary points of the KDE in cases where the stationary points may not necessarily be isolated.

stat.CO

Faster Functional Clustering via Gaussian Mixture Models

Functional data analysis (FDA) is an important modern paradigm for handling infinite-dimensional data. An important task in FDA is model-based clustering, which organizes functional populations into groups via subpopulation structures. The most common approach for model-based clustering of functional data is via mixtures of linear mixed-effects models. The mixture of linear mixed-effects models (MLMM) approach requires a computationally intensive algorithm for estimation. We provide a novel Gaussian mixture model (GMM) characterization of the model-based clustering problem. We demonstrate that this GMM-based characterization allows for improved computational speeds over the MLMM approach when applied via available functions in the R programming environment. Theoretical considerations for the GMM approach are discussed. An example application to a dataset based upon calcium imaging in the larval zebrafish brain is provided as a demonstration of the effectiveness of the simpler GMM approach.

stat.CO

Some Theoretical Results Regarding the Polygonal Distribution

The polygonal distributions are a class of distributions that can be defined via the mixture of triangular distributions over the unit interval. The class includes the uniform and trapezoidal distributions, and is an alternative to the beta distribution. We demonstrate that the polygonal densities are dense in the class of continuous and concave densities with bounded second derivatives. Pointwise consistency and Hellinger consistency results for the maximum likelihood (ML) estimator are obtained. A useful model selection theorem is stated as well as results for a related distribution that is obtained via the pointwise square of polygonal density functions.

stat.ME

Progress on a Conjecture Regarding the Triangular Distribution

Triangular distributions are a well-known class of distributions that are often used as an elementary example of a probability model. Maximum likelihood estimation of the mode parameter of the triangular distribution over the unit interval can be performed via an order statistics-based method. It had been conjectured that such a method can be conducted using only a constant number of likelihood function evaluations, on average, as the sample size becomes large. We prove two theorems that validate this conjecture. Graphical and numerical results are presented to supplement our proofs.

stat.OT

Maximum Pseudolikelihood Estimation for Model-Based Clustering of Time Series Data

Mixture of autoregressions (MoAR) models provide a model-based approach to the clustering of time series data. The maximum likelihood (ML) estimation of MoAR models requires the evaluation of products of large numbers of densities of normal random variables. In practical scenarios, these products converge to zero as the length of the time series increases, and thus the ML estimation of MoAR models becomes infeasible without the use of numerical tricks. We propose a maximum pseudolikelihood (MPL) estimation approach as an alternative to the use of numerical tricks. The MPL estimator is proved to be consistent and can be computed via an EM (expectation--maximization) algorithm. Simulations are used to assess the performance of the MPL estimator against that of the ML estimator in cases where the latter was able to be calculated. An application to the clustering of time series data arising from a resting-state fMRI experiment is presented as a demonstration of the methodology.

stat.CO

Maximum Likelihood Estimation of Triangular and Polygonal Distributions

Triangular distributions are a well-known class of distributions that are often used as elementary example of a probability model. In the past, enumeration and order statistic-based methods have been suggested for the maximum likelihood (ML) estimation of such distributions. A novel parametrization of triangular distributions is presented. The parametrization allows for the construction of an MM (minorization--maximization) algorithm for the ML estimation of triangular distributions. The algorithm is shown to both monotonically increase the likelihood evaluations, and be globally convergent. Using the parametrization is then applied to construct an MM algorithm for the ML estimation of polygonal distributions. This algorithm is shown to have the same numerical properties as that of the triangular distribution. Numerical simulation are provided to demonstrate the performances of the new algorithms against established enumeration and order statistics-based methods.

stat.CO

A Universal Approximation Theorem for Mixture of Experts Models

The mixture of experts (MoE) model is a popular neural network architecture for nonlinear regression and classification. The class of MoE mean functions is known to be uniformly convergent to any unknown target function, assuming that the target function is from Sobolev space that is sufficiently differentiable and that the domain of estimation is a compact unit hypercube. We provide an alternative result, which shows that the class of MoE mean functions is dense in the class of all continuous functions over arbitrary compact domains of estimation. Our result can be viewed as a universal approximation theorem for MoE models.

stat.ML

Spatial Clustering of Time-Series via Mixture of Autoregressions Models and Markov Random Fields

Time-series data arise in many medical and biological imaging scenarios. In such images, a time-series is obtained at each of a large number of spatially-dependent data units. It is interesting to organize these data into model-based clusters. A two-stage procedure is proposed. In Stage 1, a mixture of autoregressions (MoAR) model is used to marginally cluster the data. The MoAR model is fitted using maximum marginal likelihood (MMaL) estimation via an MM (minorization--maximization) algorithm. In Stage 2, a Markov random field (MRF) model induces a spatial structure onto the Stage 1 clustering. The MRF model is fitted using maximum pseudolikelihood (MPL) estimation via an MM algorithm. Both the MMaL and MPL estimators are proved to be consistent. Numerical properties are established for both MM algorithms. A simulation study demonstrates the performance of the two-stage procedure. An application to the segmentation of a zebrafish brain calcium image is presented.

stat.ME