SearcharxivSearch

arXiv subjects

Min Tsao

Publications and source records attributed to Min Tsao.

14 recordsLinked to original sources

Maximum Likelihood Criterion for Non-nested Model Selection

Penalization is a widely used approach to model selection with roots in information theory and Bayesian inference. We study a model selection problem involving non-nested candidate models for which penalization is counterproductive. We propose a Maximum Likelihood Criterion for this non-nested setting that selects the candidate model with the highest maximum likelihood. This criterion does not take into consideration the number of parameters of a candidate model. It is well-suited for situations where all candidate models are regarded as equal with no preference for models having fewer parameters. We establish the consistency of this criterion and compare its performance with that of existing penalization-based criteria.

stat.ME

Estimation of Bivariate Normal Distributions from Marginal Summaries in Clinical Trials

In certain privacy-sensitive scenarios within fields such as clinical trial simulations, federated learning, and distributed learning, researchers often face the challenge of estimating correlations between variables without access to individual-level data. To address this issue, we propose a novel method to estimate the correlation of bivariate normal variables using marginal information from multiple datasets. The method, based on maximum likelihood estimation (MLE), accommodates datasets with varying sample sizes and avoids reliance on sensitive information such as sample covariances, making it particularly suitable for privacy-restricted settings. Extensive simulation studies demonstrate the proposed method's effectiveness in accurately estimating correlations and its robustness across diverse data configurations.

stat.ME

Estimating the Joint Distribution of Two Binary Variables with Marginal Statistics

Clinical trial simulation (CTS) is critical in new drug development, providing insight into safety and efficacy while guiding trial design. Achieving realistic outcomes in CTS requires an accurately estimated joint distribution of the underlying variables. However, privacy concerns and data availability issues often restrict researchers to marginal summary-level data of each variable, making it challenging to estimate the joint distribution due to the lack of access to individual-level data or relational summaries between variables. We propose a novel approach based on the method of maximum likelihood that estimates the joint distribution of two binary variables using only marginal summary data. By leveraging numerical optimization and accommodating varying sample sizes across studies, our method preserves privacy while bypassing the need for granular or relational data. Through an extensive simulation study covering a diverse range of scenarios and an application to a real-world dataset, we demonstrate the accuracy, robustness, and practicality of our method. This method enhances the generation of realistic simulated data, thereby improving decision-making processes in drug development.

stat.ME

Sparse maximum likelihood estimation for regression models

For regression model selection via maximum likelihood estimation, we adopt a vector representation of candidate models and study the likelihood ratio confidence region for the regression parameter vector of a full model. We show that when its confidence level increases with the sample size at a certain speed, with probability tending to one, the confidence region consists of vectors representing models containing all active variables, including the true parameter vector of the full model. Using this result, we examine the asymptotic composition of models of maximum likelihood and find the subset of such models that contain all active variables. We then devise a consistent model selection criterion which has a sparse maximum likelihood estimation interpretation and certain advantages over popular information criteria.

math.ST

Group least squares regression for linear models with strongly correlated predictor variables

Traditionally, the least squares regression is mainly concerned with studying the effects of individual predictor variables, but strongly correlated variables generate multicollinearity which makes it difficult to study their effects. Existing methods for handling multicollinearity such as ridge regression are complicated. To resolve the multicollinearity issue without abandoning the simple least squares regression, for situations where predictor variables are in groups with strong within-group correlations but weak between-group correlations, we propose to study the effects of the groups with a group approach to the least squares regression. Using an all positive correlations arrangement of the strongly correlated variables, we first characterize group effects that are meaningful and can be accurately estimated. We then present the group approach with numerical examples and demonstrate its advantages over existing methods for handling multicollinearity. We also address a common misconception about prediction accuracy of the least squares estimated model and discuss through an example similar group effects in generalized linear models.

stat.ME

Average group effect of strongly correlated predictor variables is estimable

It is well known that individual parameters of strongly correlated predictor variables in a linear model cannot be accurately estimated by the least squares regression due to multicollinearity generated by such variables. Surprisingly, an average of these parameters can be extremely accurately estimated. We find this average and briefly discuss its applications in the least squares regression.

math.ST

Regression model selection via log-likelihood ratio and constrained minimum criterion

Although the log-likelihood is widely used in model selection, the log-likelihood ratio has had few applications in this area. We develop a log-likelihood ratio based method for selecting regression models by focusing on the set of models deemed plausible by the likelihood ratio test. We show that when the sample size is large and the significance level of the test is small, there is a high probability that the smallest model in the set is the true model; thus, we select this smallest model. The significance level of the test serves as a parameter for this method. We consider three levels of this parameter in a simulation study and compare this method with the Akaike Information Criterion and Bayesian Information Criterion to demonstrate its excellent accuracy and adaptability to different sample sizes. We also apply this method to select a logistic regression model for a South African heart disease dataset.

stat.ME

A constrained minimum criterion for model selection

We propose a hypothesis test based model selection criterion for the best subset selection of sparse linear models. We show it is consistent in that the probability of its choosing the true model approaches one and the parameter values of its chosen model converge in probability to that of the true model as the sample size goes to infinity. This criterion is capable of controlling the balance between the false active rate and false inactive rate of the selected model, and it can be applied with other methods of model selection such as the lasso. We also demonstrate its accuracy and advantages with a numerical comparison and an application.

stat.ME

Estimable group effects for strongly correlated variables in linear models

It is well known that parameters for strongly correlated predictor variables in a linear model cannot be accurately estimated. We look for linear combinations of these parameters that can be. Under a uniform model, we find such linear combinations in a neighborhood of a simple variability weighted average of these parameters. Surprisingly, this variability weighted average is more accurately estimated when the variables are more strongly correlated, and it is the only linear combination with this property. It can be easily computed for strongly correlated predictor variables in all linear models and has applications in inference and estimation concerning parameters of such variables.

math.ST

Two-sample extended empirical likelihood for estimating equations

We propose a two-sample extended empirical likelihood for inference on the difference between two p-dimensional parameters defined by estimating equations. The standard two-sample empirical likelihood for the difference is Bartlett correctable but its domain is a bounded subset of the parameter space. We expand its domain through a composite similarity transformation to derive the two-sample extended empirical likelihood which is defined on the full parameter space. The extended empirical likelihood has the same asymptotic distribution as the standard one and can also achieve the second order accuracy of the Bartlett correction. We include two applications to illustrate the use of two-sample empirical likelihood methods and to demonstrate the superior coverage accuracy of the extended empirical likelihood confidence regions.

math.ST

Empirical likelihood on the full parameter space

We extend the empirical likelihood of Owen [Ann. Statist. 18 (1990) 90-120] by partitioning its domain into the collection of its contours and mapping the contours through a continuous sequence of similarity transformations onto the full parameter space. The resulting extended empirical likelihood is a natural generalization of the original empirical likelihood to the full parameter space; it has the same asymptotic properties and identically shaped contours as the original empirical likelihood. It can also attain the second order accuracy of the Bartlett corrected empirical likelihood of DiCiccio, Hall and Romano [Ann. Statist. 19 (1991) 1053-1061]. A simple first order extended empirical likelihood is found to be substantially more accurate than the original empirical likelihood. It is also more accurate than available second order empirical likelihood methods in most small sample situations and competitive in accuracy in large sample situations. Importantly, in many one-dimensional applications this first order extended empirical likelihood is accurate for sample sizes as small as ten, making it a practical and reliable choice for small sample empirical likelihood inference.

math.ST

Multivariate two-sample extended empirical likelihood

Jing (1995) and Liu et al. (2008) studied the two-sample empirical likelihood and showed it is Bartlett correctable for the univariate and multivariate cases, respectively. We expand its domain to the full parameter space and obtain a two-sample extended empirical likelihood which is more accurate and can also achieve the second-order accuracy of the Bartlett correction.

math.ST

Extended empirical likelihood for general estimating equations

We derive an extended empirical likelihood for parameters defined by estimating equations which generalizes the original empirical likelihood for such parameters to the full parameter space. Under mild conditions, the extended empirical likelihood has all asymptotic properties of the original empirical likelihood. Its contours retain the data-driven shape of the latter. It can also attain the second order accuracy. The first order extended empirical likelihood is easy-to-use yet it is substantially more accurate than other empirical likelihoods, including second order ones. We recommend it for practical applications of the empirical likelihood method.

math.ST