SearcharxivSearch

arXiv subjects

Claudio Fuentes

Publications and source records attributed to Claudio Fuentes.

9 recordsLinked to original sources

The Importance of Discussing Assumptions when Teaching Bootstrapping

Bootstrapping and other resampling methods are increasingly appearing in the textbooks and curricula of courses that introduce undergraduate students to statistical methods. In order to teach the bootstrap well, students and instructors need to be aware of the assumptions behind these intervals. In this article we discuss important assumptions about simple non-parametric bootstrap intervals and their corresponding hypothesis tests. We present simulations that instructors can use to help students understand some of the assumptions behind these methods. The simulations will be especially relevant to instructors who desire to increase accessibility for students from non-mathematical backgrounds, including those with math anxiety.

stat.OT

A Nested Weighted Tchebycheff Multi-Objective Bayesian Optimization Approach for Flexibility of Unknown Utopia Estimation in Expensive Black-box Design Problems

We propose a nested weighted Tchebycheff Multi-objective Bayesian optimization framework where we build a regression model selection procedure from an ensemble of models, towards better estimation of the uncertain parameters of the weighted-Tchebycheff expensive black-box multi-objective function. In existing work, a weighted Tchebycheff MOBO approach has been demonstrated which attempts to estimate the unknown utopia in formulating acquisition function, through calibration using a priori selected regression model. However, the existing MOBO model lacks flexibility in selecting the appropriate regression models given the guided sampled data and therefore, can under-fit or over-fit as the iterations of the MOBO progress, reducing the overall MOBO performance. As it is too complex to a priori guarantee a best model in general, this motivates us to consider a portfolio of different families of predictive models fitted with current training data, guided by the WTB MOBO; the best model is selected following a user-defined prediction root mean-square-error-based approach. The proposed approach is implemented in optimizing a multi-modal benchmark problem and a thin tube design under constant loading of temperature-pressure, with minimizing the risk of creep-fatigue failure and design cost. Finally, the nested weighted Tchebycheff MOBO model performance is compared with different MOBO frameworks with respect to accuracy in parameter estimation, Pareto-optimal solutions and function evaluation cost. This method is generalized enough to consider different families of predictive models in the portfolio for best model selection, where the overall design architecture allows for solving any high-dimensional (multiple functions) complex black-box problems and can be extended to any other global criterion multi-objective optimization methods where prior knowledge of utopia is required.

cs.LG

A Linear Mixed Model Formulation for Spatio-Temporal Random Processes with Computational Advances for the Separable and Product-Sum Covariances

We describe spatio-temporal random processes using linear mixed models. We show how many commonly used models can be viewed as special cases of this general framework and pay close attention to models with separable or product-sum covariances. The proposed linear mixed model formulation facilitates the implementation of a novel algorithm using Stegle eigendecompositions, a recursive application of the Sherman-Morrison-Woodbury formula, and Helmert-Wolf blocking to efficiently invert separable and product-sum covariance matrices, even when every spatial location is not observed at every time point. We show our algorithm provides noticeable improvements over the standard Cholesky decomposition approach. Via simulations, we assess the performance of the separable and product-sum covariances and identify scenarios where separable covariances are noticeably inferior to product-sum covariances. We also compare likelihood-based and semivariogram-based estimation and discuss benefits and drawbacks of both. We use the proposed approach to analyze daily maximum temperature data in Oregon, USA, during the 2019 summer. We end by offering guidelines for choosing among these covariances and estimation methods based on properties of observed data.

stat.ME

A Bayesian Nonparametric Model for Predicting Pregnancy Outcomes Using Longitudinal Profiles

Across several medical fields, developing an approach for disease classification is an important challenge. The usual procedure is to fit a model for the longitudinal response in the healthy population, a different model for the longitudinal response in disease population, and then apply the Bayes' theorem to obtain disease probabilities given the responses. Unfortunately, when substantial heterogeneity exists within each population, this type of Bayes classification may perform poorly. In this paper, we develop a new approach by fitting a Bayesian nonparametric model for the joint outcome of disease status and longitudinal response, and then use the clustering induced by the Dirichlet process in our model to increase the flexibility of the method, allowing for multiple subpopulations of healthy, diseased, and possibly mixed membership. In addition, we introduce an MCMC sampling scheme that facilitates the assessment of the inference and prediction capabilities of our model. Finally, we demonstrate the method by predicting pregnancy outcomes using longitudinal profiles on the $\beta$--HCG hormone levels in a sample of Chilean women being treated with assisted reproductive therapy.

stat.AP

A Constrained Conditional Likelihood Approach for Estimating the Means of Selected Populations

Given p independent normal populations, we consider the problem of estimating the mean of those populations, that based on the observed data, give the strongest signals. We explicitly condition on the ranking of the sample means, and consider a constrained conditional maximum likelihood (CCMLE) approach, avoiding the use of any priors and of any sparsity requirement between the population means. Our results show that if the observed means are too close together, we should in fact use the grand mean to estimate the mean of the population with the larger sample mean. If they are separated by more than a certain threshold, we should shrink the observed means towards each other. As intuition suggests, it is only if the observed means are far apart that we should conclude that the magnitude of separation and consequent ranking are not due to chance. Unlike other methods, our approach does not need to pre-specify the number of selected populations and the proposed CCMLE is able to perform simultaneous inference. Our method, which is conceptually straightforward, can be easily adapted to incorporate other selection criteria. Selected populations, Maximum likelihood, Constrained MLE, Post-selection inference

stat.ME

Intrinsic Bayesian Analysis for Occupancy Models

Occupancy models are typically used to determine the probability of a species being present at a given site while accounting for imperfect detection. The survey data underlying these models often include information on several predictors that could potentially characterize habitat suitability and species detectability. Because these variables might not all be relevant, model selection techniques are necessary in this context. In practice, model selection is performed using the Akaike Information Criterion (AIC), as few other alternatives are available. This paper builds an objective Bayesian variable selection framework for occupancy models through the intrinsic prior methodology. The procedure incorporates priors on the model space that account for test multiplicity and respect the polynomial hierarchy of the predictors when higher-order terms are considered. The methodology is implemented using a stochastic search algorithm that is able to thoroughly explore large spaces of occupancy models. The proposed strategy is entirely automatic and provides control of false positives without sacrificing the discovery of truly meaningful covariates. The performance of the method is evaluated and compared to AIC through a simulation study. The method is illustrated on two datasets previously studied in the literature.

stat.ME

Bayesian Estimation of Negative Binomial Parameters with Applications to RNA-Seq Data

RNA-Seq data characteristically exhibits large variances, which need to be appropriately accounted for in the model. We first explore the effects of this variability on the maximum likelihood estimator (MLE) of the overdispersion parameter of the negative binomial distribution, and propose instead the use an estimator obtained via maximization of the marginal likelihood in a conjugate Bayesian framework. We show, via simulation studies, that the marginal MLE can better control this variation and produce a more stable and reliable estimator. We then formulate a conjugate Bayesian hierarchical model, in which the estimate of overdispersion is a marginalized estimate and use this estimator to propose a Bayesian test to detect differentially expressed genes with RNA-Seq data. We use numerical studies to show that our much simpler approach is competitive with other negative binomial based procedures, and we use a real data set to illustrate the implementation and flexibility of the procedure.

stat.ME

The matryoshka doll prior: principled multiplicity correction in Bayesian model comparison

This paper introduces a general and principled construction of model space priors with a focus on regression problems. The proposed formulation regards each model as a `local` null hypothesis whose alternatives are the set of models that nest it. Assuming constant odds between any `local` null and its alternatives provides a natural isomorphism of model spaces (like a matryoshka doll), constituting an intuitive way to correct for test multiplicity. This isomorphism yields the Poisson distribution as the unique limiting distribution over model dimension under mild assumptions. We compare this model space prior theoretically and in simulations to widely adopted Beta-Binomial constructions. We show that the proposed prior yields a `just-right` multiplicity correction that induces a desirable complexity penalization profile.

stat.ME