SearcharxivSearch

arXiv subjects

Mark J. Meyer

Publications and source records attributed to Mark J. Meyer.

9 recordsLinked to original sources

Objective Priors for the Conway-Maxwell-Poisson (COM-Poisson) Distribution

The Conway-Maxwell-Poisson (COM-Poisson) distribution is a flexible, two parameter distribution for count data that can accommodate equidispersion, overdispersion, and underdispersion in the data. Previous Bayesian evaluations of the COM-Poisson distribution have largely focused on the regression context with little discussion on objective priors for the base model. Additionally, we find the multivariate Jeffreys' prior is inadequate for this model. Motivated by this, we propose four different objective priors when using the COM-Poisson distribution including an independent Jefferys' prior model and a novel reference prior derived using the hierarchical approach for objective priors. Since the resulting reference prior is non-standard, we implement constrained nonlinear optimization by linear approximation to obtain the reference prior parameters for an Empirical Bayes approach. We also develop Stan code to perform No U-Turn Sampling on each model. We compare all four approaches in simulation and demonstrate the models in several data illustrations that span the spectrum of dispersion types.

stat.ME

Sign and signed rank tests for paired functions

Simple nonparametric tests for paired functional data are an understudied area, despite recent advances in similar tests for other types of functional data. While the sign test has received limited treatment, the signed rank-type test has not previously been examined. The aim of the present work is to develop and evaluate these types of tests for functional data. We derive a simple, theoretical framework for both sign and signed rank tests for pairs of functions. In particular, we demonstrate that doubly ranked testing -- a newly developed framework for testing hypotheses involving functional data -- is a useful conduit for examining hypotheses regarding pairs o,f functions. We briefly examine the operating characteristics of all derived tests. We also use the described approaches to re-analyze pairs of functions from a randomized crossover study of heart health during simulated flight.

stat.ME

Doubly ranked tests of location for grouped functional data

Nonparametric tests for functional data are a challenging class of tests to work with because of the potentially high dimensional nature of the data. One of the main challenges for considering rank-based tests, like the Mann-Whitney or Wilcoxon Rank Sum tests (MWW), is that the unit of observation is typically a curve. Thus any rank-based test must consider ways of ranking curves. While several procedures, including depth-based methods, have recently been used to create scores for rank-based tests, these scores are not constructed under the null and often introduce additional, uncontrolled for variability. We therefore reconsider the problem of rank-based tests for functional data and develop an alternative approach that incorporates the null hypothesis throughout. Our approach first ranks realizations from the curves at each measurement occurrence, then calculates a summary statistic for the ranks of each subject, and finally re-ranks the summary statistic in a procedure we refer to as a doubly ranked test. We propose two summaries for the middle step: a sufficient statistic and the average rank. As we demonstrate, doubly rank tests are more powerful while maintaining ideal type I error in the two sample, MWW setting. We also extend our framework to more than two samples, developing a Kruskal-Wallis test for functional data which exhibits good test characteristics as well. Finally, we illustrate the use of doubly ranked tests in functional data contexts from material science, climatology, and public health policy.

stat.ME

Global Tests for Smoothed Functions in Mean Field Variational Additive Models

Variational regression methods are an increasingly popular tool for their efficient estimation of complex. Given the mixed model representation of penalized effects, additive regression models with smoothed effects and scalar-on-function regression models can be fit relatively efficiently in a variational framework. However, inferential procedures for smoothed and functional effects in such a context is limited. We demonstrate that by using the Mean Field Variational Bayesian (MFVB) approximation to the additive model and the subsequent Coordinate Ascent Variational Inference (CAVI) algorithm, we can obtain a form of the estimated effects required of a Frequentist test for semiparametric curves. We establish MFVB approximations and CAVI algorithms for both Gaussian and binary additive models with an arbitrary number of smoothed and functional effects. We then derive a global testing framework for smoothed and functional effects. Our empirical study demonstrates that the test maintains good Frequentist properties in the variational framework and can be used to directly test results from a converged, MFVB approximation and CAVI algorithm. We illustrate the applicability of this approach in a wide range of data illustrations.

stat.ME

Model Selection in Variational Mixed Effects Models

Variational inference is an alternative estimation technique for Bayesian models. Recent work shows that variational methods provide consistent estimation via efficient, deterministic algorithms. Other tools, such as model selection using variational AICs (VAIC) have been developed and studied for the linear regression case. While mixed effects models have enjoyed some study in the variational context, tools for model selection are lacking. One important feature of model selection in mixed effects models, particularly longitudinal models, is the selection of the random effects which in turn determine the covariance structure for the repeatedly sampled outcome. To address this, we derive a VAIC specifically for variational mixed effects (VME) models. We also implement a parameter-efficient VME as part of our study which reduces any general random effects structure down to a single subject-specific score. This model accommodates a wide range of random effect structures including random intercept and slope models as well as random functional effects. Our VAIC can model and perform selection on a variety of VME models including more classic longitudinal models as well as longitudinal scalar-on-function regression. As we demonstrate empirically, our VAIC performs well in discriminating between correctly and incorrectly specified random effects structures. Finally, we illustrate the use of VAICs for VMEs on two datasets: a study of lead levels in children and a study of diffusion tensor imaging.

stat.ME

Bayesian Analysis of Multivariate Matched Proportions with Sparse Response

Multivariate matched proportions (MMP) data appears in a variety of contexts including post-market surveillance of adverse events in pharmaceuticals, disease classification, and agreement between care providers. It consists of multiple sets of paired binary measurements taken on the same subject. While recent work proposes non-Bayesian methods to address the complexities of MMP data, the issue of sparse response, where no or very few "yes" responses are recorded for one or more sets, is unaddressed. The presence of sparse response sets results in underestimates of variance, loss of coverage, and lowered power in existing methods. Bayesian methods have not previously been considered for MMP data but provide a useful framework when sparse responses are present. In particular, the Bayesian probit model provides an elegant solution to the problem of variance underestimation. We examine three approaches built on that model: a naive analysis with flat priors, a penalized analysis using half-Cauchy priors on the mean model variances, and a multivariate analysis with a Bayesian functional principal component analysis (FPCA) to model the latent covariance. We show that the multivariate analysis performs well on MMP data with sparse responses and outperforms existing non-Bayesian methods. In a re-analysis of data from a study of the system of care (SOC) framework for children with mental and behavioral disorders, we are able to provide a more complete picture of the relationships in the data. Our analysis provides additional insights into the functioning on the SOC that a previous univariate analysis missed.

stat.ME

Ordinal Probit Functional Outcome Regression with Application to Computer-Use Behavior in Rhesus Monkeys

Research in functional regression has made great strides in expanding to non-Gaussian functional outcomes, but exploration of ordinal functional outcomes remains limited. Motivated by a study of computer-use behavior in rhesus macaques (Macaca mulatta), we introduce the Ordinal Probit Functional Outcome Regression model (OPFOR). OPFOR models can be fit using one of several basis functions including penalized B-splines, wavelets, and O'Sullivan splines -- the last of which typically performs best. Simulation using a variety of underlying covariance patterns shows that the model performs reasonably well in estimation under multiple basis functions with near nominal coverage for joint credible intervals. Finally, in application, we use Bayesian model selection criteria adapted to functional outcome regression to best characterize the relation between several demographic factors of interest and the monkeys' computer use over the course of a year. In comparison with a standard ordinal longitudinal analysis, OPFOR outperforms a cumulative-link mixed-effects model in simulation and provides additional and more nuanced information on the nature of the monkeys' computer-use behavior.

stat.ME

Function-on-Function Regression for the Identification of Epigenetic Regions Exhibiting Windows of Susceptibility to Environmental Exposures

The ability to identify time periods when individuals are most susceptible to exposures, as well as the biological mechanisms through which these exposures act, is of great public health interest. Growing evidence supports an association between prenatal exposure to air pollution and epigenetic marks, such as DNA methylation, but the timing and gene-specific effects of these epigenetic changes are not well understood. Here, we present the first study that aims to identify prenatal windows of susceptibility to air pollution exposures in cord blood DNA methylation. In particular, we propose a function-on-function regression model that leverages data from nearby DNA methylation probes to identify epigenetic regions that exhibit windows of susceptibility to ambient particulate matter less than 2.5 microns (PM$_{2.5}$). By incorporating the covariance structure among both the multivariate DNA methylation outcome and the time-varying exposure under study, this framework yields greater power to detect windows of susceptibility and greater control of false discoveries than methods that model probes independently. We compare our method to a distributed lag model approach that models DNA methylation in a probe-by-probe manner, both in simulation and by application to motivating data from the Project Viva birth cohort. In two epigenetic regions selected based on prior studies of air pollution effects on epigenome-wide methylation, we identify windows of susceptibility to PM$_{2.5}$ exposure near the beginning and middle of the third trimester of pregnancy.

stat.AP

Bayesian Wavelet-packet Historical Functional Linear Models

Historical Functional Linear Models (HFLM) quantify associations between a functional predictor and functional outcome where the predictor is an exposure variable that occurs before, or at least concurrently with, the outcome. Current work on the HFLM is largely limited to frequentist estimation techniques that employ spline-based basis representations. In this work, we propose a novel use of the discrete wavelet-packet transformation, which has not previously been used in functional models, to estimate historical relationships in a fully Bayesian model. Since inference has not been an emphasis of the existing work on HFLMs, we also employ two established Bayesian inference procedures in this historical functional setting. We investigate the operating characteristics of our wavelet-packet HFLM, as well as the two inference procedures, in simulation and use the model to analyze data on the impact of lagged exposure to particulate matter finer than 2.5$μ$g on heart rate variability in a cohort of journeyman boilermakers over the course of a day's shift.

stat.ME