SearcharxivSearch

arXiv subjects

Emily C. Hector

Publications and source records attributed to Emily C. Hector.

At least 19 recordsLinked to original sources

How useful is a wrong model? Information-sharing for inference under mean misspecification in linear models

As almost all models are wrong, a mean model's usefulness is often accepted as sufficient justification for its use. In practice, however, standard statistical theory breaks when the mean model fit is imperfect, limiting this usefulness. This tension between fit and usefulness arises from a dichotomization of model fit: the model is either right or it is wrong. Motivated by the linear regression framework, we propose an alternative viewpoint that leverages the mean model's usefulness without assuming it is right or wrong. We define a new model usefulness index and use it to share information across individual observations through the mean model. The result is an estimator of each outcome's mean that shrinks individualized means towards the shared mean model, with the degree of shrinkage governed by this usefulness index. We draw connections between our estimator and the James-Stein estimator and establish when and how our estimators of the individualized means yield more efficient inference than model-based and non-model-based alternatives. We also propose a data-dependent estimate of the usefulness index that balances statistical intuition with efficiency considerations. We illustrate our method's practical value in an analysis of personal tracker data.

stat.ME

Prior- and likelihood-free probabilistic inference with finite-sample calibration guarantees

Motivated by parametric models for which the likelihood is analytically unavailable, numerically unstable, or prohibitively expensive to compute or optimize, we develop a prior- and likelihood-free framework for fully probabilistic (Bayesian-like) uncertainty quantification with finite-sample calibration guarantees. Our method, a type of inferential model, produces data-dependent degrees of belief about claims concerning the unknown parameter while controlling the frequency with which high belief is assigned to false claims, even in finite-sample settings. Our procedure is general in that it requires only the ability to simulate from the model. We first rank candidate parameter values according to how well data simulated from the model agree with the observed data, and then rescale these rankings in a way that yields the desired finite-sample calibration guarantees. The key idea is to employ a permutation-invariant function, such as a depth function, to rank parameter values. We show that such a choice yields closed-form calibration rescaling calculations, making the procedure computationally simple. We illustrate our method's broad appeal with four examples, including differential privacy and Ising models. An analysis of the spatial configuration of 2025 measles outbreaks in the U.S. showcases our method's practical advantages.

stat.ME

Redefining shared information: a heterogeneity-adaptive framework for meta-analysis

Meta-analytic methods tend to take all-or-nothing approaches to study-level heterogeneity, assuming all studies are heterogeneous or homogeneous, leading to inefficiency and/or bias in estimation and inference. In this paper, we develop a heterogeneity-adaptive meta-analysis in linear models that adapts to the amount of information shared between datasets. The primary mechanism for the information-sharing is a shrinkage of dataset-specific distributions towards a new "centroid" distribution through a Kullback-Leibler divergence penalty. The Kullback-Leibler divergence is uniquely geometrically suited for measuring relative information between datasets, and leads to relatively simple closed form estimators with intuitive interpretations. We establish our estimator's desirable inferential properties without assuming homogeneity of dataset parameters. Among other results, we show that our estimator has a provably smaller mean squared error than the dataset-specific maximum likelihood estimators, and establish asymptotically valid inference procedures. A comprehensive set of simulations highlights our estimator's versatility, and an analysis of data from the eICU Collaborative Research Database illustrates its performance in a real-world setting.

stat.ME

A new mixture model for spatiotemporal exceedances with flexible tail dependence

We propose a new model and estimation framework for spatiotemporal streamflow exceedances above a threshold that flexibly captures asymptotic dependence and independence in the tail of the distribution. We model streamflow using a mixture of processes with spatial, temporal and spatiotemporal asymptotic dependence regimes. A censoring mechanism allows us to use only observations above a threshold to estimate marginal and joint probabilities of extreme events. As the likelihood is intractable, we use simulation-based inference powered by random forests to estimate model parameters from summary statistics of the data. Simulations and modeling of streamflow data from the U.S. Geological Survey illustrate the feasibility and practicality of our approach.

stat.ME

Estimating Covariate Effects on Functional Connectivity using Voxel-Level fMRI Data

Functional connectivity (FC) analysis of resting-state fMRI data provides a framework for characterizing brain networks and their association with participant-level covariates. Due to the high dimensionality of neuroimaging data, standard approaches often average signals within regions of interest (ROIs), which ignores the underlying spatiotemporal dependence among voxels and can lead to biased or inefficient inference. We propose to use a summary statistic -- the empirical voxel-wise correlations between ROIs -- and, crucially, model the complex covariance structure among these correlations through a new positive definite covariance function. Building on this foundation, we develop a computationally efficient two-step estimation procedure that enables statistical inference on covariate effects on region-level connectivity. Simulation studies show calibrated uncertainty quantification, and substantial gains in validity of the statistical inference over the standard averaging method. With data from the Autism Brain Imaging Data Exchange, we show that autism spectrum disorder is associated with altered FC between attention-related ROIs after adjusting for age and gender. The proposed framework offers an interpretable and statistically rigorous approach to estimation of covariate effects on FC suitable for large-scale neuroimaging studies.

stat.ME

Divide-and-conquer with finite sample sizes: valid and efficient possibilistic inference

Divide-and-conquer methods use large-sample approximations to provide frequentist guarantees when each block of data is both small enough to facilitate efficient computation and large enough to support approximately valid inferences. When the overall sample size is small or moderate, likely no suitable division of the data meets both requirements, hence the resulting inference lacks validity guarantees. We propose a new approach, couched in the inferential model framework, that is fully conditional in a Bayesian sense and provably valid in a frequentist sense. The main insight is that existing divide-and-conquer approaches make use of a Gaussianity assumption twice: first in the construction of an estimator, and second in the approximation to its sampling distribution. Our proposal is to retain the first Gaussianity assumption, using a Gaussian working likelihood, but to replace the second with a validification step that uses the sampling distributions of the block summaries determined by the posited model. This latter step, a type of probability-to-possibility transform, is key to the reliability guarantees enjoyed by our approach, which are uniquely general in the divide-and-conquer literature. In addition to finite-sample validity guarantees, our proposed approach is also asymptotically efficient like the other divide-and-conquer solutions available in the literature. Our computational strategy leverages state-of-the-art black-box likelihood emulators. We demonstrate our method's performance via simulations and highlight its flexibility with an analysis of median PM2.5 in Maryborough, Queensland, during the 2023 Australian bushfire season.

stat.ME

A new block covariance regression model and inferential framework for massively large neuroimaging data

Some evidence suggests that people with autism spectrum disorder exhibit patterns of brain functional dysconnectivity relative to their typically developing peers, but specific findings have yet to be replicated. To facilitate this replication goal with data from the Autism Brain Imaging Data Exchange (ABIDE), we propose a flexible and interpretable model for participant-specific voxel-level brain functional connectivity. Our approach efficiently handles massive participant-specific whole brain voxel-level connectivity data that exceed one trillion data points. The key component of the model is to leverage the block structure induced by defined regions of interest to introduce parsimony in the high-dimensional connectivity matrix through a block covariance structure. Associations between brain functional connectivity and participant characteristics -- including eye status during the resting scan, sex, age, and their interactions -- are estimated within a Bayesian framework. A spike-and-slab prior facilitates hypothesis testing to identify voxels associated with autism diagnosis. Simulation studies are conducted to evaluate the empirical performance of the proposed model and estimation framework. In ABIDE, the method replicates key findings from the literature and suggests new associations for investigation.

stat.ME

When the whole is greater than the sum of its parts: Scaling black-box inference to large data settings through divide-and-conquer

Black-box methods such as deep neural networks are exceptionally fast at obtaining point estimates of model parameters due to their amortisation of the loss function computation, but are currently restricted to settings for which simulating training data is inexpensive. When simulating data is computationally expensive, both the training and uncertainty quantification, which typically relies on a parametric bootstrap, become intractable. We propose a black-box divide-and-conquer estimation and inference framework when data simulation is computationally expensive that trains a black-box estimation method on a partition of the multivariate data domain, estimates and bootstraps on the partitioned data, and combines estimates and inferences across data partitions. Through the divide step, only small training data need be simulated, substantially accelerating the training. Further, the estimation and bootstrapping can be conducted in parallel across multiple computing nodes to further speed up the procedure. Finally, the conquer step accounts for any dependence between data partitions through a statistically and computationally efficient weighted average. We illustrate the implementation of our framework in high-dimensional spatial settings with Gaussian and max-stable processes. Applications to modeling extremal temperature data from both a climate model and observations from the National Oceanic and Atmospheric Administration highlight the feasibility of estimation and inference of max-stable process parameters with tens of thousands of locations.

stat.ME

Density correction for multivariate spatial fields of global climate model output using deep learning

Global Climate Models (GCMs) are numerical models that simulate complex physical processes within the Earth's climate system and are essential for understanding and predicting climate change. However, GCMs suffer from systemic biases due to simplifications made to the underlying physical processes. GCM output therefore needs to be bias corrected before it can be used for future climate projections. Most common bias correction methods, however, cannot preserve spatial, temporal, or inter-variable dependencies. We propose a new semi-parametric estimation of conditional densities (SPECD) approach for density correction of the joint distribution of daily precipitation and maximum temperature data obtained from gridded GCM spatial fields. The Vecchia approximation is employed to preserve dependencies in the observed field during the density correction process, which is carried out using semi-parametric quantile regression. The ability to calibrate joint distributions of GCM projections has potential advantages not only in estimating extremes, but also in better estimating compound hazards, like heat waves and drought, under potential climate change. Illustration on historical data from 1951-2014 over two 5 x 5 spatial grids in the US indicate that SPECD can preserve key marginal and joint distribution properties of precipitation and maximum temperature, and predictions obtained using SPECD are better calibrated compared to predictions using asynchronous quantile mapping and canonical correlation analysis, two commonly used bias correction approaches.

stat.AP

Multivariate and Online Transfer Learning with Uncertainty Quantification

Untreated periodontitis causes inflammation within the supporting tissue of the teeth and can ultimately lead to tooth loss. Modeling periodontal outcomes is beneficial as they are difficult and time consuming to measure, but disparities in representation between demographic groups must be considered. There may not be enough participants to build group specific models and it can be ineffective, and even dangerous, to apply a model to participants in an underrepresented group if demographic differences were not considered during training. We propose an extension to RECaST Bayesian transfer learning framework. Our method jointly models multivariate outcomes, exhibiting significant improvement over the previous univariate RECaST method. Further, we introduce an online approach to model sequential data sets. Negative transfer is mitigated to ensure that the information shared from the other demographic groups does not negatively impact the modeling of the underrepresented participants. The Bayesian framework naturally provides uncertainty quantification on predictions. Especially important in medical applications, our method does not share data between domains. We demonstrate the effectiveness of our method in both predictive performance and uncertainty quantification on simulated data and on a database of dental records from the HealthPartners Institute.

stat.ME

Demonstrating the power and flexibility of variational assumptions for amortized neural posterior estimation in environmental applications

Classic Bayesian methods with complex models are frequently infeasible due to an intractable likelihood. Simulation-based inference methods, such as Approximate Bayesian Computing (ABC), calculate posteriors without accessing a likelihood function by leveraging the fact that data can be quickly simulated from the model, but converge slowly and/or poorly in high-dimensional settings. In this paper, we propose a framework for Bayesian posterior estimation by mapping data to posteriors of parameters using a neural network trained on data simulated from the complex model. Posterior distributions of model parameters are efficiently obtained by feeding observed data into the trained neural network. We show theoretically that our posteriors converge to the true posteriors in Kullback-Leibler divergence. Our approach yields computationally efficient and theoretically justified uncertainty quantification, which is lacking in existing simulation-based neural network approaches. Comprehensive simulation studies highlight our method's robustness and accuracy.

stat.CO

Turning the information-sharing dial: efficient inference from different data sources

A fundamental aspect of statistics is the integration of data from different sources. Classically, Fisher and others were focused on how to integrate homogeneous (or only mildly heterogeneous) sets of data. More recently, as data are becoming more accessible, the question of if data sets from different sources should be integrated is becoming more relevant. The current literature treats this as a question with only two answers: integrate or don't. Here we take a different approach, motivated by information-sharing principles coming from the shrinkage estimation literature. In particular, we deviate from the do/don't perspective and propose a dial parameter that controls the extent to which two data sources are integrated. How far this dial parameter should be turned is shown to depend, for example, on the informativeness of the different data sources as measured by Fisher information. In the context of generalized linear models, this more nuanced data integration framework leads to relatively simple parameter estimates and valid tests/confidence intervals. Moreover, we demonstrate both theoretically and empirically that setting the dial parameter according to our recommendation leads to more efficient estimation compared to other binary data integration schemes.

stat.ME

Distributed model building and recursive integration for big spatial data modeling

Motivated by the need for computationally tractable spatial methods in neuroimaging studies, we develop a distributed and integrated framework for estimation and inference of Gaussian process model parameters with ultra-high-dimensional likelihoods. We propose a shift in viewpoint from whole to local data perspectives that is rooted in distributed model building and integrated estimation and inference. The framework's backbone is a computationally and statistically efficient integration procedure that simultaneously incorporates dependence within and between spatial resolutions in a recursively partitioned spatial domain. Statistical and computational properties of our distributed approach are investigated theoretically and in simulations. The proposed approach is used to extract new insights on autism spectrum disorder from the Autism Brain Imaging Data Exchange.

stat.ME

Bayesian estimation of clustered dependence structures in functional neuroconnectivity

Motivated by the need to model the dependence between regions of interest in functional neuroconnectivity for efficient inference, we propose a new sampling-based Bayesian clustering approach for covariance structures of high-dimensional Gaussian outcomes. The key technique is based on a Dirichlet process that clusters covariance sub-matrices into independent groups of outcomes, thereby naturally inducing sparsity in the whole brain connectivity matrix. A new split-merge algorithm is employed to achieve convergence of the Markov chain that is shown empirically to recover both uniform and Dirichlet partitions with high accuracy. We investigate the empirical performance of the proposed method through extensive simulations. Finally, the proposed approach is used to group regions of interest into functionally independent groups in the Autism Brain Imaging Data Exchange participants with autism spectrum disorder and and co-occurring attention-deficit/hyperactivity disorder.

stat.ME

Functional Regression with Intensively Measured Longitudinal Outcomes: A New Lens through Data Partitioning

Estimation and inference with modern longitudinal data from wearable devices, which consist of biological signals at high-frequency time points, is burdened by massive computational costs. We propose a distributed estimation and inference procedure that efficiently estimates both functional and scalar parameters with intensively measured longitudinal outcomes. The procedure overcomes computational difficulties through a scalable divide-and-conquer algorithm that partitions the outcomes into smaller sets. We circumvent traditional basis selection problems by analyzing data using quadratic inference functions in smaller subsets such that the basis functions have a low dimension. To address the challenges of combining estimates from dependent subsets, we propose a statistically efficient one-step estimator derived from a constrained generalized method of moments objective function with a smoothing penalty. We show theoretically and numerically that the proposed estimator is as statistically efficient as non-distributed alternative approaches and more efficient computationally. We demonstrate the practicality of our approach with the analysis of accelerometer data from the National Health and Nutrition Examination Survey.

stat.ME

A statistical framework for GWAS of high dimensional phenotypes using summary statistics, with application to metabolite GWAS

The recent explosion of genetic and high dimensional biobank and 'omic' data has provided researchers with the opportunity to investigate the shared genetic origin (pleiotropy) of hundreds to thousands of related phenotypes. However, existing methods for multi-phenotype genome-wide association studies (GWAS) do not model pleiotropy, are only applicable to a small number of phenotypes, or provide no way to perform inference. To add further complication, raw genetic and phenotype data are rarely observed, meaning analyses must be performed on GWAS summary statistics whose statistical properties in high dimensions are poorly understood. We therefore developed a novel model, theoretical framework, and set of methods to perform Bayesian inference in GWAS of high dimensional phenotypes using summary statistics that explicitly model pleiotropy, beget fast computation, and facilitate the use of biologically informed priors. We demonstrate the utility of our procedure by applying it to metabolite GWAS, where we develop new nonparametric priors for genetic effects on metabolite levels that use known metabolic pathway information and foster interpretable inference at the pathway level.

stat.ME

Transfer Learning with Uncertainty Quantification: Random Effect Calibration of Source to Target (RECaST)

Transfer learning uses a data model, trained to make predictions or inferences on data from one population, to make reliable predictions or inferences on data from another population. Most existing transfer learning approaches are based on fine-tuning pre-trained neural network models, and fail to provide crucial uncertainty quantification. We develop a statistical framework for model predictions based on transfer learning, called RECaST. The primary mechanism is a Cauchy random effect that recalibrates a source model to a target population; we mathematically and empirically demonstrate the validity of our RECaST approach for transfer learning between linear models, in the sense that prediction sets will achieve their nominal stated coverage, and we numerically illustrate the method's robustness to asymptotic approximations for nonlinear models. Whereas many existing techniques are built on particular source models, RECaST is agnostic to the choice of source model. For example, our RECaST transfer learning approach can be applied to a continuous or discrete data model with linear or logistic regression, deep neural network architectures, etc. Furthermore, RECaST provides uncertainty quantification for predictions, which is mostly absent in the literature. We examine our method's performance in a simulation study and in an application to real hospital data.

stat.ME

Fused mean structure learning in data integration with dependence

Motivated by image-on-scalar regression with data aggregated across multiple sites, we consider a setting in which multiple independent studies each collect multiple dependent vector outcomes, with potential mean model parameter homogeneity between studies and outcome vectors. To determine the validity of jointly analyzing these data sources, we must learn which of these data sources share mean model parameters. We propose a new model fusion approach that delivers improved flexibility, statistical performance and computational speed over existing methods. Our proposed approach specifies a quadratic inference function within each data source and fuses mean model parameter vectors in their entirety based on a new formulation of a pairwise fusion penalty. We establish theoretical properties of our estimator and propose an asymptotically equivalent weighted oracle meta-estimator that is more computationally efficient. Simulations and application to the ABIDE neuroimaging consortium highlight the flexibility of the proposed approach. An R package is provided for ease of implementation.

stat.ME