SearcharxivSearch

arXiv subjects

Ofer Harel

Publications and source records attributed to Ofer Harel.

5 recordsLinked to original sources

Multiple Heckman Selection Model

We introduce a novel matrix-variate extension of the Heckman selection model to accommodate multiple outcomes, providing a flexible and natural generalization of classical selection models for matrix-valued data. By relying on the matrix normal distribution, the proposed model captures dependencies across both rows and columns while accounting for selection bias. An Expectation/Conditional Maximization (ECM) algorithm is developed, yielding closed-form updates for all model parameters. We investigate key theoretical properties, including the connection between sample selection models and the recently developed multivariate unified skew-normal (SUN) distribution. The performance of the proposed approach is assessed through simulation studies, and its practical utility is illustrated using two real datasets. The proposed method is implemented in the R package mvHeckman.

stat.ME

Bayesian Sparsity Modeling of Shared Neural Response in Functional Magnetic Resonance Imaging Data

Detecting shared neural activity from functional magnetic resonance imaging (fMRI) across individuals exposed to the same stimulus can reveal synchronous brain responses, functional roles of regions, and potential clinical biomarkers. Intersubject correlation (ISC) is the main method for identifying voxelwise shared responses and per-subject variability, but it relies on heavy data summarization and thousands of regional tests, leading to poor uncertainty quantification and multiple testing issues. ISC also does not directly estimate a shared neural response (SNR) function. We propose a model-based alternative applicable to both task-based and naturalistic fMRI that simultaneously identifies spatial regions of shared activity and estimates the SNR function. The model combines sparse Gaussian process estimation of the response function with a Bayesian sparsity prior inspired by the horseshoe prior to detect voxel activation. A spatially structured extension encourages neighboring voxels to exhibit similar activation patterns. We examine the model's properties, evaluate performance via simulations, and analyze two real-world fMRI datasets, including one task-based and one naturalistic dataset. The Bayesian framework provides principled uncertainty quantification for the shared response function and shows improved activation detection and response estimation compared to standard approaches. Model fits demonstrate comparable or superior performance relative to ISC, while the framework opens avenues for clinical applications.

stat.AP

Multiple imputation in functional regression with applications to EEG data in a depression study

Current source density (CSD) power asymmetry, a measure derived from electroencephalography (EEG), is a potential biomarker for major depressive disorder (MDD). Though this measure is functional in nature (defined on the frequency domain), it is typically reduced to a scalar value prior to analysis, possibly obscuring the relationship between brain function and MDD. To overcome this issue, we sought to fit a functional regression model to estimate the association between CSD power asymmetry and MDD diagnostic status, adjusting for age, sex, cognitive ability, and handedness using data from a large clinical study. Unfortunately, nearly 40\% of the observations were missing either their functional EEG data, their cognitive ability score, or both. In order to take advantage of all of the available data, we propose an extension to multiple imputation by chained equations that handles both scalar and functional data. We also propose an extension to Rubin's Rules for pooling estimates from the multiply imputed data sets in order to conduct valid inference. We investigate the performance of the proposed extensions in a simulation study and apply them to our clinical study data. Our analysis reveals that the association between CSD power asymmetry and diagnostic status depends on both age and sex.

stat.AP

Confidence intervals for the area under the receiver operating characteristic curve in the presence of ignorable missing data

Receiver operating characteristic (ROC) curves are widely used as a measure of accuracy of diagnostic tests and can be summarized using the area under the ROC curve (AUC). Often, it is useful to construct a confidence intervals for the AUC, however, since there are a number of different proposed methods to measure variance of the AUC, there are thus many different resulting methods for constructing these intervals. In this manuscript, we compare different methods of constructing Wald-type confidence interval in the presence of missing data where the missingness mechanism is ignorable. We find that constructing confidence intervals using multiple imputation (MI) based on logistic regression (LR) gives the most robust coverage probability and the choice of CI method is less important. However, when missingness rate is less severe (e.g. less than 70%), we recommend using Newcombe's Wald method for constructing confidence intervals along with multiple imputation using predictive mean matching (PMM).

stat.AP

Addressing missing data mechanism uncertainty using multiple-model multiple imputation: Application to a longitudinal clinical trial

We present a framework for generating multiple imputations for continuous data when the missing data mechanism is unknown. Imputations are generated from more than one imputation model in order to incorporate uncertainty regarding the missing data mechanism. Parameter estimates based on the different imputation models are combined using rules for nested multiple imputation. Through the use of simulation, we investigate the impact of missing data mechanism uncertainty on post-imputation inferences and show that incorporating this uncertainty can increase the coverage of parameter estimates. We apply our method to a longitudinal clinical trial of low-income women with depression where nonignorably missing data were a concern. We show that different assumptions regarding the missing data mechanism can have a substantial impact on inferences. Our method provides a simple approach for formalizing subjective notions regarding nonresponse so that they can be easily stated, communicated and compared.

stat.AP