SearcharxivSearch

arXiv subjects

Julia Wrobel

Publications and source records attributed to Julia Wrobel.

13 recordsLinked to original sources

Adaptive spatial blocking for scalable clustering inference with applications to high-throughput spatial proteomics

Ripley's K-function is a widely used spatial summary statistic for assessing clustering in point patterns. However, existing K-based methods can be computationally prohibitive for large-scale data, particularly in high-throughput spatial proteomics, because they rely on spatial information from all points in the image. To address this challenge, we propose a computationally efficient block-based testing framework that extracts disjoint local blocks from an image and aggregates clustering evidence across them. The proposed adaptive spatial blocking algorithm constructs blocks satisfying point-count and shape constraints, enabling scalable spatial clustering inference and fast p-value computation through an asymptotic normal approximation. Numerical studies demonstrate that the proposed method provides a favorable balance between statistical power and computational efficiency. In an application to healthy human intestine spatial proteomics data, our method detects strong spatial aggregation of plasma cells and colocalization between plasma cells and macrophages, while scaling favorably to large images.

stat.ME

Overcoming Barriers to Computational Reproducibility

Computational reproducibility, the possibility for independent researchers to exactly reproduce published empirical results, is fundamental to science. Despite its importance, the proportion of research articles aiming for reproducibility remains low and uneven across disciplines. Barriers include a perceived lack of incentives for researchers and journals, practical challenges in preparing reproducible materials, and the absence of harmonised standards of reproducibility processes and requirements by journals. Existing guidance is often highly technical, reaching mainly those already engaged with reproducible research. In this paper, we first synthesize evidence on the benefits of reproducibility for both authors and journals. Drawing on our extensive experience in reproducibility checking at various journals, we then put forward concise, pragmatic guidelines for creating reproducible analyses across disciplines. We further review current reproducibility policies of selected journals, illustrating the substantial heterogeneity in requirements and procedures. Motivated by the latter, we propose conceptual foundations for a harmonised multi-tier system of reproducibility standards that could support transparent, consistent assessment across journals and research communities. Our goal as journal (reproducibility) editors and contributors to the MaRDI initiative is to encourage broader adoption of reproducibility practices, in particular by lowering practical barriers for authors and journals.

cs.DL

SCoRES: An R Package for Simultaneous Confidence Region Estimates

The identification of domain sets whose outcomes belong to predefined subsets can address fundamental risk assessment challenges in climatology and medicine. Existing approaches for inverse domain estimates require restrictive assumptions, including domain density and continuity of function near thresholds, and large-sample guarantees, which limit the applicability. Besides, the estimation and coverage depend on setting a fixed threshold level, which is difficult to determine. Recently, Ren et al. (2024) proved that confidence sets of multiple levels can be simultaneously constructed with the desired confidence non-asymptotically through inverting simultaneous confidence bands. Here, we present the SCoRES R package, which implements Ren's approach for both the estimation of the inverse region and the corresponding simultaneous outer and inner confidence regions, along with visualization tools. Besides, the package also provides functions that help construct SCBs for regression data, functional data and geographical data. To illustrate its broad applicability, we present three rigorous examples that demonstrate the SCoRES workflow.

stat.CO

Functional Accelerated Failure Time Models for Predicting Time Since Cannabis Use

Cannabis consumption impairs key driving skills and increases crash risk, yet few objective, validated tools exists to identify acute cannabis use or impairment in traffic safety settings. Pupil response to light has emerged as a promising biomarker of recent cannabis use, but its predictive utility remains underexplored. We propose two functional accelerated failure time (AFT) models for predicting time since cannabis use from pupil light response curves. The linear functional AFT (lfAFT) model provides a simple and interpretable framework that summarizes the overall contribution of a functional covariate to time-since-smoking, while the additive functional AFT (afAFT) model generalizes this structure by allowing effects to vary flexibly with both magnitude and location of the functional covariate. Estimation is computationally efficient and straightforward to implement. Simulation studies show that the proposed methods achieve strong estimation accuracy and predictive performance across various scenarios and remain robust to moderate model misspecification. Application to pupillometry data from the Colorado Cannabis & Driving Study demonstrates that pupil light response curves contain meaningful predictive signal, underscoring the potential of these models for traffic safety and broader biomedical applications.

stat.ME

Identifying Neural Signatures from fMRI using Hybrid Principal Components Regression

Recent advances in neuroimaging analysis have enabled accurate decoding of mental state from brain activation patterns during functional magnetic resonance imaging scans. A commonly applied tool for this purpose is principal components regression regularized with the least absolute shrinkage and selection operator (LASSO PCR), a type of multi-voxel pattern analysis (MVPA). This model presumes that all components are equally likely to harbor relevant information, when in fact the task-related signal may be concentrated in specific components. In such cases, the model will fail to select the optimal set of principal components that maximizes the total signal relevant to the cognitive process under study. Here, we present modifications to LASSO PCR that allow for a regularization penalty tied directly to the index of the principal component, reflecting a prior belief that task-relevant signal is more likely to be concentrated in components explaining greater variance. Additionally, we propose a novel hybrid method, Joint Sparsity-Ranked LASSO (JSRL), which integrates component-level and voxel-level activity under an information parity framework and imposes ranked sparsity to guide component selection. We apply the models to brain activation during risk taking, monetary incentive, and emotion regulation tasks. Results demonstrate that incorporating sparsity ranking into LASSO PCR produces models with enhanced classification performance, with JSRL achieving up to 51.7\% improvement in cross-validated deviance $R^2$ and 7.3\% improvement in cross-validated AUC. Furthermore, sparsity-ranked models perform as well as or better than standard LASSO PCR approaches across all classification tasks and allocate predictive weight to brain regions consistent with their established functional roles, offering a robust alternative for MVPA.

stat.ML

Generalized Multilevel Functional Principal Component Analysis with Application to NHANES Active Inactive Patterns

Between 2011 and 2014 NHANES collected objectively measured physical activity data using wrist-worn accelerometers for tens of thousands of individuals for up to seven days. In this study, we analyze minute-level indicators of being active, which can be viewed as binary (since each minute is either active or inactive), multilevel (because there are multiple days of data for each participant), and functional data (because the within-day measurements can be viewed as a function of time). To identify both within- and between-participant directions of variation in these data, we introduce Generalized Multilevel Functional Principal Component Analysis (GM-FPCA), an approach based on the dimension reduction of the linear predictor. Our results indicate that specific activity patterns captured by GM-FPCA are strongly associated with mortality risk. Extensive simulation studies demonstrate that GM-FPCA accurately estimates model parameters, is computationally stable, and scales up with the number of study participants, visits, and observations per visit. R code for implementing the method is provided.

stat.ME

A robust, scalable K-statistic for quantifying immune cell clustering in spatial proteomics data

Spatial summary statistics based on point process theory are widely used to quantify the spatial organization of cell populations in single-cell spatial proteomics data. Among these, Ripley's K is a popular metric for assessing whether cells are spatially clustered or are randomly dispersed. However, the key assumption of spatial homogeneity is frequently violated in spatial proteomics data, leading to overestimates of cell clustering and colocalization. To address this, we propose a novel method, termed KAMP (K adjustment by Analytical Moments of the Permutation distribution), for quantifying the spatial organization of cells in spatial proteomics samples. KAMP leverages background cells in each sample along with a new closed-form representation of the first and second moments of the permutation null distribution of Ripley's K. Our method is robust to inhomogeneity, computationally efficient even in large datasets, and provides approximate p-values to test spatial clustering and colocalization. Methodological developments are motivated by a spatial proteomics study of women with ovarian cancer; in the subset with sufficient B cells and macrophages, KAMP provides exploratory, scale-specific evidence linking B cell-macrophage colocalization with overall patient survival. Notably, we also find evidence that using K without correcting for sample inhomogeneity may bias hazard ratio estimates in downstream analyses.

stat.ME

Generalized Conditional Functional Principal Component Analysis

We propose generalized conditional functional principal components analysis (GC-FPCA) for the joint modeling of the fixed and random effects of non-Gaussian functional outcomes. The method scales up to very large functional data sets by estimating the principal components of the covariance matrix on the linear predictor scale conditional on the fixed effects. This is achieved by combining three modeling innovations: (1) fit local generalized linear mixed models (GLMMs) conditional on covariates in windows along the functional domain; (2) conduct a functional principal component analysis (FPCA) on the person-specific functional effects obtained by assembling the estimated random effects from the local GLMMs; and (3) fit a joint functional mixed effects model conditional on covariates and the estimated principal components from the previous step. GC-FPCA was motivated by modeling the minute-level active/inactive profiles over the day ($1{,}440$ 0/1 measurements per person) for $8{,}700$ study participants in the National Health and Nutrition Examination Survey (NHANES) 2011-2014. We show that state-of-the-art approaches cannot handle data of this size and complexity, while GC-FPCA can.

stat.ME

Modeling trajectories using functional linear differential equations

We are motivated by a study that seeks to better understand the dynamic relationship between muscle activation and paw position during locomotion. For each gait cycle in this experiment, activation in the biceps and triceps is measured continuously and in parallel with paw position as a mouse trotted on a treadmill. We propose an innovative general regression method that draws from both ordinary differential equations and functional data analysis to model the relationship between these functional inputs and responses as a dynamical system that evolves over time. Specifically, our model addresses gaps in both literatures and borrows strength across curves estimating ODE parameters across all curves simultaneously rather than separately modeling each functional observation. Our approach compares favorably to related functional data methods in simulations and in cross-validated predictive accuracy of paw position in the gait data. In the analysis of the gait cycles, we find that paw speed and position are dynamically influenced by inputs from the biceps and triceps muscles, and that the effect of muscle activation persists beyond the activation itself.

stat.ME

Mistaken identities lead to missed opportunities: Testing for mean differences in partially matched data

It is increasingly common to collect pre-post data with pseudonyms or self-constructed identifiers. On survey responses from sensitive populations, identifiers may be made optional to encourage higher response rates. The ability to match responses between pre- and post-intervention phases for every participant may be impossible in such applications, leaving practitioners with a choice between the paired t-test on the matched samples and the two-sample t-test on all samples for evaluating mean differences. We demonstrate the inadequacies with both approaches, as the former test requires discarding unmatched data, while the latter test ignores correlation and assumes independence. In cases with a subset of matched samples, an opportunity to achieve limited inference about the correlation exists. We propose a novel technique for such `partially matched' data, which we refer to as the Quantile-based t-test for correlated samples, to assess mean differences using a conservative estimate of the correlation between responses based on the matched subset. Critically, our approach does not discard unmatched samples, nor does it assume independence. Our results demonstrate that the proposed method yields nominal Type I error probability while affording more power than existing approaches. Practitioners can readily adopt our approach with basic statistical programming software.

stat.ME

Analysing spatial point patterns in digital pathology: immune cells in high-grade serous ovarian carcinomas

Multiplex immunofluorescence (mIF) imaging technology facilitates the study of the tumour microenvironment in cancer patients. Due to the capabilities of this emerging bioimaging technique, it is possible to statistically analyse, for example, the co-varying location and functions of multiple different types of immune cells. Complex spatial relationships between different immune cells have been shown to correlate with patient outcomes and may reveal new pathways for targeted immunotherapy treatments. This tutorial reviews methods and procedures relating to spatial point patterns for complex data analysis. We consider tissue cells as a realisation of a spatial point process for each patient. We focus on proper functional descriptors for each observation and techniques that allow us to obtain information about inter-patient variation. Ovarian cancer is the deadliest gynaecological malignancy and can resist chemotherapy treatment effective in cancers. We use a dataset of high-grade serous ovarian cancer samples from 51 patients. We examine the immune cell composition (T cells, B cells, macrophages) within tumours and additional information such as cell classification (tumour or stroma) and other patient clinical characteristics. Our analyses, supported by reproducible software, apply to other digital pathology datasets.

stat.AP

Fast Generalized Functional Principal Components Analysis

We propose a new fast generalized functional principal components analysis (fast-GFPCA) algorithm for dimension reduction of non-Gaussian functional data. The method consists of: (1) binning the data within the functional domain; (2) fitting local random intercept generalized linear mixed models in every bin to obtain the initial estimates of the person-specific functional linear predictors; (3) using fast functional principal component analysis to smooth the linear predictors and obtain their eigenfunctions; and (4) estimating the global model conditional on the eigenfunctions of the linear predictors. An extensive simulation study shows that fast-GFPCA performs as well or better than existing state-of-the-art approaches, it is orders of magnitude faster than existing general purpose GFPCA methods, and scales up well with both the number of observed curves and observations per curve. Methods were motivated by and applied to a study of active/inactive physical activity profiles obtained from wearable accelerometers in the NHANES 2011-2014 study. The method can be implemented by any user familiar with mixed model software, though the R package fastGFPCA is provided for convenience.

stat.ME

Interactive graphics for functional data analyses

Although there are established graphics that accompany the most common functional data analyses, generating these graphics for each dataset and analysis can be cumbersome and time consuming. Often, the barriers to visualization inhibit useful exploratory data analyses and prevent the development of intuition for a method and its application to a particular dataset. The refund.shiny package was developed to address these issues for several of the most common functional data analyses. After conducting an analysis, the plot_shiny() function is used to generate an interactive visualization environment that contains several distinct graphics, many of which are updated in response to user input. These visualizations reduce the burden of exploratory analyses and can serve as a useful tool for the communication of results to non-statisticians.

stat.OT