SearcharxivSearch

arXiv subjects

Thomas E. Nichols

Publications and source records attributed to Thomas E. Nichols.

At least 19 recordsLinked to original sources

Task-guided cross-subject latent alignment: a multi-encoder-decoder VAE

Aligning neural activity across subjects offers the promise of discovering shared computational principles and generalizable decoders. However, traditional alignment methods require shared stimuli across subjects, a constraint that limits applicability to naturalistic paradigms with limited or non-overlapping data. We introduce a Multi-Encoder-Decoder Variational Autoencoder (MED-VAE) that achieves cross-subject alignment without shared stimuli by anchoring representations to a common scaffold provided by a pretrained ANN. Using the Natural Scenes Dataset, we show that MED-VAE creates common latent spaces with superior semantic organisation, achieving higher cross-subject alignment than common methods while maintaining robust generalisation to held-out stimuli where traditional methods degrade. Reconstructing from these common spaces back to each subject's original neural space, MED-VAE preserves equal stimulus-driven signal in its cross-subject latent space. Finally, we show that this superior alignment directly enables cross-subject neural prediction, as demonstrated via cross-subject image decoding. In summary, we introduce a framework to identify generalisable common subspaces for cross-subject predictions and downstream tasks, demonstrated here for visual cortex responses to static images.

q-bio.NC

eTFCE: Exact Threshold-Free Cluster Enhancement via Fast Cluster Retrieval

Threshold-free cluster enhancement (TFCE) is widely used for cluster-based inference in neuroimaging, but existing implementations typically rely on discretized approximations that may introduce numerical variability. We present eTFCE, an efficient framework that provides a numerically exact evaluation of the TFCE integral using an optimized cluster retrieval algorithm. Across multiple datasets, eTFCE and the standard implementation produce highly consistent inference results. Voxel-wise comparisons reveal a systematic asymmetry: the standard method yields smaller p-values for more voxels, while eTFCE concentrates stronger statistical evidence within a smaller subset. These differences are primarily confined to voxels near the inference boundary and have minimal impact on overall inference. This pattern is consistent with discretization effects in standard implementations, where the TFCE integral is approximated using a finite set of threshold levels, introducing subtle biases in statistical evidence accumulation across thresholds. Furthermore, eTFCE improves computational efficiency (71.3% of runtime on average) and enables unified computation of multiple cluster-based statistics within a single permutation framework. Overall, eTFCE provides an exact, efficient, and extensible approach to nonparametric neuroimaging inference.

stat.ME

Excessive data censoring in fMRI undermines individual precision and weakens brain-behavior associations

Censoring high-motion volumes in fMRI is common practice to reduce effects of head motion on functional connectivity (FC). Although aggressive censoring removes more noise, it causes extensive data loss, creating a tradeoff that may ultimately improve or degrade FC accuracy. Here, we evaluate how censoring affects FC estimation and downstream brain-wide association studies (BWAS). Using extensively sampled participants from the Human Connectome Project (HCP) Retest dataset, we establish individual "ground truth" FC and assess the accuracy of FC estimated from 5-30 minute scans. We find that censoring degrades FC accuracy, with more aggressive censoring being more detrimental, particularly among participants exhibiting above-average motion. In these participants, aggressive censoring reduces FC accuracy by 30% for 30-minute scans denoised with ICA-FIX, an advanced denoising method, and by 3% for scans denoised with conventional confound regression. These effects reflect substantial data loss (34%) that outweighs comparatively modest noise reductions: 7% with ICA-FIX and 18% with confound regression. Compensating for this would require substantially longer scans (62% with confound regression; 76% with ICA-FIX), inflating data collection budgets. Introducing a repeated measures framework to separate motion trait from artifact, we find that standard QC metrics are dominated by motion trait and overstate motion bias, which is effectively mitigated with less aggressive censoring. Finally, using data from nearly 1,000 HCP participants, we demonstrate that unreliable FC substantially attenuates BWAS correlations: by ~30% under optimal conditions (longer ICA-FIX scans with no censoring) but exceeding 75% in short, aggressively censored scans. Our findings support the use of advanced denoising methods, limiting censoring, and collecting longer scans to maximize fidelity of FC and BWAS.

stat.AP

Scalable Bayesian Image-on-Scalar Regression for Population-Scale Neuroimaging Data Analysis

Bayesian Image-on-Scalar Regression (ISR) provides flexible, uncertainty-aware neuroimaging analysis. However, applying ISR to large-scale datasets such as the UK Biobank is challenging due to intensive computational demands and the need to handle subject-specific brain masks rather than a common mask. We propose a novel Bayesian ISR model that scales efficiently while accommodating these inconsistent masks. Our method leverages Gaussian process priors with salience area indicators and introduces a scalable posterior computation algorithm using stochastic gradient Langevin dynamics combined with memory mapping. This approach achieves linear scaling with subsample size and constrains memory usage to the batch size, facilitating direct spatial posterior inferences on brain activation regions. Simulation studies and analysis of UK Biobank task fMRI data (38,639 subjects; over 120,000 voxels per image) demonstrate a 4- to 11-fold speed increase and an 8-18% enhancement in statistical power compared to traditional Gibbs sampling with zero-imputation. Our analysis reveals a subregion of the amygdala where emotion-related brain activation decreases by approximately 58% between ages 50 and 60.

stat.AP

False Discovery Rate and Localizing Power

False discovery rate (FDR) is commonly used for correction for multiple testing in neuroimaging studies. However, when using two-tailed tests, making directional inferences about the results can lead to a vastly inflated error rate, even approaching 100% in some cases. This happens because FDR controls the error rate only globally, over all tests, not within subsets, such as among those in only one or another direction. Here we consider and evaluate different strategies for FDR control in such cases, using both synthetic and real imaging data. Approaches that separate the tests by direction of the hypothesis test, or by the direction of the resulting test statistic, more properly control the directional error rate and preserve FDR benefits, albeit with a doubled risk of errors under complete absence of signal. Strategies that combine tests in both directions, or that use simple two-tailed p-values, can lead to invalid directional conclusions, even if these tests remain globally valid. A solution to this problem is through the use of selective inference, whereby positive and negative tails are treated as sets (families), which are screened locally, then subjected to FDR at a modified level that controls average FDR over those that survive the initial screening. Moreover, the BKY procedure can be used in place of the well-known Benjamini-Hochberg, yielding additional power. These methods are easy to implement. Finally, to enable valid thresholding for directional inference, we suggest that imaging software should allow the user to set asymmetrical thresholds for the two sides of the statistical map. While FDR continues to be a valid, powerful procedure for multiple testing correction, care is needed when making directional inferences for two-tailed tests, or more broadly, when making any localized inference.

stat.ME

A hierarchical modelling approach for Bayesian Causal Forests on longitudinal data: A Case Study in Multiple Sclerosis Clinical Trials

Long-running clinical trials offer a unique opportunity to study disease progression and treatment response over time, enabling questions about how and when interventions alter patient trajectories. However, drawing causal conclusions in this setting is challenging due to irregular follow-up, individual-level heterogeneity, and time-varying confounding. Bayesian Additive Regression Trees (BART) and their extension, Bayesian Causal Forests (BCF), have proven powerful for flexible causal inference in observational data, especially for heterogeneous treatment effects and non-linear outcome surfaces. Yet, both models assume independence across observations and are fundamentally limited in their ability to model within-individual correlation over time. This limits their use in real-world longitudinal settings where repeated measures are the norm. Motivated by the NO.MS dataset, the largest and most comprehensive clinical trial dataset in Multiple Sclerosis (MS), with more than 35,000 patients and up to 15 years follow-up, we develop BCFLong, a hierarchical model that preserves BART's strengths while extending it for longitudinal analysis. Inspired by BCF, we decompose the mean into prognostic and treatment effects, modelling the former on Image Quality Metrics (IQMs) to account for scanner effects, and introduce individual-specific random effects, including intercepts and slope, with a sparsity-inducing horseshoe prior. Simulations confirm BCFLong's superior performance and robustness to sparsity, significantly improving outcome and treatment effect estimation. On NO.MS, BCFLong captures clinically meaningful longitudinal patterns in brain volume change, which would have otherwise remained undetected. These findings highlight the importance of adaptively accounting for within-individual correlations and position BCFLong as a flexible framework for causal inference in longitudinal data.

stat.AP

Normative brain mapping of 3-dimensional morphometry imaging data using skewed functional data analysis

Tensor-based morphometry (TBM) aims at showing local differences in brain volumes with respect to a common template. TBM images are smooth but they exhibit (especially in diseased groups) higher values in some brain regions called lateral ventricles. More specifically, our voxelwise analysis shows both a mean-variance relationship in these areas and evidence of spatially dependent skewness. We propose a model for 3-dimensional functional data where mean, variance, and skewness functions vary smoothly across brain locations. We model the voxelwise distributions as skew-normal. The smooth effects of age and sex are estimated on a reference population of cognitively normal subjects from the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset and mapped across the whole brain. The three parameter functions allow to transform each TBM image (in the reference population as well as in a test set) into a Gaussian process. These subject-specific normative maps are used to derive indices of deviation from a healthy condition to assess the individual risk of pathological degeneration.

stat.AP

Scalable Scalar-on-Image Cortical Surface Regression with a Relaxed-Thresholded Gaussian Process Prior

In addressing the challenge of analysing the large-scale Adolescent Brain Cognition Development (ABCD) fMRI dataset, involving over 5,000 subjects and extensive neuroimaging data, we propose a scalable Bayesian scalar-on-image regression model for computational feasibility and efficiency. Our model employs a relaxed-thresholded Gaussian process (RTGP), integrating piecewise-smooth, sparse, and continuous functions capable of both hard- and soft-thresholding. This approach introduces additional flexibility in feature selection in scalar-on-image regression and leads to scalable posterior computation by adopting a variational approximation and utilising the Karhunen-Loève expansion for Gaussian processes. This advancement substantially reduces the computational costs in vertex-wise analysis of cortical surface data in large-scale Bayesian spatial models. The model's parameter estimation and prediction accuracy and feature selection performance are validated through extensive simulation studies and an application to the ABCD study. Here, we perform regression analysis correlating intelligence scores with task-based functional MRI data, taking into account confounding factors including age, sex, and parental education level. This validation highlights our model's capability to handle large-scale neuroimaging data while maintaining computational feasibility and accuracy.

stat.ME

The Past, Present, and Future of the Brain Imaging Data Structure (BIDS)

The Brain Imaging Data Structure (BIDS) is a community-driven standard for the organization of data and metadata from a growing range of neuroscience modalities. This paper is meant as a history of how the standard has developed and grown over time. We outline the principles behind the project, the mechanisms by which it has been extended, and some of the challenges being addressed as it evolves. We also discuss the lessons learned through the project, with the aim of enabling researchers in other domains to learn from the success of BIDS.

q-bio.OT

Robust FWER control in Neuroimaging using Random Field Theory: Riding the SuRF to Continuous Land Part 2

Historically, applications of RFT in fMRI have relied on assumptions of smoothness, stationarity and Gaussianity. The first two assumptions have been addressed in Part 1 of this article series. Here we address the severe non-Gaussianity of (real) fMRI data to greatly improve the performance of voxelwise RFT in fMRI group analysis. In particular, we introduce a transformation which accelerates the convergence of the Central Limit Theorem allowing us to rely on limiting Gaussianity of the test-statistic. We shall show that, when the GKF is combined with the Gaussianization transformation, we are able to accurately estimate the EEC of the excursion set of the transformed test-statistic even when the data is non-Gaussian. This allows us to drop the key assumptions of RFT inference and enables us to provide a fast approach which correctly controls the voxelwise false positive rate in fMRI. We employ a big data \cite{Eklund2016} style validation in which we process resting state data from 7000 subjects from the UK BioBank with fake task designs. We resample from this data to create realistic noise and use this to demonstrate that the error rate is correctly controlled.

stat.AP

Bayesian Lesion Estimation with a Structured Spike-and-Slab Prior

Neural demyelination and brain damage accumulated in white matter appear as hyperintense areas on T2-weighted MRI scans in the form of lesions. Modeling binary images at the population level, where each voxel represents the existence of a lesion, plays an important role in understanding aging and inflammatory diseases. We propose a scalable hierarchical Bayesian spatial model, called BLESS, capable of handling binary responses by placing continuous spike-and-slab mixture priors on spatially-varying parameters and enforcing spatial dependency on the parameter dictating the amount of sparsity within the probability of inclusion. The use of mean-field variational inference with dynamic posterior exploration, which is an annealing-like strategy that improves optimization, allows our method to scale to large sample sizes. Our method also accounts for underestimation of posterior variance due to variational inference by providing an approximate posterior sampling approach based on Bayesian bootstrap ideas and spike-and-slab priors with random shrinkage targets. Besides accurate uncertainty quantification, this approach is capable of producing novel cluster size based imaging statistics, such as credible intervals of cluster size, and measures of reliability of cluster occurrence. Lastly, we validate our results via simulation studies and an application to the UK Biobank, a large-scale lesion mapping study with a sample size of 40,000 subjects.

stat.ME

Neuroimaging Meta Regression for Coordinate Based Meta Analysis Data with a Spatial Model

Coordinate-based meta-analysis combines evidence from a collection of Neuroimaging studies to estimate brain activation. In such analyses, a key practical challenge is to find a computationally efficient approach with good statistical interpretability to model the locations of activation foci. In this article, we propose a generative coordinate-based meta-regression (CBMR) framework to approximate smooth activation intensity function and investigate the effect of study-level covariates (e.g., year of publication, sample size). We employ spline parameterization to model spatial structure of brain activation and consider four stochastic models for modelling the random variation in foci. To examine the validity of CBMR, we estimate brain activation on $20$ meta-analytic datasets, conduct spatial homogeneity tests at voxel level, and compare to results generated by existing kernel-based approaches.

stat.ME

Neuroscience needs Network Science

The brain is a complex system comprising a myriad of interacting elements, posing significant challenges in understanding its structure, function, and dynamics. Network science has emerged as a powerful tool for studying such intricate systems, offering a framework for integrating multiscale data and complexity. Here, we discuss the application of network science in the study of the brain, addressing topics such as network models and metrics, the connectome, and the role of dynamics in neural networks. We explore the challenges and opportunities in integrating multiple data streams for understanding the neural transitions from development to healthy function to disease, and discuss the potential for collaboration between network science and neuroscience communities. We underscore the importance of fostering interdisciplinary opportunities through funding initiatives, workshops, and conferences, as well as supporting students and postdoctoral fellows with interests in both disciplines. By uniting the network science and neuroscience communities, we can develop novel network-based methods tailored to neural circuits, paving the way towards a deeper understanding of the brain and its functions.

q-bio.NC

Cluster extent inference revisited: quantification and localization of brain activity

Cluster inference based on spatial extent thresholding is the most popular analysis method for finding activated brain areas in neuroimaging. However, the method has several well-known issues. While powerful for finding brain regions with some activation, the method as currently defined does not allow any further quantification or localization of signal. In this paper we repair this gap. We show that cluster-extent inference can be used (1.) to infer the presence of signal in anatomical regions of interest and (2.) to quantify the percentage of active voxels in any cluster or region of interest. These additional inferences come for free, i.e. they do not require any further adjustment of the alpha-level of tests, while retaining full familywise error control. We achieve this extension of the possibilities of cluster inference by an embedding of the method into a closed testing procedure, and solving the graph-theoretic k-separator problem that results from this embedding. The new method can be used in combination with random field theory or permutations. We demonstrate the usefulness of the method in a large-scale application to neuroimaging data from the Neurovault database.

stat.ME

Confidence regions for the location of peaks of a smooth random field

Local maxima of random processes are useful for finding important regions and are routinely used, for summarising features of interest (e.g. in neuroimaging). In this work we provide confidence regions for the location of local maxima of the mean and standardized effect size (i.e. Cohen's d) given multiple realisations of a random process. We prove central limit theorems for the location of the maximum of mean and t-statistic random fields and use these to provide asymptotic confidence regions for the location of peaks of the mean and Cohen's d. Under the assumption of stationarity we develop Monte Carlo confidence regions for the location of peaks of the mean that have better finite sample coverage than regions derived based on classical asymptotic normality. We illustrate our methods on 1D MEG data and 2D fMRI data from the UK Biobank.

math.ST

Spatial Confidence Regions for Combinations of Excursion Sets in Image Analysis

The analysis of excursion sets in imaging data is essential to a wide range of scientific disciplines such as neuroimaging, climatology and cosmology. Despite growing literature, there is little published concerning the comparison of processes that have been sampled across the same spatial region but which reflect different study conditions. Given a set of asymptotically Gaussian random fields, each corresponding to a sample acquired for a different study condition, this work aims to provide confidence statements about the intersection, or union, of the excursion sets across all fields. Such spatial regions are of natural interest as they directly correspond to the questions "all random fields exceed a predetermined threshold?", or "Where does at least one random field exceed a predetermined threshold?". To assess the degree of spatial variability present, we develop a method that provides, with a desired confidence, subsets and supersets of spatial regions defined by logical conjunctions (i.e. set intersections) or disjunctions (i.e. set unions), without any assumption on the dependence between the different fields. The method is verified by extensive simulations and demonstrated using a task-fMRI dataset to identify brain regions with activation common to four variants of a working memory task.

stat.ME

Fisher Scoring for crossed factor Linear Mixed Models

The analysis of longitudinal, heterogeneous or unbalanced clustered data is of primary importance to a wide range of applications. The Linear Mixed Model (LMM) is a popular and flexible extension of the linear model specifically designed for such purposes. Historically, a large proportion of material published on the LMM concerns the application of popular numerical optimization algorithms, such as Newton-Raphson, Fisher Scoring and Expectation Maximization to single-factor LMMs (i.e. LMMs that only contain one "factor" by which observations are grouped). However, in recent years, the focus of the LMM literature has moved towards the development of estimation and inference methods for more complex, multi-factored designs. In this paper, we present and derive new expressions for the extension of an algorithm classically used for single-factor LMM parameter estimation, Fisher Scoring, to multiple, crossed-factor designs. Through simulation and real data examples, we compare five variants of the Fisher Scoring algorithm with one another, as well as against a baseline established by the R package lmer, and find evidence of correctness and strong computational efficiency for four of the five proposed approaches. Additionally, we provide a new method for LMM Satterthwaite degrees of freedom estimation based on analytical results, which does not require iterative gradient estimation. Via simulation, we find that this approach produces estimates with both lower bias and lower variance than the existing methods.

stat.ME

Permutation Inference for Canonical Correlation Analysis

Canonical correlation analysis (CCA) has become a key tool for population neuroimaging, allowing investigation of associations between many imaging and non-imaging measurements. As other variables are often a source of variability not of direct interest, previous work has used CCA on residuals from a model that removes these effects, then proceeded directly to permutation inference. We show that such a simple permutation test leads to inflated error rates. The reason is that residualisation introduces dependencies among the observations that violate the exchangeability assumption. Even in the absence of nuisance variables, however, a simple permutation test for CCA also leads to excess error rates for all canonical correlations other than the first. The reason is that a simple permutation scheme does not ignore the variability already explained by previous canonical variables. Here we propose solutions for both problems: in the case of nuisance variables, we show that transforming the residuals to a lower dimensional basis where exchangeability holds results in a valid permutation test; for more general cases, with or without nuisance variables, we propose estimating the canonical correlations in a stepwise manner, removing at each iteration the variance already explained, while dealing with different number of variables in both sides. We also discuss how to address the multiplicity of tests, proposing an admissible test that is not conservative, and provide a complete algorithm for permutation inference for CCA.

stat.ME