SearcharxivSearch

arXiv subjects

Samuel Davenport

Publications and source records attributed to Samuel Davenport.

15 recordsLinked to original sources

Regularity Conditions for Critical Point Convergence

We focus on a sequence of functions $\{f_n\}$, defined on a compact manifold with boundary $S$, converging in the $C^k$ metric to a limit $f$. A common assumption implicitly made in the empirical sciences is that when such functions represent random processes derived from data, the topological features of $f_n$ will eventually resemble those of $f$. In this work, we investigate the validity of this claim under various regularity assumptions, with the goal of finding conditions sufficient for the number of local maxima, minima and saddle of such functions to converge. In the $C^1$ setting, we do so by employing lesser-known variants of the Poincaré-Hopf and mountain pass theorems, and in the $C^2$ setting we pursue an approach inspired by the homotopy-based proof of the Morse Lemma. To aid practical use, we end by reformulating our central theorems in the language of the empirical processes.

math.GN

False Discovery Controlled Regions for Excursion Sets

Identifying areas where the signal is prominent is an important task in image analysis, with particular applications in brain mapping. In this work, we develop False Discovery Controlled Regions (FDCRs) for excursion sets, the set of locations where the values are above or below a given level. We achieve this by treating the confidence procedure as a testing problem at the given level, allowing control of the False Discovery Rate (FDR). Methods are developed to control the FDR, separately for positive and negative excursions, as well as jointly over both. Furthermore, power is increased by incorporating a two-stage adaptive procedure. Simulation results in a variety of different settings show that the FDCRs successfully control the FDR under the nominal alpha level. We showcase our methods with an application to functional magnetic resonance imaging (fMRI) data from the Human Connectome Project illustrating the improvement in statistical power over existing approaches.

stat.ME

Consistency of heritability estimation from summary statistics in high-dimensional linear models

In Genome-Wide Association Studies (GWAS), heritability is defined as the fraction of variance of an outcome explained by a large number of genetic predictors in a high-dimensional polygenic linear model. This work studies the asymptotic properties of the most common estimator of heritability from summary statistics called linkage disequilibrium score (LDSC) regression, together with a simpler and closely related estimator called GWAS heritability (GWASH). These estimators are analyzed in their basic versions and under various modifications used in practice including weighting and standardization. We show that, with some variations, two conditions which we call weak dependence (WD) and bounded-kurtosis effects (BKE) are sufficient for consistency of both the basic LDSC with fixed intercept and GWASH estimators, for both Gaussian and non-Gaussian predictors. For Gaussian predictors it is shown that these conditions are also necessary for consistency of GWASH (with truncation) and simulations suggest that necessity holds too when the predictors are non-Gaussian. We also show that, with properly truncated weights, weighting does not change the consistency results, but standardization of the predictors and outcome, as done in practice, introduces bias in both LDSC and GWASH if the two essential conditions are violated. Finally, we show that, when population stratification is present, all the estimators considered are biased, and the bias is not remedied by using the LDSC regression estimator with free intercept, as originally suggested by the authors of that estimator.

math.ST

On the peak height distribution of non-stationary Gaussian random fields: 1D general covariance and scale space

We study the peak height distribution of certain non-stationary Gaussian random fields. The explicit peak height distribution of smooth, non-stationary Gaussian processes in 1D with general covariance is derived. The formula is determined by two parameters, each of which has a clear statistical meaning. For multidimensional non-stationary Gaussian random fields, we generalize these results to the setting of scale space fields, which play an important role in peak detection by helping to handle peaks of different spatial extents. We demonstrate that these properties not only offer a better interpretation of the scale space field but also simplify the computation of the peak height distribution. Finally, two efficient numerical algorithms are proposed as a general solution for computing the peak height distribution of smooth multidimensional Gaussian random fields in applications.

stat.ME

Peak Inference for Gaussian Random Fields on a Lattice

In this work we develop a Monte Carlo method to compute the height distribution of local maxima of a stationary Gaussian or Gaussian-related random field that is observed on a regular lattice. We show that our method can be used to provide valid peak based inference in datasets with low levels of smoothness, where existing formulae derived for continuous domains are not accurate. We also extend the methods in Worsley (2005) and Taylor et al. (2007) to compute the peak height distribution and compare them with our approach. Lastly, we apply our method to a task fMRI dataset to show how it can be used in practice.

stat.ME

Conformal confidence sets for biomedical image segmentation

We develop confidence sets which provide spatial uncertainty guarantees for the output of a black-box machine learning model designed for image segmentation. To do so we adapt conformal inference to the imaging setting, obtaining thresholds on a calibration dataset based on the distribution of the maximum of the transformed logit scores within and outside of the ground truth masks. We prove that these confidence sets, when applied to new predictions of the model, are guaranteed to contain the true unknown segmented mask with desired probability. We show that learning appropriate score transformations on a learning dataset before performing calibration is crucial for optimizing performance. We illustrate and validate our approach on a polpys tumor dataset. To do so we obtain the logit scores from a deep neural network trained for polpys segmentation and show that using distance transformed scores to obtain outer confidence sets and the original scores for inner confidence sets enables tight bounds on tumor location whilst controlling the false coverage rate.

stat.ML

Permutation-based multiple testing when fitting many generalized linear models

In many applied sciences a popular analysis strategy for high-dimensional data is to fit many multivariate generalized linear models in parallel. This paper presents a novel approach to address the resulting multiple testing problem by combining a recently developed sign-flip test with permutation-based multiple-testing procedures. Our method builds upon the univariate standardized flip-scores test which offers robustness against misspecified variances in generalized linear models, a crucial feature in high-dimensional settings where comprehensive model validation is particularly challenging. We extend this approach to the multivariate setting, enabling adaptation to unknown response correlation structures. This approach yields relevant power improvements over conventional multiple testing methods when correlation is present.

math.ST

Inference in generalized linear models with robustness to misspecified variances

Generalized linear models usually assume a common dispersion parameter, an assumption that is seldom true in practice. Consequently, standard parametric methods may suffer appreciable loss of type I error control. As an alternative, we present a semi-parametric group-invariance method based on sign flipping of score contributions. Our method requires only the correct specification of the mean model, but is robust against any misspecification of the variance. We present tests for single as well as multiple regression coefficients. The test is asymptotically valid but shows excellent performance in small samples. We illustrate the method using RNA sequencing count data, for which it is difficult to model the overdispersion correctly. The method is available in the R library flipscores.

stat.ME

Object detection under the linear subspace model with application to cryo-EM images

Detecting multiple unknown objects in noisy data is a key problem in many scientific fields, such as electron microscopy imaging. A common model for the unknown objects is the linear subspace model, which assumes that the objects can be expanded in some known basis (such as the Fourier basis). In this paper, we develop an object detection algorithm that under the linear subspace model is asymptotically guaranteed to detect all objects, while controlling the family wise error rate or the false discovery rate. Numerical simulations show that the algorithm also controls the error rate with high power in the non-asymptotic regime, even in highly challenging regimes. We apply the proposed algorithm to experimental electron microscopy data set, and show that it outperforms existing standard software.

math.ST

Precise FWER Control for Gaussian Related Fields: Riding the SuRF to continuous land -- Part 1

The Gaussian Kinematic Formula (GKF) is a powerful and computationally efficient tool to perform statistical inference on random fields and became a well-established tool in the analysis of neuroimaging data. Using realistic error models, recent articles show that GKF based methods for \emph{voxelwise inference} lead to conservative control of the familywise error rate (FWER) and for cluster-size inference lead to inflated false positive rates. In this series of articles we identify and resolve the main causes of these shortcomings in the traditional usage of the GKF for voxelwise inference. This first part removes the \textit{good lattice assumption} and allows the data to be non-stationary, yet still assumes the data to be Gaussian. The latter assumption is resolved in part 2, where we also demonstrate that our GKF based methodology is non-conservative under realistic error models.

stat.ME

Robust FWER control in Neuroimaging using Random Field Theory: Riding the SuRF to Continuous Land Part 2

Historically, applications of RFT in fMRI have relied on assumptions of smoothness, stationarity and Gaussianity. The first two assumptions have been addressed in Part 1 of this article series. Here we address the severe non-Gaussianity of (real) fMRI data to greatly improve the performance of voxelwise RFT in fMRI group analysis. In particular, we introduce a transformation which accelerates the convergence of the Central Limit Theorem allowing us to rely on limiting Gaussianity of the test-statistic. We shall show that, when the GKF is combined with the Gaussianization transformation, we are able to accurately estimate the EEC of the excursion set of the transformed test-statistic even when the data is non-Gaussian. This allows us to drop the key assumptions of RFT inference and enables us to provide a fast approach which correctly controls the voxelwise false positive rate in fMRI. We employ a big data \cite{Eklund2016} style validation in which we process resting state data from 7000 subjects from the UK BioBank with fake task designs. We resample from this data to create realistic noise and use this to demonstrate that the error rate is correctly controlled.

stat.AP

FDP control in multivariate linear models using the bootstrap

In this article we develop a method for performing post hoc inference of the False Discovery Proportion (FDP) over multiple contrasts of interest in the multivariate linear model. To do so we use the bootstrap to simulate from the distribution of the null contrasts. We combine the bootstrap with the post hoc inference bounds of Blanchard (2020) and prove that doing so provides simultaneous asymptotic control of the FDP over all subsets of hypotheses. This requires us to demonstrate consistency of the multivariate bootstrap in the linear model, which we do via the Lindeberg Central Limit Theorem, providing a simpler proof of this result than that of Eck (2018). We demonstrate, via simulations, that our approach provides simultaneous control of the FDP over all subsets and is typically more powerful than existing, state of the art, parametric methods. We illustrate our approach on functional Magnetic Resonance Imaging data from the Human Connectome project and on a transcriptomic dataset of chronic obstructive pulmonary disease.

stat.ME

Confidence regions for the location of peaks of a smooth random field

Local maxima of random processes are useful for finding important regions and are routinely used, for summarising features of interest (e.g. in neuroimaging). In this work we provide confidence regions for the location of local maxima of the mean and standardized effect size (i.e. Cohen's d) given multiple realisations of a random process. We prove central limit theorems for the location of the maximum of mean and t-statistic random fields and use these to provide asymptotic confidence regions for the location of peaks of the mean and Cohen's d. Under the assumption of stationarity we develop Monte Carlo confidence regions for the location of peaks of the mean that have better finite sample coverage than regions derived based on classical asymptotic normality. We illustrate our methods on 1D MEG data and 2D fMRI data from the UK Biobank.

math.ST

Functional delta residuals and applications to simultaneous confidence bands of moment based statistics

Given a functional central limit (fCLT) for an estimator and a parameter transformation, we construct random processes, called functional delta residuals, which asymptotically have the same covariance structure as the limit process of the functional delta method. An explicit construction of these residuals for transformations of moment-based estimators and a multiplier bootstrap fCLT for the resulting functional delta residuals are proven. The latter is used to consistently estimate the quantiles of the maximum of the limit process of the functional delta method in order to construct asymptotically valid simultaneous confidence bands for the transformed functional parameters. Performance of the coverage rate of the developed construction, applied to functional versions of Cohen's d, skewness and kurtosis, is illustrated in simulations and their application to test Gaussianity is discussed.

math.ST

On the finiteness of the second moment of the number of critical points of Gaussian random fields

We prove that the second moment of the number of critical points of any sufficiently regular random field, for example with almost surely $ C^3 $ sample paths, defined over a compact Whitney stratified manifold is finite. Our results hold without the assumption of stationarity - which has traditionally been assumed in other work. Under stationarity we demonstrate that our imposed conditions imply the generalized Geman condition of Estrade 2016.

math.PR