Searcharxiv⌕ Search

arXiv subjects

David Bolin

Publications and source records attributed to David Bolin.

At least 55 records · Page 3Linked to original sources

Gaussian Whittle-Matérn fields on metric graphs

We define a new class of Gaussian processes on compact metric graphs such as street or river networks. The proposed models, the Whittle--Matérn fields, are defined via a fractional stochastic differential equation on the compact metric graph and are a natural extension of Gaussian fields with Matérn covariance functions on Euclidean domains to the non-Euclidean metric graph setting. Existence of the processes, as well as some of their main properties, such as sample path regularity are derived. The model class in particular contains differentiable processes. To the best of our knowledge, this is the first construction of a differentiable Gaussian process on general compact metric graphs. Further, we prove an intrinsic property of these processes: that they do not change upon addition or removal of vertices with degree two. Finally, we obtain Karhunen--Loève expansions of the processes, provide numerical experiments, and compare them to Gaussian processes with isotropic covariance functions.

math.ST↗

A diffusion-based spatio-temporal extension of Gaussian Matérn fields

Gaussian random fields with Matérn covariance functions are popular models in spatial statistics and machine learning. In this work, we develop a spatio-temporal extension of the Gaussian Matérn fields formulated as solutions to a stochastic partial differential equation. The spatially stationary subset of the models have marginal spatial Matérn covariances, and the model also extends to Whittle-Matérn fields on curved manifolds, and to more general non-stationary fields. In addition to the parameters of the spatial dependence (variance, smoothness, and practical correlation range) it additionally has parameters controlling the practical correlation range in time, the smoothness in time, and the type of non-separability of the spatio-temporal covariance. Through the separability parameter, the model also allows for separable covariance functions. We provide a sparse representation based on a finite element approximation, that is well suited for statistical inference and which is implemented in the R-INLA software. The flexibility of the model is illustrated in an application to spatio-temporal modeling of global temperature data.

stat.ME↗

Fitting latent non-Gaussian models using variational Bayes and Laplace approximations

Latent Gaussian models (LGMs) are perhaps the most commonly used class of models in statistical applications. Nevertheless, in areas ranging from longitudinal studies in biostatistics to geostatistics, it is easy to find datasets that contain inherently non-Gaussian features, such as sudden jumps or spikes, that adversely affect the inferences and predictions made from an LGM. These datasets require more general latent non-Gaussian models (LnGMs) that can handle these non-Gaussian features automatically. However, fast implementation and easy-to-use software are lacking, which prevent LnGMs from becoming widely applicable. In this paper, we derive variational Bayes algorithms for fast and scalable inference of LnGMs. The approximation leads to an LGM that downweights extreme events in the latent process, reducing their impact and leading to more robust inferences. It can be applied to a wide range of models, such as autoregressive processes for time series, simultaneous autoregressive models for areal data, and spatial Matérn models. To facilitate Bayesian inference, we introduce the ngvb package, where LGMs implemented in R-INLA can be easily extended to LnGMs by adding a single line of code.

stat.ME↗

Fast Bayesian estimation of brain activation with cortical surface fMRI data using EM

Task functional magnetic resonance imaging (fMRI) is a type of neuroimaging data used to identify areas of the brain that activate during specific tasks or stimuli. These data are conventionally modeled using a massive univariate approach across all data locations, which ignores spatial dependence at the cost of model power. We previously developed and validated a spatial Bayesian model leveraging dependencies along the cortical surface of the brain in order to improve accuracy and power. This model utilizes stochastic partial differential equation spatial priors with sparse precision matrices to allow for appropriate modeling of spatially-dependent activations seen in the neuroimaging literature, resulting in substantial increases in model power. Our original implementation relies on the computational efficiencies of the integrated nested Laplace approximation (INLA) to overcome the computational challenges of analyzing high-dimensional fMRI data while avoiding issues associated with variational Bayes implementations. However, this requires significant memory resources, extra software, and software licenses to run. In this article, we develop an exact Bayesian analysis method for the general linear model, employing an efficient expectation-maximization algorithm to find maximum a posteriori estimates of task-based regressors on cortical surface fMRI data. Through an extensive simulation study of cortical surface-based fMRI data, we compare our proposed method to the existing INLA implementation, as well as a conventional massive univariate approach employing ad-hoc spatial smoothing. We also apply the method to task fMRI data from the Human Connectome Project and show that our proposed implementation produces similar results to the validated INLA implementation. Both the INLA and EM-based implementations are available through our open-source BayesfMRI R package.

stat.ME↗

Controlling the flexibility of non-Gaussian processes through shrinkage priors

The normal inverse Gaussian (NIG) and generalized asymmetric Laplace (GAL) distributions can be seen as skewed and semi-heavy-tailed extensions of the Gaussian distribution. Models driven by these more flexible noise distributions are then regarded as flexible extensions of simpler Gaussian models. Inferential procedures tend to overestimate the degree of non-Gaussianity in the data and therefore we propose controlling the flexibility of these non-Gaussian models by adding sensible priors in the inferential framework that contract the model towards Gaussianity. In our venture to derive sensible priors, we also propose a new intuitive parameterization of the non-Gaussian models and discuss how to implement them efficiently in $Stan$. The methods are derived for a generic class of non-Gaussian models that include spatial Matérn fields, autoregressive models for time series, and simultaneous autoregressive models for aerial data. The results are illustrated with a simulation study and geostatistics application, where priors that penalize model complexity were shown to lead to more robust estimation and give preference to the Gaussian model, while at the same time allowing for non-Gaussianity if there is sufficient evidence in the data.

stat.ME↗

Local scale invariance and robustness of proper scoring rules

Averages of proper scoring rules are often used to rank probabilistic forecasts. In many cases, the individual terms in these averages are based on observations and forecasts from different distributions. We show that some of the most popular proper scoring rules, such as the continuous ranked probability score (CRPS), give more importance to observations with large uncertainty which can lead to unintuitive rankings. To describe this issue, we define the concept of local scale invariance for scoring rules. A new class of generalized proper kernel scoring rules is derived and as a member of this class we propose the scaled CRPS (SCRPS). This new proper scoring rule is locally scale invariant and therefore works in the case of varying uncertainty. Like CRPS it is computationally available for output from ensemble forecasts, and does not require the ability to evaluate densities of forecasts. We further define robustness of scoring rules, show why this also is an important concept for average scores, and derive new proper scoring rules that are robust against outliers. The theoretical findings are illustrated in three different applications from spatial statistics, stochastic volatility models, and regression for count data.

math.ST↗

Fast Bayesian estimation of brain activation with cortical surface and subcortical fMRI data using EM

Analysis of brain imaging scans is critical to understanding the way the human brain functions, which can be leveraged to treat injuries and conditions that affect the quality of life for a significant portion of the human population. In particular, functional magnetic resonance imaging (fMRI) scans give detailed data on a living subject at high spatial and temporal resolutions. Due to the high cost involved in the collection of these scans, robust methods of analysis are of critical importance in order to produce meaningful inference. Bayesian methods in particular allow for the inclusion of expected behavior from prior study into an analysis, increasing the power of the results while circumventing problems that arise in classical analyses, including the effects of smoothing results and sensitivity to multiple comparison testing corrections. Recent development of a surface-based spatial Bayesian general linear model for cortical surface fMRI (cs-fMRI) data provides the desired power increase in task fMRI data using stochastic partial differential equation (SPDE) priors. This model relies on the computational efficiencies of the integrated nested Laplace approximation (INLA) to perform powerful analyses that have been validated to outperform classical analyses. In this article, we develop an exact Bayesian analysis method for the GLM, employing an expectation-maximization (EM) algorithm to find maximum a posteriori (MAP) estimates of task-based regressors on cs-fMRI and subcortical fMRI data while using minimal computational resources. Our proposed method is compared to the INLA implementation of the Bayesian GLM, as well as a classical GLM on simulated data. A validation of the method on data from the Human Connectome Project is also provided.

stat.ME↗

Equivalence of measures and asymptotically optimal linear prediction for Gaussian random fields with fractional-order covariance operators

We consider Gaussian measures $μ, \tildeμ$ on a separable Hilbert space, with fractional-order covariance operators $A^{-2β}$ resp. $\tilde{A}^{-2\tildeβ}$, and derive necessary and sufficient conditions on $A, \tilde{A}$ and $β, \tildeβ > 0$ for I. equivalence of the measures $μ$ and $\tildeμ$, and II. uniform asymptotic optimality of linear predictions for $μ$ based on the misspecified measure $\tildeμ$. These results hold, e.g., for Gaussian processes on compact metric spaces. As an important special case, we consider the class of generalized Whittle-Matérn Gaussian random fields, where $A$ and $\tilde{A}$ are elliptic second-order differential operators, formulated on a bounded Euclidean domain $\mathcal{D}\subset\mathbb{R}^d$ and augmented with homogeneous Dirichlet boundary conditions. Our outcomes explain why the predictive performances of stationary and non-stationary models in spatial statistics often are comparable, and provide a crucial first step in deriving consistency results for parameter estimation of generalized Whittle-Matérn fields.

math.PR↗

Necessary and sufficient conditions for asymptotically optimal linear prediction of random fields on compact metric spaces

Optimal linear prediction (aka. kriging) of a random field $\{Z(x)\}_{x\in\mathcal{X}}$ indexed by a compact metric space $(\mathcal{X},d_{\mathcal{X}})$ can be obtained if the mean value function $m\colon\mathcal{X}\to\mathbb{R}$ and the covariance function $\varrho\colon\mathcal{X}\times\mathcal{X}\to\mathbb{R}$ of $Z$ are known. We consider the problem of predicting the value of $Z(x^*)$ at some location $x^*\in\mathcal{X}$ based on observations at locations $\{x_j\}_{j=1}^n$ which accumulate at $x^*$ as $n\to\infty$ (or, more generally, predicting $φ(Z)$ based on $\{φ_j(Z)\}_{j=1}^n$ for linear functionals $φ,φ_1,\ldots,φ_n$). Our main result characterizes the asymptotic performance of linear predictors (as $n$ increases) based on an incorrect second order structure $(\tilde{m},\tilde{\varrho})$, without any restrictive assumptions on $\varrho,\tilde{\varrho}$ such as stationarity. We, for the first time, provide necessary and sufficient conditions on $(\tilde{m},\tilde{\varrho})$ for asymptotic optimality of the corresponding linear predictor holding uniformly with respect to $φ$. These general results are illustrated by weakly stationary random fields on $\mathcal{X}\subset\mathbb{R}^d$ with Matérn or periodic covariance functions, and on the sphere $\mathcal{X}=\mathbb{S}^2$ for the case of two isotropic covariance functions.

math.ST↗

The SPDE approach for Gaussian and non-Gaussian fields: 10 years and still running

Gaussian processes and random fields have a long history, covering multiple approaches to representing spatial and spatio-temporal dependence structures, such as covariance functions, spectral representations, reproducing kernel Hilbert spaces, and graph based models. This article describes how the stochastic partial differential equation approach to generalising Matérn covariance models via Hilbert space projections connects with several of these approaches, with each connection being useful in different situations. In addition to an overview of the main ideas, some important extensions, theory, applications, and other recent developments are discussed. The methods include both Markovian and non-Markovian models, non-Gaussian random fields, non-stationary fields and space-time fields on arbitrary manifolds, and practical computational considerations.

stat.ME↗

Spatial Bayesian GLM on the cortical surface produces reliable task activations in individuals and groups

The general linear model (GLM) is a widely popular and convenient tool for estimating the functional brain response and identifying areas of significant activation during a task or stimulus. However, the classical GLM is based on a massive univariate approach that does not explicitly leverage the similarity of activation patterns among neighboring brain locations. As a result, it tends to produce noisy estimates and be underpowered to detect significant activations, particularly in individual subjects and small groups. A recent alternative, a cortical surface-based spatial Bayesian GLM, leverages spatial dependencies among neighboring cortical vertices to produce more accurate estimates and areas of functional activation. The spatial Bayesian GLM can be applied to individual and group-level analysis. In this study, we assess the reliability and power of individual and group-average measures of task activation produced via the surface-based spatial Bayesian GLM. We analyze motor task data from 45 subjects in the Human Connectome Project (HCP) and HCP Retest datasets. We also extend the model to multi-run analysis and employ subject-specific cortical surfaces rather than surfaces inflated to a sphere for more accurate distance-based modeling. Results show that the surface-based spatial Bayesian GLM produces highly reliable activations in individual subjects and is powerful enough to detect trait-like functional topologies. Additionally, spatial Bayesian modeling enhances reliability of group-level analysis even in moderately sized samples (n=45). The power of the spatial Bayesian GLM to detect activations above a scientifically meaningful effect size is nearly invariant to sample size, exhibiting high power even in small samples (n=10). The spatial Bayesian GLM is computationally efficient in individuals and groups and is convenient to implement with the open-source BayesfMRI R package.

stat.ME↗

Efficient methods for Gaussian Markov random fields under sparse linear constraints

Methods for inference and simulation of linearly constrained Gaussian Markov Random Fields (GMRF) are computationally prohibitive when the number of constraints is large. In some cases, such as for intrinsic GMRFs, they may even be unfeasible. We propose a new class of methods to overcome these challenges in the common case of sparse constraints, where one has a large number of constraints and each only involves a few elements. Our methods rely on a basis transformation into blocks of constrained versus non-constrained subspaces, and we show that the methods greatly outperform existing alternatives in terms of computational cost. By combining the proposed methods with the stochastic partial differential equation approach for Gaussian random fields, we also show how to formulate Gaussian process regression with linear constraints in a GMRF setting to reduce computational cost. This is illustrated in two applications with simulated data.

stat.ME↗

Spatial 3D Matérn priors for fast whole-brain fMRI analysis

Bayesian whole-brain functional magnetic resonance imaging (fMRI) analysis with three-dimensional spatial smoothing priors has been shown to produce state-of-the-art activity maps without pre-smoothing the data. The proposed inference algorithms are computationally demanding however, and the proposed spatial priors have several less appealing properties, such as being improper and having infinite spatial range. We propose a statistical inference framework for whole-brain fMRI analysis based on the class of Matérn covariance functions. The framework uses the Gaussian Markov random field (GMRF) representation of possibly anisotropic spatial Matérn fields via the stochastic partial differential equation (SPDE) approach of Lindgren et al. (2011). This allows for more flexible and interpretable spatial priors, while maintaining the sparsity required for fast inference in the high-dimensional whole-brain setting. We develop an accelerated stochastic gradient descent (SGD) optimization algorithm for empirical Bayes (EB) inference of the spatial hyperparameters. Conditionally on the inferred hyperparameters, we make a fully Bayesian treatment of the brain activity. The Matérn prior is applied to both simulated and experimental task-fMRI data and clearly demonstrates that it is a more reasonable choice than the previously used priors, using comparisons of activity maps, prior simulation and cross-validation.

stat.ME↗

Deformed SPDE models with an application to spatial modeling of significant wave height

A non-stationary Gaussian random field model is developed based on a combination of the stochastic partial differential equation (SPDE) approach and the classical deformation method. With the deformation method, a stationary field is defined on a domain which is deformed so that the field becomes non-stationary. We show that if the stationary field is a Mat'ern field defined as a solution to a fractional SPDE, the resulting non-stationary model can be represented as the solution to another fractional SPDE on the deformed domain. By defining the model in this way, the computational advantages of the SPDE approach can be combined with the deformation method's more intuitive parameterisation of non-stationarity. In particular it allows for independent control over the non-stationary practical correlation range and the variance, which has not been possible with previously proposed non-stationary SPDE models. The model is tested on spatial data of significant wave height, a characteristic of ocean surface conditions which is important when estimating the wear and risks associated with a planned journey of a ship. The model parameters are estimated to data from the north Atlantic using a maximum likelihood approach. The fitted model is used to compute wave height exceedance probabilities and the distribution of accumulated fatigue damage for ships traveling a popular shipping route. The model results agree well with the data, indicating that the model could be used for route optimization in naval logistics.

stat.AP↗

A spatial template independent component analysis model for subject-level brain network estimation and inference

Independent component analysis is commonly applied to functional magnetic resonance imaging (fMRI) data to extract independent components (ICs) representing functional brain networks. While ICA produces reliable group-level estimates, single-subject ICA often produces noisy results. Template ICA (tICA) is a hierarchical ICA model using empirical population priors to produce reliable subject-level IC estimates. However, this and other hierarchical ICA models assume unrealistically that subject effects are spatially independent. Here, we propose spatial template ICA (stICA), which incorporates spatial process priors into tICA. This results in greater estimation efficiency of ICs and subject effects. Additionally, the joint posterior distribution can be used to identify engaged areas using an excursions set approach. By leveraging spatial dependencies and avoiding massive multiple comparisons, stICA has high power to detect true effects. We derive an efficient expectation-maximization algorithm to obtain maximum likelihood estimates of the model parameters and posterior moments of the latent fields. Based on analysis of simulated data and fMRI data from the Human Connectome Project, we find that stICA produces estimates that are more accurate and reliable than benchmark approaches, and identifies larger and more reliable areas of engagement. The algorithm is quite tractable, achieving convergence within 7 hours in our fMRI analysis.

stat.ME↗

The COST IRACON Geometry-based Stochastic Channel Model for Vehicle-to-Vehicle Communication in Intersections

Vehicle-to-vehicle (V2V) wireless communications can improve traffic safety at road intersections and enable congestion avoidance. However, detailed knowledge about the wireless propagation channel is needed for the development and realistic assessment of V2V communication systems. We present a novel geometry-based stochastic MIMO channel model with support for frequencies in the band of 5.2-6.2 GHz. The model is based on extensive high-resolution measurements at different road intersections in the city of Berlin, Germany. We extend existing models, by including the effects of various obstructions, higher order interactions, and by introducing an angular gain function for the scatterers. Scatterer locations have been identified and mapped to measured multi-path trajectories using a measurement-based ray tracing method and a subsequent RANSAC algorithm. The developed model is parameterized, and using the measured propagation paths that have been mapped to scatterer locations, model parameters are estimated. The time variant power fading of individual multi-path components is found to be best modeled by a Gamma process with an exponential autocorrelation. The path coherence distance is estimated to be in the range of 0-2 m. The model is also validated against measurement data, showing that the developed model accurately captures the behavior of the measured channel gain, Doppler spread, and delay spread. This is also the case for intersections that have not been used when estimating model parameters.

eess.SP↗

Multivariate type G Matérn stochastic partial differential equation random fields

For many applications with multivariate data, random field models capturing departures from Gaussianity within realisations are appropriate. For this reason, we formulate a new class of multivariate non-Gaussian models based on systems of stochastic partial differential equations with additive type G noise whose marginal covariance functions are of Matérn type. We consider four increasingly flexible constructions of the noise, where the first two are similar to existing copula-based models. In contrast to these, the latter two constructions can model non-Gaussian spatial data without replicates. Computationally efficient methods for likelihood-based parameter estimation and probabilistic prediction are proposed, and the flexibility of the suggested models is illustrated by numerical examples and two statistical applications.

stat.ME↗

The rational SPDE approach for Gaussian random fields with general smoothness

A popular approach for modeling and inference in spatial statistics is to represent Gaussian random fields as solutions to stochastic partial differential equations (SPDEs) of the form $L^βu = \mathcal{W}$, where $\mathcal{W}$ is Gaussian white noise, $L$ is a second-order differential operator, and $β>0$ is a parameter that determines the smoothness of $u$. However, this approach has been limited to the case $2β\in\mathbb{N}$, which excludes several important models and makes it necessary to keep $β$ fixed during inference. We propose a new method, the rational SPDE approach, which in spatial dimension $d\in\mathbb{N}$ is applicable for any $β>d/4$, and thus remedies the mentioned limitation. The presented scheme combines a finite element discretization with a rational approximation of the function $x^{-β}$ to approximate $u$. For the resulting approximation, an explicit rate of convergence to $u$ in mean-square sense is derived. Furthermore, we show that our method has the same computational benefits as in the restricted case $2β\in\mathbb{N}$. Several numerical experiments and a statistical application are used to illustrate the accuracy of the method, and to show that it facilitates likelihood-based inference for all model parameters including $β$.

stat.ME↗