SearcharxivSearch

arXiv subjects

David Bolin

Publications and source records attributed to David Bolin.

At least 37 records · Page 2Linked to original sources

Leveraging population information in brain connectivity via Bayesian ICA with a novel informative prior for correlation matrices

Brain functional connectivity (FC), the temporal synchrony between brain networks, is essential to understand the functional organization of the brain and to identify changes due to neurological disorders, development, treatment, and other phenomena. Independent component analysis (ICA) is a matrix decomposition method used extensively for simultaneous estimation of functional brain topography and connectivity. However, estimation of FC via ICA is often sub-optimal due to the use of ad-hoc estimation methods or temporal dimension reduction prior to ICA. Bayesian ICA can avoid dimension reduction, estimate latent variables and model parameters more accurately, and facilitate posterior inference. In this paper, we develop a novel, computationally feasible Bayesian ICA method with population-derived priors on both the spatial ICs and their temporal correlation, i.e. FC. For the latter we consider two priors: the inverse-Wishart, which is conjugate but is not ideally suited for modeling correlation matrices; and a novel informative prior for correlation matrices. For each prior, we derive a variational Bayes algorithm to estimate the model variables and facilitate posterior inference. Through extensive simulation studies, we evaluate the performance of the proposed methods and benchmark against existing approaches. We also analyze fMRI data from over 400 healthy adults in the Human Connectome Project. We find that our Bayesian ICA model and algorithms result in more accurate measures of functional connectivity and spatial brain features. Our novel prior for correlation matrices is more computationally intensive than the inverse-Wishart but provides improved accuracy and inference. The proposed framework is applicable to single-subject analysis, making it potentially clinically viable.

stat.AP

Incorporating Correlated Nugget Effects in Multivariate Spatial Models: An Application to Argo Ocean Data

Accurate analysis of global oceanographic data, such as temperature and salinity profiles from the Argo program, requires geostatistical models capable of capturing complex spatial dependencies. This study introduces Gaussian and non-Gaussian hierarchical multivariate Matérn-SPDE models with correlated nugget effects to account for small-scale variability and measurement error correlations. Using simulations and Argo data, we demonstrate that incorporating correlated nugget effects significantly improves the accuracy of parameter estimation and spatial prediction in both Gaussian and non-Gaussian multivariate spatial processes. When applied to global ocean temperature and salinity data, our model yields lower correlation estimates between fields compared to models that assume independent noise. This suggests that traditional models may overestimate the underlying field correlation. By separating these effects, our approach captures fine-scale oceanic patterns more effectively. These findings show the importance of relaxing the assumption of independent measurement errors in multivariate hierarchical models.

stat.ME

Wasserstein complexity penalization priors: a new class of penalizing complexity priors

Penalizing complexity (PC) priors provide a principled framework for reducing model complexity by penalizing the Kullback--Leibler Divergence (KLD) between a ``simple'' base model and a more complex model. However, constructing priors by penalizing the KLD becomes impossible in many cases because the KLD is infinite, and alternative principles often lose interpretability in terms of KLD. We propose a new class of priors, the Wasserstein complexity penalization (WCP) priors, which replace the KLD with the Wasserstein distance in the PC prior framework. WCP priors avoid the issue of infinite model distances and retain interpretability by adhering to adjusted principles. Additionally, we introduce the concept of base measures, removing the parameter dependency on the base model, and extend the framework to joint WCP priors for multiple parameters. These priors can be constructed analytically and we have both analytical and numerical implementations in R programming language. We demonstrate their use in previous PC prior applications and as well as new multivariate settings.

stat.ME

Linear cost and exponentially convergent approximation of Gaussian Matérn processes on intervals

The computational cost for inference and prediction of statistical models based on Gaussian processes with Matérn covariance functions scales cubicly with the number of observations, limiting their applicability to large data sets. The cost can be reduced in certain special cases, but there are currently no generally applicable exact methods with linear cost. Several approximate methods have been introduced to reduce the cost, but most of these lack theoretical guarantees for the accuracy. We consider Gaussian processes on bounded intervals with Matérn covariance functions and for the first time develop a generally applicable method with linear cost and with a covariance error that decreases exponentially fast in the order $m$ of the proposed approximation. The method is based on an optimal rational approximation of the spectral density and results in an approximation that can be represented as a sum of $m$ independent Gaussian Markov processes, which facilitates easy usage in general software for statistical inference, enabling its efficient implementation in general statistical inference software packages. Besides the theoretical justifications, we demonstrate the accuracy empirically through carefully designed simulation studies which show that the method outperforms all state-of-the-art alternatives in terms of accuracy for a fixed computational cost in statistical tasks such as Gaussian process regression.

math.ST

rSPDE: tools for statistical modeling using fractional SPDEs

The R software package rSPDE contains methods for approximating Gaussian random fields based on fractional-order stochastic partial differential equations (SPDEs). A common example of such fields are Whittle-Matérn fields on bounded domains in $\mathbb{R}^d$, manifolds, or metric graphs. The package also implements various other models which are briefly introduced in this article. Besides the approximation methods, the package contains methods for simulation, prediction, and statistical inference for such models, as well as interfaces to INLA, inlabru and MetricGraph. With these interfaces, fractional-order SPDEs can be used as model components in general latent Gaussian models, for which full Bayesian inference can be performed, also for fractional models on metric graphs. This includes estimation of the smoothness parameter of the fields. This article describes the computational methods used in the package and summarizes the theoretical basis for these. The main functions of the package are introduced, and their usage is illustrated through various examples.

stat.CO

Log-Gaussian Cox Processes on General Metric Graphs

The modeling of spatial point processes has advanced considerably, yet extending these models to non-Euclidean domains, such as road networks, remains a challenging problem. We propose a novel framework for log-Gaussian Cox processes on general compact metric graphs by leveraging the Gaussian Whittle-Matérn fields, which are solutions to fractional-order stochastic differential equations on metric graphs. To achieve computationally efficient likelihood-based inference, we introduce a numerical approximation of the likelihood that eliminates the need to approximate the Gaussian process. This method, coupled with the exact evaluation of finite-dimensional distributions for Whittle-Matérn fields with integer smoothness, ensures scalability and theoretical rigour, with derived convergence rates for posterior distributions. The framework is implemented in the open-source MetricGraph R package, which integrates seamlessly with R-INLA to support fully Bayesian inference. We demonstrate the applicability and scalability of this approach through an analysis of road accident data from Al-Ahsa, Saudi Arabia, consisting of over 150,000 road segments. By identifying high-risk road segments using exceedance probabilities and excursion sets, our framework provides localized insights into accident hotspots and offers a powerful tool for modeling spatial point processes directly on complex networks.

stat.ME

An explicit link between graphical models and Gaussian Markov random fields on metric graphs

We derive an explicit link between Gaussian Markov random fields on metric graphs and graphical models, and in particular show that a Markov random field restricted to the vertices of the graph is, under mild regularity conditions, a Gaussian graphical model with a distribution which is faithful to its pairwise independence graph, which coincides with the neighbor structure of the metric graph. This is used to show that there are no Gaussian random fields on general metric graphs which are both Markov and isotropic in some suitably regular metric on the graph, such as the geodesic or resistance metrics.

math.PR

Markov properties of Gaussian random fields on compact metric graphs

There has recently been much interest in Gaussian fields on linear networks and, more generally, on compact metric graphs. One proposed strategy for defining such fields on a metric graph $Γ$ is through a covariance function that is isotropic in a metric on the graph. Another is through a fractional-order differential equation $L^{α/2} (τu) = \mathcal{W}$ on $Γ$, where $L = κ^2 - \nabla(a\nabla)$ for (sufficiently nice) functions $κ, a$, and $\mathcal{W}$ is Gaussian white noise. We study Markov properties of these two types of fields. First, we show that no Gaussian random fields exist on general metric graphs that are both isotropic and Markov. Then, we show that the second type of fields, the generalized Whittle--Matérn fields, are Markov if $α\in\mathbb{N}$, and conversely, if $a$ and $κ$ are constant and $u$ is Markov, then $α\in\mathbb{N}$. Further, if $α\in\mathbb{N}$, a generalized Whittle--Matérn field $u$ is Markov of order $α$, which means that the field $u$ in one region $S\subsetΓ$ is conditionally independent of $u$ in $Γ\setminus S$ given the values of $u$ and its $α-1$ derivatives on $\partial S$. Finally, we provide two results as consequences of the theory developed: first we prove that the Markov property implies an explicit characterization of $u$ on a fixed edge $e$, revealing that the conditional distribution of $u$ on $e$ given the values at the two vertices connected to $e$ is independent of the geometry of $Γ$; second, we show that the solution to $L^{1/2}(τu) = \mathcal{W}$ on $Γ$ can obtained by conditioning independent generalized Whittle--Matérn processes on the edges, with $α=1$ and Neumann boundary conditions, on being continuous at the vertices.

math.PR

Fast and robust cross-validation-based scoring rule inference for spatial statistics

Scoring rules are aimed at evaluation of the quality of predictions, but can also be used for estimation of parameters in statistical models. We propose estimating parameters of multivariate spatial models by maximising the average leave-one-out cross-validation score. This method, LOOS, thus optimises predictions instead of maximising the likelihood. The method allows for fast computations for Gaussian models with sparse precision matrices, such as spatial Markov models. It also makes it possible to tailor the estimator's robustness to outliers and their sensitivity to spatial variations of uncertainty through the choice of the scoring rule which is used in the maximisation. The effects of the choice of scoring rule which is used in LOOS are studied by simulation in terms of computation time, statistical efficiency, and robustness. Various popular scoring rules and a new scoring rule, the root score, are compared to maximum likelihood estimation. The results confirmed that for spatial Markov models the computation time for LOOS was much smaller than for maximum likelihood estimation. Furthermore, the standard deviations of parameter estimates were smaller for maximum likelihood estimation, although the differences often were small. The simulations also confirmed that the usage of a robust scoring rule results in robust LOOS estimates and that the robustness provides better predictive quality for spatial data with outliers. Finally, the new inference method was applied to ERA5 temperature reanalysis data for the contiguous United States and the average July temperature for the years 1940 to 2023, and this showed that the LOOS estimator provided parameter estimates that were more than a hundred times faster to compute compared to maximum-likelihood estimation, and resulted in a model with better predictive performance.

stat.ME

Efficient Modeling of Spatial Extremes over Large Geographical Domains

Various natural phenomena exhibit spatial extremal dependence at short spatial distances. However, existing models proposed in the spatial extremes literature often assume that extremal dependence persists across the entire domain. This is a strong limitation when modeling extremes over large geographical domains, and yet it has been mostly overlooked in the literature. We here develop a more realistic Bayesian framework based on a novel Gaussian scale mixture model, with the Gaussian process component defined by a stochastic partial differential equation yielding a sparse precision matrix, and the random scale component modeled as a low-rank Pareto-tailed or Weibull-tailed spatial process determined by compactly-supported basis functions. We show that our proposed model is approximately tail-stationary and that it can capture a wide range of extremal dependence structures. Its inherently sparse structure allows fast Bayesian computations in high spatial dimensions based on a customized Markov chain Monte Carlo algorithm prioritizing calibration in the tail. We fit our model to analyze heavy monsoon rainfall data in Bangladesh. Our study shows that our model outperforms natural competitors and that it fits precipitation extremes well. We finally use the fitted model to draw inference on long-term return levels for marginal precipitation and spatial aggregates.

stat.ME

Spatial confounding under infill asymptotics

The estimation of regression parameters in spatially referenced data plays a crucial role across various scientific domains. A common approach involves employing an additive regression model to capture the relationship between observations and covariates, accounting for spatial variability not explained by the covariates through a Gaussian random field. While theoretical analyses of such models have predominantly focused on prediction and covariance parameter inference, recent attention has shifted towards understanding the theoretical properties of regression coefficient estimates, particularly in the context of spatial confounding. This article studies the effect of misspecified covariates, in particular when the misspecification changes the smoothness. We analyze the theoretical properties of the generalize least-square estimator under infill asymptotics, and show that the estimator can have counter-intuitive properties. In particular, the estimated regression coefficients can converge to zero as the number of observations increases, despite high correlations between observations and covariates. Perhaps even more surprising, the estimates can diverge to infinity under certain conditions. Through an application to temperature and precipitation data, we show that both behaviors can be observed for real data. Finally, we propose a simple fix to the problem by adding a smoothing step in the regression.

math.ST

Locally tail-scale invariant scoring rules for evaluation of extreme value forecasts

Statistical analysis of extremes can be used to predict the probability of future extreme events, such as large rainfalls or devastating windstorms. The quality of these forecasts can be measured through scoring rules. Locally scale invariant scoring rules give equal importance to the forecasts at different locations regardless of differences in the prediction uncertainty. This is a useful feature when computing average scores but can be an unnecessarily strict requirement when mostly concerned with extremes. We propose the concept of local weight-scale invariance, describing scoring rules fulfilling local scale invariance in a certain region of interest, and as a special case local tail-scale invariance, for large events. Moreover, a new version of the weighted Continuous Ranked Probability score (wCRPS) called the scaled wCRPS (swCRPS) that possesses this property is developed and studied. The score is a suitable alternative for scoring extreme value models over areas with varying scale of extreme events, and we derive explicit formulas of the score for the Generalised Extreme Value distribution. The scoring rules are compared through simulation, and their usage is illustrated in modelling of extreme water levels, annual maximum rainfalls, and in an application to non-extreme forecast for the prediction of air pollution.

stat.ME

Enhanced spatial modeling on linear networks using Gaussian Whittle-Matérn fields

Spatial statistics is traditionally based on stationary models on $\mathbb{R^d}$ like Matérn fields. The adaptation of traditional spatial statistical methods, originally designed for stationary models in Euclidean spaces, to effectively model phenomena on linear networks such as stream systems and urban road networks is challenging. The current study aims to analyze the incidence of traffic accidents on road networks using three different methodologies and compare the model performance for each methodology. Initially, we analyzed the application of spatial triangulation precisely on road networks instead of traditional continuous regions. However, this approach posed challenges in areas with complex boundaries, leading to the emergence of artificial spatial dependencies. To address this, we applied an alternative computational method to construct nonstationary barrier models. Finally, we explored a recently proposed class of Gaussian processes on compact metric graphs, the Whittle-Matérn fields, defined by a fractional SPDE on the metric graph. The latter fields are a natural extension of Gaussian fields with Matérn covariance functions on Euclidean domains to non-Euclidean metric graph settings. A ten-year period (2010-2019) of daily traffic-accident records from Barcelona, Spain have been used to evaluate the three models referred above. While comparing model performance we observed that the Whittle-Matérn fields defined directly on the network outperformed the network triangulation and barrier models. Due to their flexibility, the Whittle-Matérn fields can be applied to a wide range of environmental problems on linear networks such as spatio-temporal modeling of water contamination in stream networks or modeling air quality or accidents on urban road networks.

stat.AP

Regularity and numerical approximation of fractional elliptic differential equations on compact metric graphs

The fractional differential equation $L^βu = f$ posed on a compact metric graph is considered, where $β>0$ and $L = κ^2 - \nabla(a\nabla)$ is a second-order elliptic operator equipped with certain vertex conditions and sufficiently smooth and positive coefficients $κ, a$. We demonstrate the existence of a unique solution for a general class of vertex conditions and derive the regularity of the solution in the specific case of Kirchhoff vertex conditions. These results are extended to the stochastic setting when $f$ is replaced by Gaussian white noise. For the deterministic and stochastic settings under generalized Kirchhoff vertex conditions, we propose a numerical solution based on a finite element approximation combined with a rational approximation of the fractional power $L^{-β}$. For the resulting approximation, the strong error is analyzed in the deterministic case, and the strong mean squared error as well as the $L_2(Γ\times Γ)$-error of the covariance function of the solution are analyzed in the stochastic setting. Explicit rates of convergences are derived for all cases. Numerical experiments for ${L = κ^2 - Δ, κ>0}$ are performed to illustrate the results.

math.NA

Statistical inference for Gaussian Whittle-Matérn fields on metric graphs

Whittle-Matérn fields are a recently introduced class of Gaussian processes on metric graphs, which are specified as solutions to a fractional-order stochastic differential equation. Unlike earlier covariance-based approaches for specifying Gaussian fields on metric graphs, the Whittle-Matérn fields are well-defined for any compact metric graph and can provide Gaussian processes with differentiable sample paths. We derive the main statistical properties of the model class, particularly the consistency and asymptotic normality of maximum likelihood estimators of model parameters and the necessary and sufficient conditions for asymptotic optimality properties of linear prediction based on the model with misspecified parameters. The covariance function of the Whittle-Matérn fields is generally unavailable in closed form, and they have therefore been challenging to use for statistical inference. However, we show that for specific values of the fractional exponent, when the fields have Markov properties, likelihood-based inference and spatial prediction can be performed exactly and computationally efficiently. This facilitates using the Whittle-Matérn fields in statistical applications involving big datasets without the need for any approximations. The methods are illustrated via an application to modeling of traffic data, where allowing for differentiable processes dramatically improves the results.

stat.ME

Extremal Dependence of Moving Average Processes Driven by Exponential-Tailed Lévy Noise

Moving average processes driven by exponential-tailed Lévy noise are important extensions of their Gaussian counterparts in order to capture deviations from Gaussianity, more flexible dependence structures, and sample paths with jumps. Popular examples include non-Gaussian Ornstein--Uhlenbeck processes and type G Matérn stochastic partial differential equation random fields. This paper is concerned with the open problem of determining their extremal dependence structure. We leverage the fact that such processes admit approximations on grids or triangulations that are used in practice for efficient simulations and inference. These approximations can be expressed as special cases of a class of linear transformations of independent, exponential-tailed random variables, that bridge asymptotic dependence and independence in a novel, tractable way. This result is of independent interest since models that can capture both extremal dependence regimes are scarce and the construction of such flexible models is an active area of research. This new fundamental result allows us to show that the integral approximation of general moving average processes with exponential-tailed Lévy noise is asymptotically independent when the mesh is fine enough. Under mild assumptions on the kernel function we also derive the limiting residual tail dependence function. For the popular exponential-tailed Ornstein--Uhlenbeck process we prove that it is asymptotically independent, but with a different residual tail dependence function than its Gaussian counterpart. Our results are illustrated through simulation studies.

math.ST

Covariance-based rational approximations of fractional SPDEs for computationally efficient Bayesian inference

The stochastic partial differential equation (SPDE) approach is widely used for modeling large spatial datasets. It is based on representing a Gaussian random field $u$ on $\mathbb{R}^d$ as the solution of an elliptic SPDE $L^βu = \mathcal{W}$ where $L$ is a second-order differential operator, $2β$ (belongs to natural number starting from 1) is a positive parameter that controls the smoothness of $u$ and $\mathcal{W}$ is Gaussian white noise. A few approaches have been suggested in the literature to extend the approach to allow for any smoothness parameter satisfying $β>d/4$. Even though those approaches work well for simulating SPDEs with general smoothness, they are less suitable for Bayesian inference since they do not provide approximations which are Gaussian Markov random fields (GMRFs) as in the original SPDE approach. We address this issue by proposing a new method based on approximating the covariance operator $L^{-2β}$ of the Gaussian field $u$ by a finite element method combined with a rational approximation of the fractional power. This results in a numerically stable GMRF approximation which can be combined with the integrated nested Laplace approximation (INLA) method for fast Bayesian inference. A rigorous convergence analysis of the method is performed and the accuracy of the method is investigated with simulated data. Finally, we illustrate the approach and corresponding implementation in the R package rSPDE via an application to precipitation data which is analyzed by combining the rSPDE package with the R-INLA software for full Bayesian inference.

stat.ME

Robustness, model checking and latent Gaussian models

Model checking is essential to evaluate the adequacy of statistical models and the validity of inferences drawn from them. Particularly, hierarchical models such as latent Gaussian models (LGMs) pose unique challenges as it is difficult to check assumptions about the distribution of the latent parameters. Discrepancy measures are often used to quantify the degree to which a model fit deviates from the observed data. We construct discrepancy measures by (a) defining an alternative model with relaxed assumptions and (b) deriving the discrepancy measure most sensitive to discrepancies induced by this alternative model. We also promote a workflow for model criticism that combines model checking with subsequent robustness analysis. As a result, we obtain a general recipe to check assumptions in LGMs and the impact of these assumptions on the results. We demonstrate the ideas by assessing the latent Gaussianity assumption, a crucial but often overlooked assumption in LGMs. We illustrate the methods via examples utilising Stan and provide functions for easy usage of the methods for general models fitted through R-INLA.

stat.ME