SearcharxivSearch

arXiv subjects

Haakon Bakka

Publications and source records attributed to Haakon Bakka.

15 recordsLinked to original sources

Generalised logistic regression with vine copulas

We propose a generalisation of the logistic regression model, that aims to account for non-linear main effects and complex interactions, while keeping the model inherently explainable. This is obtained by starting with log-odds that are linear in the covariates, and adding non-linear terms that depend on at least two covariates. More specifically, we use a generative specification of the model, consisting of a combination of certain margins on natural exponential form, combined with vine copulas. The estimation of the model is however based on the discriminative likelihood, and dependencies between covariates are included in the model, only if they contribute significantly to the distinction between the two classes. Further, a scheme for model selection and estimation is presented. The methods described in this paper are implemented in the R package LogisticCopula. In order to assess the performance of our model, we ran an extensive simulation study. The results from the study, as well as from a couple of examples on real data, showed that our model performs at least as well as natural competitors, especially in the presence of non-linearities and complex interactions, even when $n$ is not large compared to $p$.

stat.ME

High stakes classification with multiple unknown classes based on imperfect data

High stakes classification refers to classification problems where erroneously predicting the wrong class is very bad, but assigning "unknown" is acceptable. We make the argument that these problems require us to give multiple unknown classes, to get the most information out of our analysis. With imperfect data we refer to covariates with a large number of missing values, large noise variance, and some errors in the data. The combination of high stakes classification and imperfect data is very common in practice, but it is very difficult to work on using current methods. We present a one-class classifier (OCC) to solve this problem, and we call it NBP. The classifier is based on Naive Bayes, simple to implement, and interpretable. We show that NBP gives both good predictive performance, and works for high stakes classification based on imperfect data. The model we present is quite simple; it is just an OCC based on density estimation. However, we have always felt a big gap between the applied classification problems we have worked on and the theory and models we use for classification, and this model closes that gap. Our main contribution is the motivation for why this model is a good approach, and we hope that this paper will inspire further development down this path.

stat.AP

A diffusion-based spatio-temporal extension of Gaussian Matérn fields

Gaussian random fields with Matérn covariance functions are popular models in spatial statistics and machine learning. In this work, we develop a spatio-temporal extension of the Gaussian Matérn fields formulated as solutions to a stochastic partial differential equation. The spatially stationary subset of the models have marginal spatial Matérn covariances, and the model also extends to Whittle-Matérn fields on curved manifolds, and to more general non-stationary fields. In addition to the parameters of the spatial dependence (variance, smoothness, and practical correlation range) it additionally has parameters controlling the practical correlation range in time, the smoothness in time, and the type of non-separability of the spatio-temporal covariance. Through the separability parameter, the model also allows for separable covariance functions. We provide a sparse representation based on a finite element approximation, that is well suited for statistical inference and which is implemented in the R-INLA software. The flexibility of the model is illustrated in an application to spatio-temporal modeling of global temperature data.

stat.ME

How to solve the stochastic partial differential equation that gives a Matérn random field using the finite element method

This tutorial teaches parts of the finite element method (FEM), and solves a stochastic partial differential equation (SPDE). The contents herein are considered "known" in the numerics literature, but for statisticians it is very difficult to find a resource for learning these ideas in a timely manner (without doing a year's worth of courses in numerics). The goal of this tutorial is to be pedagogical and explain the computations/theory to a statistician. This is not a practical tutorial, there is little computer code, and no data analysis.

stat.CO

Approximate Bayesian inference for analysis of spatio-temporal flood frequency data

Extreme floods cause casualties, and widespread damage to property and vital civil infrastructure. We here propose a Bayesian approach for predicting extreme floods using the generalized extreme-value (GEV) distribution within gauged and ungauged catchments. A major methodological challenge is to find a suitable parametrization for the GEV distribution when covariates or latent spatial effects are involved. Other challenges involve balancing model complexity and parsimony using an appropriate model selection procedure, and making inference using a reliable and computationally efficient approach. Our approach relies on a latent Gaussian modeling framework with a novel multivariate link function designed to separate the interpretation of the parameters at the latent level and to avoid unreasonable estimates of the shape and time trend parameters. Structured additive regression models are proposed for the four parameters at the latent level. For computational efficiency with large datasets and richly parametrized models, we exploit an accurate and fast approximate Bayesian inference approach. We applied our proposed methodology to annual peak river flow data from 554 catchments across the United Kingdom (UK). Our model performed well in terms of flood predictions for both gauged and ungauged catchments. The results show that the spatial model components for the transformed location and scale parameters, and the time trend, are all important. Posterior estimates of the time trend parameters correspond to an average increase of about $1.5\%$ per decade and reveal a spatial structure across the UK. To estimate return levels for spatial aggregates, we further develop a novel copula-based post-processing approach of posterior predictive samples, in order to mitigate the effect of the conditional independence assumption at the data level, and we show that our approach provides accurate results.

stat.ME

High-resolution Bayesian mapping of landslide hazard with unobserved trigger event

Statistical models for landslide hazard enable mapping of risk factors and landslide occurrence intensity by using geomorphological covariates available at high spatial resolution. However, the spatial distribution of the triggering event (e.g., precipitation or earthquakes) is often not directly observed. In this paper, we develop Bayesian spatial hierarchical models for point patterns of landslide occurrences using different types of log-Gaussian Cox processes. Starting from a competitive baseline model that captures the unobserved precipitation trigger through a spatial random effect at slope unit resolution, we explore novel complex model structures that take clusters of events arising at small spatial scales into account, as well as nonlinear or spatially-varying covariate effects. For a 2009 event of around 4000 precipitation-triggered landslides in Sicily, Italy, we show how to fit our proposed models efficiently using the integrated nested Laplace approximation (INLA), and rigorously compare the performance of our models both from a statistical and applied perspective. In this context, we argue that model comparison should not be based on a single criterion, and that different models of various complexity may provide insights into complementary aspects of the same applied problem. In our application, our models are found to have mostly the same spatial predictive performance, implying that key to successful prediction is the inclusion of a slope-unit resolved random effect capturing the precipitation trigger. Interestingly, a parsimonious formulation of space-varying slope effects reflects a physical interpretation of the precipitation trigger: in subareas with weak trigger, the slope steepness is shown to be mostly irrelevant.

stat.ME

A principled distance-based prior for the shape of the Weibull model

The use of flat or weakly informative priors is popular due to the objective a priori belief in the absence of strong prior information. In the case of the Weibull model the improper uniform, equal parameter gamma and joint Jeffrey's priors for the shape parameter are popular choices. The effects and behaviors of these priors have yet to be established from a modeling viewpoint, especially their ability to reduce to the simpler exponential model. In this work we propose a new principled prior for the shape parameter of the Weibull model, originating from a prior on the distance function, and advocate this new prior as a principled choice in the absence of strong prior information. This new prior can then be used in models with a Weibull modeling component, like competing risks, joint and spatial models, to mention a few. This prior is available in the R-INLA for use, and is applied in a joint longitudinal-survival model framework using the INLA method.

stat.ME

Max-and-Smooth: a two-step approach for approximate Bayesian inference in latent Gaussian models

With modern high-dimensional data, complex statistical models are necessary, requiring computationally feasible inference schemes. We introduce Max-and-Smooth, an approximate Bayesian inference scheme for a flexible class of latent Gaussian models (LGMs) where one or more of the likelihood parameters are modeled by latent additive Gaussian processes. Max-and-Smooth consists of two-steps. In the first step (Max), the likelihood function is approximated by a Gaussian density with mean and covariance equal to either (a) the maximum likelihood estimate and the inverse observed information, respectively, or (b) the mean and covariance of the normalized likelihood function. In the second step (Smooth), the latent parameters and hyperparameters are inferred and smoothed with the approximated likelihood function. The proposed method ensures that the uncertainty from the first step is correctly propagated to the second step. Since the approximated likelihood function is Gaussian, the approximate posterior density of the latent parameters of the LGM (conditional on the hyperparameters) is also Gaussian, thus facilitating efficient posterior inference in high dimensions. Furthermore, the approximate marginal posterior distribution of the hyperparameters is tractable, and as a result, the hyperparameters can be sampled independently of the latent parameters. In the case of a large number of independent data replicates, sparse precision matrices, and high-dimensional latent vectors, the speedup is substantial in comparison to an MCMC scheme that infers the posterior density from the exact likelihood function. The proposed inference scheme is demonstrated on one spatially referenced real dataset and on simulated data mimicking spatial, temporal, and spatio-temporal inference problems. Our results show that Max-and-Smooth is accurate and fast.

stat.ME

Competing risks joint models using R-INLA

The methodological advancements made in the field of joint models are numerous. None the less, the case of competing risks joint models have largely been neglected, especially from a practitioner's point of view. In the relevant works on competing risks joint models, the assumptions of Gaussian linear longitudinal series and proportional cause-specific hazard functions, amongst others, have remained unchallenged. In this paper, we provide a framework based on R-INLA to apply competing risks joint models in a unifying way such that non-Gaussian longitudinal data, spatial structures, time dependent splines and various latent association structures, to mention a few, are all embraced in our approach. Our motivation stems from the SANAD trial which exhibits non-linear longitudinal trajectories and competing risks for failure of treatment. We also present a discrete competing risks joint model for longitudinal count data as well as a spatial competing risks joint model, as specific examples.

stat.ME

New frontiers in Bayesian modeling using the INLA package in R

The INLA package provides a tool for computationally efficient Bayesian modeling and inference for various widely used models, more formally the class of latent Gaussian models. It is a non-sampling based framework which provides approximate results for Bayesian inference, using sparse matrices. The swift uptake of this framework for Bayesian modeling is rooted in the computational efficiency of the approach and catalyzed by the demand presented by the big data era. In this paper, we present new developments within the INLA package with the aim to provide a computationally efficient mechanism for the Bayesian inference of relevant challenging situations.

stat.ME

Joint models as latent Gaussian models - not reinventing the wheel

Joint models have received increasing attention during recent years with extensions into various directions; numerous hazard functions, different association structures, linear and non-linear longitudinal trajectories amongst others. Many of these resulted in new R packages and new formulations of the joint model. However, a joint model with a linear bivariate Gaussian association structure is still a latent Gaussian model (LGM) and thus can be implemented using most existing packages for LGM's. In this paper, we will show that these joint models can be implemented from a LGM viewpoint using the R-INLA package. As a particular example, we will focus on the joint model with a non-linear longitudinal trajectory, recently developed and termed the partially linear joint model. Instead of the usual spline approach, we argue for using a Bayesian smoothing spline framework for the joint model that is stable with respect to knot selection and hence less cumbersome for practitioners.

stat.ME

Non-stationary Gaussian models with physical barriers

The classical tools in spatial statistics are stationary models, like the Matérn field. However, in some applications there are boundaries, holes, or physical barriers in the study area, e.g. a coastline, and stationary models will inappropriately smooth over these features, requiring the use of a non-stationary model. We propose a new model, the Barrier model, which is different from the established methods as it is not based on the shortest distance around the physical barrier, nor on boundary conditions. The Barrier model is based on viewing the Matérn correlation, not as a correlation function on the shortest distance between two points, but as a collection of paths through a Simultaneous Autoregressive (SAR) model. We then manipulate these local dependencies to cut off paths that are crossing the physical barriers. To make the new SAR well behaved, we formulate it as a stochastic partial differential equation (SPDE) that can be discretised to represent the Gaussian field, with a sparse precision matrix that is automatically positive definite. The main advantage with the Barrier model is that the computational cost is the same as for the stationary model. The model is easy to use, and can deal with both sparse data and very complex barriers, as shown in an application in the Finnish Archipelago Sea. Additionally, the Barrier model is better at reconstructing the modified Horseshoe test function than the standard models used in R-INLA.

stat.AP

Geostatistical modeling to capture seismic-shaking patterns from earthquake-induced landslides

In this paper, we investigate earthquake-induced landslides using a geostatistical model that includes a latent spatial effect (LSE). The LSE represents the spatially structured residuals in the data, which are complementary to the information carried by the covariates. To determine whether the LSE can capture the residual signal from a given trigger, we test whether the LSE is able to capture the pattern of seismic shaking caused by an earthquake from the distribution of seismically induced landslides, without prior knowledge of the earthquake being included in the statistical model. We assess the landslide intensity, i.e., the expected number of landslide activations per mapping unit, for the area in which landslides triggered by the Wenchuan (M 7.9, May 12, 2008) and Lushan (M 6.6, April 20, 2013) earthquakes overlap. We chose an area of overlapping landslides in order to test our method on landslide inventories located in the near and far fields of the earthquake. We generated three different models for each earthquake-induced landslide scenario: i) seismic parameters only (as a proxy for the trigger); ii) the LSE only; and iii) both seismic parameters and the LSE. The three configurations share the same set of morphometric covariates. This allowed us to study the pattern in the LSE and assess whether it adequately approximated the effects of seismic wave propagation. Moreover, it allowed us to check whether the LSE captured effects that are not explained by the shaking levels, such as topographic amplification. Our results show that the LSE reproduced the shaking patterns in space for both earthquakes with a level of spatial detail even greater than the seismic parameters. In addition, the models including the LSE perform better than conventional models featuring seismic parameters only.

stat.AP

Spatial modelling with R-INLA: A review

Coming up with Bayesian models for spatial data is easy, but performing inference with them can be challenging. Writing fast inference code for a complex spatial model with realistically-sized datasets from scratch is time-consuming, and if changes are made to the model, there is little guarantee that the code performs well. The key advantages of R-INLA are the ease with which complex models can be created and modified, without the need to write complex code, and the speed at which inference can be done even for spatial problems with hundreds of thousands of observations. R-INLA handles latent Gaussian models, where fixed effects, structured and unstructured Gaussian random effects are combined linearly in a linear predictor, and the elements of the linear predictor are observed through one or more likelihoods. The structured random effects can be both standard areal model such as the Besag and the BYM models, and geostatistical models from a subset of the Matérn Gaussian random fields. In this review, we discuss the large success of spatial modelling with R-INLA and the types of spatial models that can be fitted, we give an overview of recent developments for areal models, and we give an overview of the stochastic partial differential equation (SPDE) approach and some of the ways it can be extended beyond the assumptions of isotropy and separability. In particular, we describe how slight changes to the SPDE approach leads to straight-forward approaches for non-stationary spatial models and non-separable space-time models.

stat.ME

INLA goes extreme: Bayesian tail regression for the estimation of high spatio-temporal quantiles

This work has been motivated by the challenge of the 2017 conference on Extreme-Value Analysis (EVA2017), with the goal of predicting daily precipitation quantiles at the $99.8\%$ level for each month at observed and unobserved locations. We here develop a Bayesian generalized additive modeling framework tailored to estimate complex trends in marginal extremes observed over space and time. Our approach is based on a set of regression equations linked to the exceedance probability above a high threshold and to the size of the excess, the latter being modeled using the generalized Pareto (GP) distribution suggested by Extreme-Value Theory. Latent random effects are modeled additively and semi-parametrically using Gaussian process priors, which provides high flexibility and interpretability. Fast and accurate estimation of posterior distributions may be performed thanks to the Integrated Nested Laplace approximation (INLA), efficiently implemented in the R-INLA software, which we also use for determining a nonstationary threshold based on a model for the body of the distribution. We show that the GP distribution meets the theoretical requirements of INLA, and we then develop a penalized complexity prior specification for the tail index, which is a crucial parameter for extrapolating tail event probabilities. This prior concentrates mass close to a light exponential tail while allowing heavier tails by penalizing the distance to the exponential distribution. We illustrate this methodology through the modeling of spatial and seasonal trends in daily precipitation data provided by the EVA2017 challenge. Capitalizing on R-INLA's fast computation capacities and large distributed computing resources, we conduct an extensive cross-validation study to select model parameters governing the smoothness of trends. Our results outperform simple benchmarks and are comparable to the best-scoring approach.

stat.ME