SearcharxivSearch

arXiv subjects

Jarno Vanhatalo

Publications and source records attributed to Jarno Vanhatalo.

18 recordsLinked to original sources

On Asymptotic Outlier Rejection in Bayesian Mixed Poisson Regression Models Under Extreme Target and Covariate Values

Bayesian models are defined to be fully robust against outliers if observations infinitely far from the other data do not influence the posterior. In regression models, this entails a need to consider outliers in both target and covariate values. While in linear regression these cases are interchangeable, as both lead to anomalously large residuals, it has remained unclear whether this symmetry applies to generalized linear models. Moreover, only recently, the theoretical understanding of generalized linear models' robustness to outliers in target values has progressed significantly. Importantly, Hamura et al. (2025, arXiv:2106.10503) presented sufficient conditions for mixed Poisson count regression models to be robust against infinitely large target values and proposed a mixed Poisson-Rescaled Beta model fulfilling these conditions. We continue from their work and study the robustness properties of mixed Poisson regression models with Gaussian latent variables in the presence of outliers in covariates. We show that in count regression the symmetry between covariate and target outliers breaks: mixed Poisson models are not robust to outlier covariates even if they were robust to target outliers. Furthermore, we show that, as a covariate gets infinitely large, the corresponding regression coefficient posterior collapses to a point-mass distribution concentrated around zero. We then summarize the theoretical asymptotic outlier rejection properties of Gamma, log-Student's-$t$, and Rescaled Beta mixed Poisson models in the presence of outliers in either target or covariate values. We also study the properties of these three mixed Poisson models in the presence of moderate outliers with simulations and a real world case study. Experimental results indicate that all mixed Poisson models are less sensitive to moderate outliers than (non-mixed) Poisson.

stat.ME

Integrated population model reveals human and environment driven changes in Baltic ringed seal (Pusa hispida botnica) demography and behavior

Integrated population models (IPMs) are a promising approach to test ecological theories and assess wildlife populations in dynamic and uncertain conditions. By combining multiple data sources into a unified model, they enable the parametrization of versatile, mechanistic models that can predict population dynamics in novel circumstances. Here, we present a Bayesian IPM for the ringed seal (Pusa hispida botnica) population inhabiting the Bothnian Bay in the Baltic Sea. Despite the availability of long-term monitoring data, traditional assessment methods have faltered due to dynamic environmental conditions, varying reproductive rates, and the recently re-introduced hunting, thus limiting the quality of information available to managers. We fit our model to census and various demographic, reproductive, and harvest data from 1988 to 2023 to provide a comprehensive assessment of past population trends, and predict population response to alternative hunting scenarios. We estimated that 20,000 to 36,000 ringed seals inhabited the Bothnian Bay in 2024, increasing at a rate of 3% to 6% per year. Reproductive rates have increased since 1988, leading to a substantial increase in the growth rate up until 2015. However, the re-introduction of hunting has since reduced the growth rate, and even minor quota increases are likely to reduce it further. Our results also support the hypothesis that a greater proportion of the population hauls out under lower ice cover circumstances, leading to higher aerial survey results in such years. In general, our study demonstrates the value of IPMs for monitoring wildlife populations under changing environments, and supporting science-based management decisions.

q-bio.PE

Bayesian Calibration and Uncertainty Quantification for a Large Nutrient Load Impact Model

Nutrient load simulators are large, deterministic, models that simulate the hydrodynamics and biogeochemical processes in aquatic ecosystems. They are central tools for planning cost efficient actions to fight eutrophication since they allow scenario predictions on impacts of nutrient load reductions to, e.g., harmful algal biomass growth. Due to being computationally heavy, the uncertainties related to these predictions are typically not rigorously assessed though. In this work, we developed a novel Bayesian computational approach for estimating the uncertainties in predictions of the Finnish coastal nutrient load model FICOS. First, we constructed a likelihood function for the multivariate spatiotemporal outputs of the FICOS model. Then, we used Bayes optimization to locate the posterior mode for the model parameters conditional on long term monitoring data. After that, we constructed a space filling design for FICOS model runs around the posterior mode and used it to train a Gaussian process emulator for the (log) posterior density of the model parameters. We then integrated over this (approximate) parameter posterior to produce probabilistic predictions for algal biomass and chlorophyll a concentration under alternative nutrient load reduction scenarios. Our computational algorithm allowed for fast posterior inference and the Gaussian process emulator had good predictive accuracy within the highest posterior probability mass region. The posterior predictive scenarios showed that the probability to reach the EUs Water Framework Directive objectives in the Finnish Archipelago Sea is generally low even under large load reductions.

stat.ME

Joint Species Distribution Modeling of Percentage Cover Data with Exclusive Competition for Space

Joint species distribution models (JSDM) are among the most important statistical tools in community ecology. They are routinely used for inference and various prediction tasks, such as to build species distribution maps or biomass estimation over spatial areas. Existing JSDM's cannot, however, model mutual exclusion between species, which may happen in some species groups, such as mosses in the bottom layer of a peatland site. We tackle this deficiency in the context of modeling plant percentage cover data, where mutual exclusion arises from limited growing space and competition for light. We propose a hierarchical JSDM where multivariate latent Gaussian variable model describes species' niche preferences and Dirichlet-Multinomial distribution models the observation process and exclusive competition for space between species. We use both stationary and non-stationary multivariate Gaussian processes to model residual phenomena. We also propose a decision theoretic model comparison and validation approach to assess the goodness of JSDMs in four different types of predictive tasks. We apply our models and methods to a case study on modeling vegetation cover in a boreal peatland. Our results show that ignoring the interspecific interactions and competition for space significantly reduces models' predictive performance and leads to biased estimates for total percentage cover both for individual species and over all species combined. A model's relative predictive performance also depends on the model comparison methods highlighting that model comparison and assessment should resemble the true predictive task. Our results also demonstrate that the proposed joint species distribution model can be used to simultaneously infer interspecific correlations in niche preference as well as mutual exclusive competition for space and through that provide novel insight into ecological research.

stat.AP

Species Distribution Modeling with Expert Elicitation and Bayesian Calibration

Species distribution models (SDMs) are key tools in ecology, conservation and management of natural resources. They are commonly trained by scientific survey data but, since surveys are expensive, there is a need for complementary sources of information to train them. To this end, several authors have proposed to use expert elicitation since local citizen and substance area experts can hold valuable information on species distributions. Expert knowledge has been incorporated within SDMs, for example, through informative priors. However, existing approaches pose challenges related to assessment of the reliability of the experts. Since expert knowledge is inherently subjective and prone to biases, we should optimally calibrate experts' assessments and make inference on their reliability. Moreover, demonstrated examples of improved species distribution predictions using expert elicitation compared to using only survey data are few as well. In this work, we propose a novel approach to use expert knowledge on species distribution within SDMs and demonstrate that it leads to significantly better predictions. First, we propose expert elicitation process where experts summarize their belief on a species occurrence proability with maps. Second, we collect survey data to calibrate the expert assessments. Third, we propose a hierarchical Bayesian model that combines the two information sources and can be used to make predictions over the study area. We apply our methods to study the distribution of spring spawning pikeperch larvae in a coastal area of the Gulf of Finland. According to our results, the expert information significantly improves species distribution predictions compared to predictions conditioned on survey data only. However, experts' reliability also varies considerably, and even generally reliable experts had spatially structured biases in their assessments.

stat.ME

Additive multivariate Gaussian processes for joint species distribution modeling with heterogeneous data

Species distribution models (SDM) are a key tool in ecology, conservation and management of natural resources. Two key components of the state-of-the-art SDMs are the description for species distribution response along environmental covariates and the spatial random effect. Joint species distribution models (JSDMs) additionally include interspecific correlations which have been shown to improve their descriptive and predictive performance compared to single species models. Current JSDMs are restricted to hierarchical generalized linear modeling framework. These parametric models have trouble in explaining changes in abundance due, e.g., highly non-linear physical tolerance limits which is particularly important when predicting species distribution in new areas or under scenarios of environmental change. On the other hand, semi-parametric response functions have been shown to improve the predictive performance of SDMs in these tasks in single species models. Here, we propose JSDMs where the responses to environmental covariates are modeled with additive multivariate Gaussian processes coded as linear models of coregionalization. These allow inference for wide range of functional forms and interspecific correlations between the responses. We propose also an efficient approach for inference with Laplace approximation and parameterization of the interspecific covariance matrices on the euclidean space. We demonstrate the benefits of our model with two small scale examples and one real world case study. We use cross-validation to compare the proposed model to analogous semi-parametric single species models and parametric single and joint species models in interpolation and extrapolation tasks. The proposed model outperforms the alternative models in all cases. We also show that the proposed model can be seen as an extension of the current state-of-the-art JSDMs to semi-parametric models.

stat.ME

EcoMem: An R package for quantifying ecological memory

Ecological processes may exhibit memory to past disturbances affecting the resilience of ecosystems to future disturbance. Understanding the role of ecological memory in shaping ecosystem responses to disturbance under global change is a critical step toward developing effective adaptive management strategies to maintain ecosystem function and biodiversity. We developed EcoMem, an R package for quantifying ecological memory functions using common environmental time series data (continuous, count, proportional) applying a Bayesian hierarchical framework. The package estimates memory functions for continuous and binary (e.g., disturbance chronology) variables making no a priori assumption on the form of the functions. EcoMem allows users to quantify ecological memory for a wide range of ecosystem processes and responses. The utility of the package to advance understanding of the memory of ecosystems to environmental drivers is demonstrated using a simulated dataset and a case study assessing the memory of boreal tree growth to insect defoliation.

stat.CO

Non-stationary Gaussian models with physical barriers

The classical tools in spatial statistics are stationary models, like the Matérn field. However, in some applications there are boundaries, holes, or physical barriers in the study area, e.g. a coastline, and stationary models will inappropriately smooth over these features, requiring the use of a non-stationary model. We propose a new model, the Barrier model, which is different from the established methods as it is not based on the shortest distance around the physical barrier, nor on boundary conditions. The Barrier model is based on viewing the Matérn correlation, not as a correlation function on the shortest distance between two points, but as a collection of paths through a Simultaneous Autoregressive (SAR) model. We then manipulate these local dependencies to cut off paths that are crossing the physical barriers. To make the new SAR well behaved, we formulate it as a stochastic partial differential equation (SPDE) that can be discretised to represent the Gaussian field, with a sparse precision matrix that is automatically positive definite. The main advantage with the Barrier model is that the computational cost is the same as for the stationary model. The model is easy to use, and can deal with both sparse data and very complex barriers, as shown in an application in the Finnish Archipelago Sea. Additionally, the Barrier model is better at reconstructing the modified Horseshoe test function than the standard models used in R-INLA.

stat.AP

Correcting boundary over-exploration deficiencies in Bayesian optimization with virtual derivative sign observations

Bayesian optimization (BO) is a global optimization strategy designed to find the minimum of an expensive black-box function, typically defined on a compact subset of $\mathcal{R}^d$, by using a Gaussian process (GP) as a surrogate model for the objective. Although currently available acquisition functions address this goal with different degree of success, an over-exploration effect of the contour of the search space is typically observed. However, in problems like the configuration of machine learning algorithms, the function domain is conservatively large and with a high probability the global minimum does not sit on the boundary of the domain. We propose a method to incorporate this knowledge into the search process by adding virtual derivative observations in the \gp at the boundary of the search space. We use the properties of GPs to impose conditions on the partial derivatives of the objective. The method is applicable with any acquisition function, it is easy to use and consistently reduces the number of evaluations required to optimize the objective irrespective of the acquisition used. We illustrate the benefits of our approach in an extensive experimental comparison.

stat.ML

Bayesian model-based spatiotemporal survey design for log-Gaussian Cox process

In geostatistics, the design for data collection is central for accurate prediction and parameter inference. One important class of geostatistical models is log-Gaussian Cox process (LGCP) which is used extensively, for example, in ecology. However, there are no formal analyses on optimal designs for LGCP models. In this work, we develop a novel model-based experimental design for LGCP modeling of spatiotemporal point process data. We propose a new spatially balanced rejection sampling design which directs sampling to spatiotemporal locations that are a priori expected to provide most information. We compare the rejection sampling design to traditional balanced and uniform random designs using the average predictive variance loss function and the Kullback-Leibler divergence between prior and posterior for the LGCP intensity function. Our results show that the rejection sampling method outperforms the corresponding balanced and uniform random sampling designs for LGCP whereas the latter work better for models with Gaussian models. We perform a case study applying our new sampling design to plan a survey for species distribution modeling on larval areas of two commercially important fish stocks on Finnish coastal areas. The case study results show that rejection sampling designs give considerable benefit compared to traditional designs. Results show also that best performing designs may vary considerably between target species.

stat.ME

Laplace approximation and the natural gradient for Gaussian process regression with the heteroscedastic Student-t model

This paper considers the Laplace method to derive approximate inference for the Gaussian process (GP) regression in the location and scale parameters of the Student-t probabilistic model. This allows both mean and variance of the data to vary as a function of covariates with the attractive feature that the Student-t model has been widely used as a useful tool for robustifying data analysis. The challenge in the approximate inference for the GP regression with the Student-t probabilistic model, lies in the analytical intractability of the posterior distribution and the lack of concavity of the log-likelihood function. We present the natural gradient adaptation for the estimation process which primarily relies on the property that the Student-t model naturally has orthogonal parametrization with respect to the location and scale paramaters. Due to this particular property of the model, we also introduce an alternative Laplace approximation by using the Fisher information matrix in place of the Hessian matrix of the negative log-likelihood function. According to experiments this alternative approximation provides very similar posterior approximations and predictive performance when compared to the traditional Laplace approximation. We also compare both of these Laplace approximations with the Monte Carlo Markov Chain (MCMC) method. Moreover, we compare our heteroscedastic Student-t model and the GP regression with the heteroscedastic Gaussian model. We also discuss how our approach can improve the inference algorithm in cases where the probabilistic model assumed for the data is not log-concave.

stat.ME

Optimal design of observational studies: overview and synthesis

We review typical design problems encountered in the planning of observational studies and propose a unifying framework that allows us to use the same concepts and notation for different problems. In the framework, the design is defined as a probability measure in the space of observational processes that determine whether the value of a variable is observed for a specific unit at the given time. The optimal design is then defined, according to Bayesian decision theory, to be the one that maximizes the expected utility related to the design. We present examples on the use of the framework and discuss methods for deriving optimal or approximately optimal designs.

math.ST

A Bayesian length-based population dynamics model for northern shrimp (Pandalus Borealis)

We introduce a fully length-based Bayesian model for the population dynamics of northern shrimp (Pandalus Borealis). This has the advantage of structuring the population in terms of a directly observable quantity, requiring no indirect estimation of age distributions from measurements of size. The introduced model is intended as a simplistic prototype around which further developments and refinements can be built. As a case study, we use the model to analyze the population of Skagerrak and the Norwegian Deep in the years 1988-2012.

stat.AP

Bayesian Modeling with Gaussian Processes using the GPstuff Toolbox

Gaussian processes (GP) are powerful tools for probabilistic modeling purposes. They can be used to define prior distributions over latent functions in hierarchical Bayesian models. The prior over functions is defined implicitly by the mean and covariance function, which determine the smoothness and variability of the function. The inference can then be conducted directly in the function space by evaluating or approximating the posterior process. Despite their attractive theoretical properties GPs provide practical challenges in their implementation. GPstuff is a versatile collection of computational tools for GP models compatible with Linux and Windows MATLAB and Octave. It includes, among others, various inference methods, sparse approximations and tools for model assessment. In this work, we review these tools and demonstrate the use of GPstuff in several models.

stat.ML

Experiences in Bayesian Inference in Baltic Salmon Management

We review a success story regarding Bayesian inference in fisheries management in the Baltic Sea. The management of salmon fisheries is currently based on the results of a complex Bayesian population dynamic model, and managers and stakeholders use the probabilities in their discussions. We also discuss the technical and human challenges in using Bayesian modeling to give practical advice to the public and to government officials and suggest future areas in which it can be applied. In particular, large databases in fisheries science offer flexible ways to use hierarchical models to learn the population dynamics parameters for those by-catch species that do not have similar large stock-specific data sets like those that exist for many target species. This information is required if we are to understand the future ecosystem risks of fisheries.

stat.ME

Modelling local and global phenomena with sparse Gaussian processes

Much recent work has concerned sparse approximations to speed up the Gaussian process regression from the unfavorable O(n3) scaling in computational time to O(nm2). Thus far, work has concentrated on models with one covariance function. However, in many practical situations additive models with multiple covariance functions may perform better, since the data may contain both long and short length-scale phenomena. The long length-scales can be captured with global sparse approximations, such as fully independent conditional (FIC), and the short length-scales can be modeled naturally by covariance functions with compact support (CS). CS covariance functions lead to naturally sparse covariance matrices, which are computationally cheaper to handle than full covariance matrices. In this paper, we propose a new sparse Gaussian process model with two additive components: FIC for the long length-scales and CS covariance function for the short length-scales. We give theoretical and experimental results and show that under certain conditions the proposed model has the same computational complexity as FIC. We also compare the model performance of the proposed model to additive models approximated by fully and partially independent conditional (PIC). We use real data sets and show that our model outperforms FIC and PIC approximations for data sets with two additive phenomena.

cs.LG

Speeding up the binary Gaussian process classification

Gaussian processes (GP) are attractive building blocks for many probabilistic models. Their drawbacks, however, are the rapidly increasing inference time and memory requirement alongside increasing data. The problem can be alleviated with compactly supported (CS) covariance functions, which produce sparse covariance matrices that are fast in computations and cheap to store. CS functions have previously been used in GP regression but here the focus is in a classification problem. This brings new challenges since the posterior inference has to be done approximately. We utilize the expectation propagation algorithm and show how its standard implementation has to be modified to obtain computational benefits from the sparse covariance matrices. We study four CS covariance functions and show that they may lead to substantial speed up in the inference time compared to globally supported functions.

stat.ML

Gaussian Process Regression with a Student-t Likelihood

This paper considers the robust and efficient implementation of Gaussian process regression with a Student-t observation model. The challenge with the Student-t model is the analytically intractable inference which is why several approximative methods have been proposed. The expectation propagation (EP) has been found to be a very accurate method in many empirical studies but the convergence of the EP is known to be problematic with models containing non-log-concave site functions such as the Student-t distribution. In this paper we illustrate the situations where the standard EP fails to converge and review different modifications and alternative algorithms for improving the convergence. We demonstrate that convergence problems may occur during the type-II maximum a posteriori (MAP) estimation of the hyperparameters and show that the standard EP may not converge in the MAP values in some difficult cases. We present a robust implementation which relies primarily on parallel EP updates and utilizes a moment-matching-based double-loop algorithm with adaptively selected step size in difficult cases. The predictive performance of the EP is compared to the Laplace, variational Bayes, and Markov chain Monte Carlo approximations.

stat.ML