SearcharxivSearch

arXiv subjects

Ruth King

Publications and source records attributed to Ruth King.

At least 19 recordsLinked to original sources

Incorporating Animal Movement into Continuous-Time Spatial Capture-Recapture Models

Estimation of wildlife population size and spatial dynamics is central to ecology and conservation. Spatial capture-recapture (SCR) models estimate abundance, using detections at sensors such as camera traps, by linking detection probability to the distance between detectors and latent individual activity centres. However, standard SCR formulations assume detections are conditionally independent over time given activity centres without explicitly modelling movement between detection events. This assumption can be problematic when individuals exhibit movement-driven dependence in detections, potentially leading to biased inference on population size. We address this important issue by developing a continuous-time framework that integrates movement into spatial capture-recapture. Individual movement is modelled as a continuous-time Markov chain over a discretised landscape, and detections arise as state-dependent Poisson events. This yields a Markov-modulated marked Poisson process representation, in which detections provide information about an individual's latent location at the time of observation and allow likelihood-based inference in continuous time. We show through simulation studies that ignoring movement-driven dependence can lead to positively biased estimates of population size, whereas the proposed model recovers unbiased estimates and provides additional inference on space use. An application to camera-trap data of American martens illustrates how the framework yields new insights into movement and density. These results demonstrate that explicitly modelling movement is critical for reliable inference in spatial capture-recapture studies.

stat.ME

Logistic regression models: practical induced prior specification

Bayesian inference for statistical models with a hierarchical structure is often characterized by specification of priors for parameters at different levels of the hierarchy. When higher level parameters are functions of the lower level parameters, specifying a prior on the lower level parameters leads to induced priors on the higher level parameters. However, what are deemed uninformative priors for lower level parameters can induce strikingly non-vague priors for higher level parameters. Depending on the sample size and specific model parameterization, these priors can then have unintended effects on the posterior distribution of the higher level parameters. Here we focus on Bayesian inference for the Bernoulli distribution parameter $\theta$ which is modeled as a function of covariates via a logistic regression, where the coefficients are the lower level parameters for which priors are specified. A specific area of interest and application is the modeling of survival probabilities in capture-recapture data and occupancy and detection probabilities in presence-absence data. In particular we propose alternative priors for the coefficients that yield specific induced priors for $\theta$. We address three induced prior cases. The simplest is when the induced prior for $\theta$ is Uniform(0,1). The second case is when the induced prior for $\theta$ is an arbitrary Beta($\alpha$, $\beta$) distribution. The third case is one where the intercept in the logistic model is to be treated distinct from the partial slope coefficients; e.g., $E[\theta]$ equals a specified value on (0,1) when all covariates equal 0. Simulation studies were carried out to evaluate performance of these priors and the methods were applied to a real presence/absence data set and occupancy modelling.

stat.ME

Grid Particle Gibbs with Ancestor Sampling for State-Space Models

We consider the challenge of estimating the model parameters and latent states of general state-space models within a Bayesian framework. We extend the commonly applied particle Gibbs framework by proposing an efficient particle generation scheme for the latent states. The approach efficiently samples particles using an approximate hidden Markov model (HMM) representation of the general state-space model via a deterministic grid on the state space. We refer to the approach as the grid particle Gibbs with ancestor sampling algorithm. We discuss several computational and practical aspects of the algorithm in detail and highlight further computational adjustments that improve the efficiency of the algorithm. The efficiency of the approach is investigated via challenging regime-switching models, including a post-COVID tourism demand model, and we demonstrate substantial computational gains compared to previous particle Gibbs with ancestor sampling methods.

stat.CO

Incorporating Memory into Continuous-Time Spatial Capture-Recapture Models

Obtaining reliable and precise estimates of wildlife species abundance and distribution is essential for the conservation and management of animal populations and natural reserves. Spatial capture-recapture (SCR) models provide estimates of population size and spatial density from data collected from remote sensors such as camera traps. Such data contain spatial correlation between observations of the same individual, which SCR models partly account for through a latent individual-specific activity centre, a location near which the individual is more likely detected. However, SCR models assume that the observations of an individual are independent over time and space, conditional on its activity centre, so that observed sightings at a given time and location do not influence the probability of being seen at future times and/or locations. This assumption is ecologically unrealistic given the smooth movement of animals over space through time. We propose a new continuous-time modelling framework that incorporates both an individual's (latent) activity centre and its (known) previous location and time of detection. By formulating the detections of an individual as an inhomogeneous temporal Poisson process, we develop a model drawing inspiration from the Ornstein-Uhlenbeck process, which is commonly used to model animal movement. Applying our model to a camera-trap survey of American martens, we observe a substantial improvement in model fit and notable differences in the estimated spatial distribution of activity centres. A simulation study shows that standard SCR models can produce substantially biased population estimates when spatio-temporal dependence is ignored, while the memory-based model remains robust. These findings highlight the importance of accounting for memory of previous detections in SCR models to improve ecological interpretation and inference.

stat.ME

A Bayesian functional model with multilevel partition priors for group studies in neuroscience

The statistical analysis of group studies in neuroscience is particularly challenging due to the complex spatio-temporal nature of the data, its multiple levels and the inter-individual variability in brain responses. In this respect, traditional ANOVA-based studies and linear mixed effects models typically provide only limited exploration of the dynamic of the group brain activity and variability of the individual responses potentially leading to overly simplistic conclusions and/or missing more intricate patterns. In this study we propose a novel Bayesian model-based clustering method for functional data to simultaneously assess group effects and individual deviations over the most important temporal features in the data. To this aim, we develop an innovative multilevel partition prior to model the functional scores of a functional Principal Components decomposition of neuroscientific recordings; this approach returns a thorough exploration of group differences and individual deviations without compromising on the spatio-temporal nature of the data. By means of a simulation study we demonstrate that the proposed model returns correct classification in different clustering scenarios under low and high noise levels in the data. Finally we consider a case study using Electroencephalogram data recorded during an object recognition task where our approach provides new insights into the underlying brain mechanisms generating the data and their variability.

stat.ME

Large Data and (Not Even Very) Complex Ecological Models: When Worlds Collide

We consider the challenges that arise when fitting complex ecological models to 'large' data sets. In particular, we focus on random effect models which are commonly used to describe individual heterogeneity, often present in ecological populations under study. In general, these models lead to a likelihood that is expressible only as an analytically intractable integral. Common techniques for fitting such models to data include, for example, the use of numerical approximations for the integral, or a Bayesian data augmentation approach. However, as the size of the data set increases (i.e. the number of individuals increases), these computational tools may become computationally infeasible. We present an efficient Bayesian model-fitting approach, whereby we initially sample from the posterior distribution of a smaller subsample of the data, before correcting this sample to obtain estimates of the posterior distribution of the full dataset, using an importance sampling approach. We consider several practical issues, including the subsampling mechanism, computational efficiencies (including the ability to parallelise the algorithm) and combining subsampling estimates using multiple subsampled datasets. We demonstrate the approach in relation to individual heterogeneity capture-recapture models. We initially demonstrate the feasibility of the approach via simulated data before considering a challenging real dataset of approximately 30,000 guillemots, and obtain posterior estimates in substantially reduced computational time.

stat.ME

A Point Mass Proposal Method for Bayesian State-Space Model Fitting

State-space models (SSMs) are commonly used to model time series data where the observations depend on an unobserved latent process. However, inference on the model parameters of an SSM can be challenging, especially when the likelihood of the data given the parameters is not available in closed-form. One approach is to jointly sample the latent states and model parameters via Markov chain Monte Carlo (MCMC) and/or sequential Monte Carlo approximation. These methods can be inefficient, mixing poorly when there are many highly correlated latent states or parameters, or when there is a high rate of sample impoverishment in the sequential Monte Carlo approximations. We propose a novel block proposal distribution for Metropolis-within-Gibbs sampling on the joint latent state and parameter space. The proposal distribution is informed by a deterministic hidden Markov model (HMM), defined such that the usual theoretical guarantees of MCMC algorithms apply. We discuss how the HMMs are constructed, the generality of the approach arising from the tuning parameters, and how these tuning parameters can be chosen efficiently in practice. We demonstrate that the proposed algorithm using HMM approximations provides an efficient alternative method for fitting state-space models, even for those that exhibit near-chaotic behavior.

stat.CO

Parameter clustering in Bayesian functional PCA of fMRI data

The extraordinary advancements in neuroscientific technology for brain recordings over the last decades have led to increasingly complex spatio-temporal datasets. To reduce oversimplifications, new models have been developed to be able to identify meaningful patterns and new insights within a highly demanding data environment. To this extent, we propose a new model called parameter clustering functional Principal Component Analysis (PCl-fPCA) that merges ideas from Functional Data Analysis and Bayesian nonparametrics to obtain a flexible and computationally feasible signal reconstruction and exploration of spatio-temporal neuroscientific data. In particular, we use a Dirichlet process Gaussian mixture model to cluster functional principal component scores within the standard Bayesian functional PCA framework. This approach captures the spatial dependence structure among smoothed time series (curves) and its interaction with the time domain without imposing a prior spatial structure on the data. Moreover, by moving the mixture from data to functional principal component scores, we obtain a more general clustering procedure, thus allowing a higher level of intricate insight and understanding of the data. We present results from a simulation study showing improvements in curve and correlation reconstruction compared with different Bayesian and frequentist fPCA models and we apply our method to functional Magnetic Resonance Imaging and Electroencephalogram data analyses providing a rich exploration of the spatio-temporal dependence in brain time series.

stat.ME

A primer on coupled state-switching models for multiple interacting time series

State-switching models such as hidden Markov models or Markov-switching regression models are routinely applied to analyse sequences of observations that are driven by underlying non-observable states. Coupled state-switching models extend these approaches to address the case of multiple observation sequences whose underlying state variables interact. In this paper, we provide an overview of the modelling techniques related to coupling in state-switching models, thereby forming a rich and flexible statistical framework particularly useful for modelling correlated time series. Simulation experiments demonstrate the relevance of being able to account for an asynchronous evolution as well as interactions between the underlying latent processes. The models are further illustrated using two case studies related to a) interactions between a dolphin mother and her calf as inferred from movement data; and b) electronic health record data collected on 696 patients within an intensive care unit.

stat.ME

Continuous-time multi-state capture-recapture models

Multi-state capture-recapture data comprise individual-specific sighting histories together with information on individuals' states related, for example, to breeding status, infection level, or geographical location. Such data are often analysed using the Arnason-Schwarz model, where transitions between states are modelled using a discrete-time Markov chain, making the model most easily applicable to regular time series. When time intervals between capture occasions are not of equal length, more complex time-dependent constructions may be required, increasing the number of parameters to estimate, decreasing interpretability, and potentially leading to reduced precision. Here we develop a novel continuous-time multi-state model that can be regarded as an analogue of the Arnason-Schwarz model for irregularly sampled data. Statistical inference is carried out by regarding the capture-recapture data as realisations from a continuous-time hidden Markov model, which allows the associated efficient algorithms to be used for maximum likelihood estimation and state decoding. To illustrate the feasibility of the modelling framework, we use a long-term survey of bottlenose dolphins where capture occasion are not regularly spaced through time. Here we are particularly interested in seasonal effects on the movement rates of the dolphins along the Scottish east coast. The results reveal seasonal movement patterns between two core areas of their range, providing information that will inform conservation management.

stat.AP

Parameter Redundancy and the Existence of Maximum Likelihood Estimates in Log-linear Models

Log-linear models are typically fitted to contingency table data to describe and identify the relationship between different categorical variables. However, the data may include observed zero cell entries. The presence of zero cell entries can have an adverse effect on the estimability of parameters, due to parameter redundancy. We describe a general approach for determining whether a given log-linear model is parameter redundant for a pattern of observed zeros in the table, prior to fitting the model to the data. We derive the estimable parameters or functions of parameters and also explain how to reduce the unidentifiable model to an identifiable one. Parameter redundant models have a flat ridge in their likelihood function. We further explain when this ridge imposes some additional parameter constraints on the model, which can lead to obtaining unique maximum likelihood estimates for parameters that otherwise would not have been estimable. In contrast to other frameworks, the proposed novel approach informs on those constraints, elucidating the model that is actually being fitted.

stat.ME

Estimating abundance from multiple sampling capture-recapture data via a multi-state multi-period stopover model

The collection of capture-recapture data often involves collecting data on numerous capture occasions over a relatively short period of time. For many study species this process is repeated, for example annually, resulting in capture information spanning multiple sampling periods. The robust design class of models provide a convenient framework in which to analyse all of the available capture data in a single likelihood expression. However, these models typically rely either upon the assumption of closure within a sampling period (the closed robust design) or condition on the number of individuals captured within a sampling period (the open robust design). The models we develop in this paper require neither assumption by explicitly modelling the movement of individuals into the population both within and between the sampling periods, which in turn permits the estimation of abundance. These models are further extended to allow parameters to depend not only on capture occasion but also the amount of time since joining the population and to the case of multi-state data where there is individual time-varying discrete covariate information. We derive an efficient likelihood expression for the new multi-state multi-period stopover model using the hidden Markov model framework. We demonstrate the new model through a simulation study before considering a dataset on great crested newts, Triturus cristatus.

stat.ME

Efficient sequential Monte Carlo algorithms for integrated population models

State-space models are commonly used to describe different forms of ecological data. We consider the case of count data with observation errors. For such data the system process is typically multi-dimensional consisting of coupled Markov processes, where each component corresponds to a different characterisation of the population, such as age group, gender or breeding status. The associated system process equations describe the biological mechanisms under which the system evolves over time. However, there is often limited information in the count data alone to sensibly estimate demographic parameters of interest, so these are often combined with additional ecological observations leading to an integrated data analysis. Unfortunately, fitting these models to the data can be challenging, especially if the state-space model for the count data is non-linear or non-Gaussian. We propose an efficient particle Markov chain Monte Carlo algorithm to estimate the demographic parameters without the need for resorting to linear or Gaussian approximations. In particular, we exploit the integrated model structure to enhance the efficiency of the algorithm. We then incorporate the algorithm into a sequential Monte Carlo sampler in order to perform model comparison with regards to the dependence structure of the demographic parameters. Finally, we demonstrate the applicability and computational efficiency of our algorithms on two real datasets.

stat.ME

Estimation of population size when capture probability depends on individual states

We develop a multi-state model to estimate the size of a closed population from ecological capture-recapture studies. We consider the case where capture-recapture data are not of a simple binary form, but where the state of an individual is also recorded upon every capture as a discrete variable. The proposed multi-state model can be regarded as a generalisation of the commonly applied set of closed population models to a multi-state form. The model permits individuals to move between the different discrete states, whilst allowing heterogeneity within the capture probabilities. A closed-form expression for the likelihood is presented in terms of a set of sufficient statistics. The link between existing models for capture heterogeneity are established, and simulation is used to show that the estimate of population size can be biased when movement between states is not accounted for. The proposed unconditional approach is also compared to a conditional approach to assess estimation bias. The model derived in this paper is motivated by a real ecological data set on great crested newts, Triturus cristatus.

stat.AP

Statistical modelling of individual animal movement: an overview of key methods and a discussion of practical challenges

With the influx of complex and detailed tracking data gathered from electronic tracking devices, the analysis of animal movement data has recently emerged as a cottage industry amongst biostatisticians. New approaches of ever greater complexity are continue to be added to the literature. In this paper, we review what we believe to be some of the most popular and most useful classes of statistical models used to analyze individual animal movement data. Specifically we consider discrete-time hidden Markov models, more general state-space models and diffusion processes. We argue that these models should be core components in the toolbox for quantitative researchers working on stochastic modelling of individual animal movement. The paper concludes by offering some general observations on the direction of statistical analysis of animal movement. There is a trend in movement ecology toward what are arguably overly-complex modelling approaches which are inaccessible to ecologists, unwieldy with large data sets or not based in mainstream statistical practice. Additionally, some analysis methods developed within the ecological community ignore fundamental properties of movement data, potentially leading to misleading conclusions about animal movement. Corresponding approaches, e.g. based on Lévy walk-type models, continue to be popular despite having been largely discredited. We contend that there is a need for an appropriate balance between the extremes of either being overly complex or being overly simplistic, whereby the discipline relies on models of intermediate complexity that are usable by general ecologists, but grounded in well-developed statistical practice and efficient to fit to large data sets.

stat.AP

Capture-recapture abundance estimation using a semi-complete data likelihood approach

Capture-recapture data are often collected when abundance estimation is of interest. In the presence of unobserved individual heterogeneity, specified on a continuous scale for the capture probabilities, the likelihood is not generally available in closed form, but expressible only as an analytically intractable integral. Model-fitting algorithms to estimate abundance most notably include a numerical approximation for the likelihood or use of a Bayesian data augmentation technique considering the complete data likelihood. We consider a Bayesian hybrid approach, defining a "semi-complete" data likelihood, composed of the product of a complete data likelihood component for individuals seen at least once within the study and a marginal data likelihood component for the individuals not seen within the study, approximated using numerical integration. This approach combines the advantages of the two different approaches, with the semi-complete likelihood component specified as a single integral (over the dimension of the individual heterogeneity component). In addition, the models can be fitted within BUGS/JAGS (commonly used for the Bayesian complete data likelihood approach) but with significantly improved computational efficiency compared to the commonly used super-population data augmentation approaches (between about 10 and 77 times more efficient in the two examples we consider). The semi-complete likelihood approach is flexible and applicable to a range of models, including spatially explicit capture-recapture models. The model-fitting approach is applied to two different datasets corresponding to the closed population model $M_h$ for snowshoe hare data and a spatially explicit capture-recapture model applied to gibbon data.

stat.ME

Maximum penalized likelihood estimation in semiparametric capture-recapture models

We discuss the semiparametric modeling of mark-recapture-recovery data where the temporal and/or individual variation of model parameters is explained via covariates. Typically, in such analyses a fixed (or mixed) effects parametric model is specified for the relationship between the model parameters and the covariates of interest. In this paper, we discuss the modeling of the relationship via the use of penalized splines, to allow for considerably more flexible functional forms. Corresponding models can be fitted via numerical maximum penalized likelihood estimation, employing cross-validation to choose the smoothing parameters in a data-driven way. Our contribution builds on and extends the existing literature, providing a unified inferential framework for semiparametric mark-recapture-recovery models for open populations, where the interest typically lies in the estimation of survival probabilities. The approach is applied to two real datasets, corresponding to grey herons (Ardea Cinerea), where we model the survival probability as a function of environmental condition (a time-varying global covariate), and Soay sheep (Ovis Aries), where we model the survival probability as a function of individual weight (a time-varying individual-specific covariate). The proposed semiparametric approach is compared to a standard parametric (logistic) regression and new interesting underlying dynamics are observed in both cases.

stat.AP

Semi-Markov Arnason-Schwarz models

We consider multi-state capture-recapture-recovery data where observed individuals are recorded in a set of possible discrete states. Traditionally, the Arnason-Schwarz model has been fitted to such data where the state process is modeled as a first-order Markov chain, though second-order models have also been proposed and fitted to data. However, low-order Markov models may not accurately represent the underlying biology. For example, specifying a (time-independent) first-order Markov process assumes that the dwell time in each state (i.e., the duration of a stay in a given state) has a geometric distribution, and hence that the modal dwell time is one. Specifying time-dependent or higher-order processes provides additional flexibility, but at the expense of a potentially significant number of additional model parameters. We extend the Arnason-Schwarz model by specifying a semi-Markov model for the state process, where the dwell-time distribution is specified more generally, using for example a shifted Poisson or negative binomial distribution. A state expansion technique is applied in order to represent the resulting semi-Markov Arnason-Schwarz model in terms of a simpler and computationally tractable hidden Markov model. Semi-Markov Arnason-Schwarz models come with only a very modest increase in the number of parameters, yet permit a significantly more flexible state process. Model selection can be performed using standard procedures, and in particular via the use of information criteria. The semi-Markov approach allows for important biological inference to be drawn on the underlying state process, for example on the times spent in the different states. The feasibility of the approach is demonstrated in a simulation study, before being applied to real data corresponding to house finches where the states correspond to the presence or absence of conjunctivitis.

stat.AP