SearcharxivSearch

arXiv subjects

Ottmar Cronie

Publications and source records attributed to Ottmar Cronie.

14 recordsLinked to original sources

New density/likelihood representations for Gibbs models based on generating functionals of point processes

Deriving exact density functions for Gibbs point processes has been challenging due to their general intractability, stemming from the intractability of their normalising constants/partition functions. This paper offers a solution to this open problem by exploiting a recent alternative representation of point process densities. Here, for a finite point process, the density is expressed as the void probability multiplied by a higher-order Papangelou conditional intensity function. By leveraging recent results on dependent thinnings, exact expressions for generating functionals and void probabilities of locally stable point processes are derived. Consequently, exact expressions for density/likelihood functions, partition functions and posterior densities are also obtained. The paper finally extends the results to locally stable Gibbsian random fields on lattices by representing them as point processes.

math.PR

Comparison of Point Process Learning and its special case Takacs-Fiksel estimation

Recently, Cronie et al. (2024) introduced the notion of cross-validation for point processes and a new statistical methodology called Point Process Learning (PPL). In PPL one splits a point process/pattern into a training and a validation set, and then predicts the latter from the former through a parametrised Papangelou conditional intensity. The model parameters are estimated by minimizing a point process prediction error; this notion was introduced as the second building block of PPL. It was shown that PPL outperforms the state-of-the-art in both kernel intensity estimation and estimation of the parameters of the Gibbs hard-core process. In the latter case, the state-of-the-art was represented by pseudolikelihood estimation. In this paper we study PPL in relation to Takacs-Fiksel estimation, of which pseudolikelihood is a special case. We show that Takacs-Fiksel estimation is a special case of PPL in the sense that PPL with a specific loss function asymptotically reduces to Takacs-Fiksel estimation if we let the cross-validation regime tend to leave-one-out cross-validation. Moreover, PPL involves a certain type of hyperparameter given by a weight function which ensures that the prediction errors have expectation zero if and only if we have the correct parametrisation. We show that the weight function takes an explicit but intractable form for general Gibbs models. Consequently, we propose different approaches to estimate the weight function in practice. In order to assess how the general PPL setup performs in relation to its special case Takacs-Fiksel estimation, we conduct a simulation study where we find that for common Gibbs models we can find loss functions and hyperparameters so that PPL typically outperforms Takacs-Fiksel estimation significantly in terms of mean square error. Here, the hyperparameters are the cross-validation parameters and the weight function estimate.

stat.ME

Semi-parametric profile pseudolikelihood via local summary statistics for spatial point pattern intensity estimation

Second-order statistics play a crucial role in analysing point processes. Previous research has specifically explored locally weighted second-order statistics for point processes, offering diagnostic tests in various spatial domains. However, there remains a need to improve inference for complex intensity functions, especially when the point process likelihood is intractable and in the presence of interactions among points. This paper addresses this gap by proposing a method that exploits local second-order characteristics to account for local dependencies in the fitting procedure. Our approach utilises the Papangelou conditional intensity function for general Gibbs processes, avoiding explicit assumptions about the degree of interaction and homogeneity. We provide simulation results and an application to real data to assess the proposed method's goodness-of-fit. Overall, this work contributes to advancing statistical techniques for point process analysis in the presence of spatial interactions.

stat.ME

Local inhomogeneous weighted summary statistics for marked point processes

We introduce a family of local inhomogeneous mark-weighted summary statistics, of order two and higher, for general marked point processes. Depending on how the involved weight function is specified, these summary statistics capture different kinds of local dependence structures. We first derive some basic properties and show how these new statistical tools can be used to construct most existing summary statistics for (marked) point processes. We then propose a local test of random labelling. This procedure allows us to identify points, and consequently regions, where the random labelling assumption does not hold, e.g.~when the (functional) marks are spatially dependent. Through a simulation study we show that the test is able to detect local deviations from random labelling. We also provide an application to an earthquake point pattern with functional marks given by seismic waveforms.

stat.ME

Spatial point process via regularisation modelling of ambulance call risk

This study investigates the spatial distribution of emergency alarm call events to identify spatial covariates associated with the events and discern hotspot regions for the events. The study is motivated by the problem of developing optimal dispatching strategies for prehospital resources such as ambulances. To achieve our goals, we model the spatially varying call occurrence risk as an intensity function of an inhomogeneous spatial Poisson process that we assume is a log-linear function of some underlying spatial covariates. The spatial covariates used in this study are related to road network coverage, population density, and the socio-economic status of the population in Skellefteå, Sweden. A new heuristic algorithm has been developed to select an optimal estimate of the kernel bandwidth in order to obtain the non-parametric intensity estimate of the events and to generate other covariates. Since we consider a large number of spatial covariates as well as their products, and since some of them may be strongly correlated, lasso-like elastic-net regularisation has been used in the log-likelihood intensity modeling to perform variable selection and reduce variance inflation from overfitting and bias from underfitting. As a result of the variable selection, the fitted model structure contains individual covariates of both road network and demographic types. We discovered that hotspot regions of calls have been observed along dense parts of the road network. Evaluation of the model also suggests that the estimated model is stable and can be used to generate a reliable intensity estimate over the region, which can be used as an input in the problem of designing prehospital resource dispatching strategies.

stat.AP

Statistical learning and cross-validation for point processes

This paper presents the first general (supervised) statistical learning framework for point processes in general spaces. Our approach is based on the combination of two new concepts, which we define in the paper: i) bivariate innovations, which are measures of discrepancy/prediction-accuracy between two point processes, and ii) point process cross-validation (CV), which we here define through point process thinning. The general idea is to carry out the fitting by predicting CV-generated validation sets using the corresponding training sets; the prediction error, which we minimise, is measured by means of bivariate innovations. Having established various theoretical properties of our bivariate innovations, we study in detail the case where the CV procedure is obtained through independent thinning and we apply our statistical learning methodology to three typical spatial statistical settings, namely parametric intensity estimation, non-parametric intensity estimation and Papangelou conditional intensity fitting. Aside from deriving theoretical properties related to these cases, in each of them we numerically show that our statistical learning approach outperforms the state of the art in terms of mean (integrated) squared error.

stat.ME

Large-scale modelling and forecasting of ambulance calls in northern Sweden using spatio-temporal log-Gaussian Cox processes

Although ambulance call data typically come in the form of spatio-temporal point patterns, point process-based modelling approaches presented in the literature are scarce. In this paper, we study a unique set of Swedish spatio-temporal ambulance call data, which consist of the spatial (GPS) locations of the calls (within the four northernmost regions of Sweden) and the associated days of occurrence of the calls (January 1, 2014, to December 31, 2018). Motivated by the nature of the data, we here employ log-Gaussian Cox processes (LGCPs) for the spatio-temporal modelling and forecasting of the calls. To this end, we propose a K-means clustering based bandwidth selection method for the kernel estimation of the spatial component of the separable spatio-temporal intensity function. The temporal component of the intensity function is modelled using Poisson regression, using different calendar covariates, and the spatio-temporal random field component of the random intensity of the LGCP is fitted using the Metropolis-adjusted Langevin algorithm. Spatial hot-spots have been found in the south-eastern part of the study region, where most people in the region live and our fitted model/forecasts manage to capture this behavior quite well. Also, there is a significant association between the expected number of calls and the day-of-the-week and the season-of-the-year. A non-parametric second-order analysis indicates that LGCPs seem to be reasonable models for the data. Finally, we find that the fitted forecasts generate simulated future spatial event patterns that quite well resemble the actual future data.

stat.AP

Functional marked point processes -- A natural structure to unify spatio-temporal frameworks and to analyse dependent functional data

This paper treats functional marked point processes (FMPPs), which are defined as marked point processes where the marks are random elements in some (Polish) function space. Such marks may represent e.g. spatial paths or functions of time. To be able to consider e.g. multivariate FMPPs, we also attach an additional, Euclidean, mark to each point. We indicate how FMPPs quite naturally connect the point process framework with both the functional data analysis framework and the geostatistical framework. We further show that various existing models fit well into the FMPP framework. In addition, we introduce a new family of summary statistics, weighted marked reduced moment measures, together with their non-parametric estimators, in order to study features of the functional marks. We further show how they generalise other summary statistics and we finally apply these tools to analyse population structures, such as demographic evolution and sex ratio over time, in Spanish provinces.

math.ST

Inhomogeneous higher-order summary statistics for linear network point processes

We introduce the notion of intensity reweighted moment pseudostationary point processes on linear networks. Based on arbitrary general regular linear network distances, we propose geometrically corrected versions of different higher-order summary statistics, including the inhomogeneous empty space function, the inhomogeneous nearest neighbour distance distribution function and the inhomogeneous $J$-function. We also discuss their non-parametric estimators. Through a simulation study, considering models with different types of spatial interaction, we study the performance of our proposed summary statistics. Finally, we make use of our methodology to analyse two datasets: motor vehicle traffic accidents and spider data.

stat.ME

Adaptive Algorithm for Sparse Signal Recovery

Spike and slab priors play a key role in inducing sparsity for sparse signal recovery. The use of such priors results in hard non-convex and mixed integer programming problems. Most of the existing algorithms to solve the optimization problems involve either simplifying assumptions, relaxations or high computational expenses. We propose a new adaptive alternating direction method of multipliers (AADMM) algorithm to directly solve the presented optimization problem. The algorithm is based on the one-to-one mapping property of the support and non-zero element of the signal. At each step of the algorithm, we update the support by either adding an index to it or removing an index from it and use the alternating direction method of multipliers to recover the signal corresponding to the updated support. Experiments on synthetic data and real-world images show that the proposed AADMM algorithm provides superior performance and is computationally cheaper, compared to the recently developed iterative convex refinement (ICR) algorithm.

stat.ME

Resample-smoothing of Voronoi intensity estimators

Voronoi intensity estimators, which are non-parametric estimators for intensity functions of point processes, are both parameter-free and adaptive; the intensity estimate at a given location is given by the reciprocal size of the Voronoi/Dirichlet cell containing that location. Their major drawback, however, is that they tend to under-smooth the data in regions where the point density of the observed point pattern is high and over-smooth in regions where the point density is low. To remedy this problem, i.e. to find some middle-ground between over- and under-smoothing, we propose an additional smoothing technique for Voronoi intensity estimators for point processes in arbitrary metric spaces, which is based on repeated independent thinnings of the point process/pattern. Through a simulation study we show that our resample-smoothing technique improves the estimation significantly. In addition, we study statistical properties such as unbiasedness and variance, and propose a rule-of-thumb and a data-driven cross-validation approach to choose the amount of thinning/smoothing to apply. We finally apply our proposed intensity estimation scheme to two datasets: locations of pine saplings (planar point pattern) and motor vehicle traffic accidents (linear network point pattern).

stat.ME

The second-order analysis of marked spatio-temporal point processes, with an application to earthquake data

To analyse interaction in marked spatio-temporal point processes (MSTPPs), we introduce marked (cross) second-order reduced moment measures and K-functions for general inhomogeneous second-order intensity reweighted stationary MSTPPs. These summary statistics, which allow us to quantify dependence between different mark categories of the points, are depending on the specific mark space and mark reference measure chosen. We also look closer at how the summary statistics reduce under assumptions such as the MSTPP being multivariate and/or stationary. A new test for independent marking is devised and unbiased minus-sampling estimators are derived for all statistics considered. In addition, we treat Voronoi intensity estimators for MSTPPs and indicate their unbiasedness. These new statistics are finally employed to analyse the well-known Andaman sea earthquake dataset. We find that clustering takes place between main and fore-/aftershocks at virtually all space and time scales. In addition, we find evidence that, conditionally on the space-time locations of the earthquakes, the magnitudes do not behave like an iid sequence.

stat.ME

Spatio-temporal càdlàg functional marked point processes: Unifying spatio-temporal frameworks

This paper defines the class of càdlàg functional marked point processes (CFMPPs). These are (spatio-temporal) point processes marked by random elements which take values in a càdlàg function space, i.e. the marks are given by càdlàg stochastic processes. We generalise notions of marked (spatio-temporal) point processes and indicate how this class, in a sensible way, connects the point process framework with the random fields framework. We also show how they can be used to construct a class of spatio-temporal Boolean models, how to construct different classes of these models by choosing specific mark functions, and how càdlàg functional marked Cox processes have a double connection to random fields. We also discuss finite CFMPPs, purely temporally well-defined CFMPPs and Markov CFMPPs. Furthermore, we define characteristics such as product densities, Palm distributions and conditional intensities, in order to develop statistical inference tools such as likelihood estimation schemes.

math.ST

Likelihood Inference for a Functional Marked Point Process with Cox-Ingersoll-Ross Process Marks

This paper considers maximum likelihood inference for a functional marked point process - the stochastic growth-interaction process - which is an extension of the spatio-temporal growth-interaction process to the stochastic mark setting. As a pilot study we here consider a particular version of this extended process, which has a homogenous Poisson process as unmarked point process and shifted independent Cox-Ingersoll-Ross processes as functional marks. These marks have supports determined by the lifetimes generated by an immigration-death process. By considering a (temporally) discrete sample scheme for the marks and by considering the process' alternative evolutionary representation as a multivariate diffusion (Markovian) with jumps, the likelihood function is expressed as a product of the process' closed form transition densities. Additionally, under the assumption that the mark processes are started in their common stationary distribution, and under some restrictions on the underlying parameters, consistency and asymptotic normality of the maximum likelihood (ML) estimators are proved. The ML-estimators derived from the stationarity assumption are then compared numerically to the ML-estimators derived under non-stationarity, in order to investigate the robustness of the stationarity assumption. To illustrate the model's use in forestry, it is fitted to a data set of Scots pines.

math.ST