SearcharxivSearch

arXiv subjects

Jorge Mateu

Publications and source records attributed to Jorge Mateu.

At least 19 recordsLinked to original sources

Bayesian Effect Selection for Additive Quantile Regression with an Application to Air Pollution Thresholds

Air pollution regulatory limits are typically defined in terms of exceedances of concentration thresholds which are naturally related to conditional quantiles of the pollutant distribution and are therefore of direct relevance for assessing severe pollution events. At the same time, it is important to determine not only whether a covariate affects air pollution but also whether this effect is linear, nonlinear, or both. We address these issues by developing a Bayesian effect selection approach for additive quantile regression. While commonly used mixed model representations (MMRs) of penalized splines allow for flexible nonlinear effects, they do not provide a meaningful separation of linear and nonlinear effect components. We therefore employ a Demmler-Reinsch basis expansion, which yields an orthogonal decomposition of each additive effect into linear and nonlinear parts and show theoretically that both effect components can be estimated consistently. To facilitate data-driven model building, we propose Bayesian effect selection with separate spike and slab priors on the scalar importance parameters associated with the linear and nonlinear components and implement an efficient Gibbs sampler. Through simulation studies, we demonstrate robustness to the misspecification induced by the employed asymmetric Laplace working likelihood and show superior performance relative to the MMR. In a detailed analysis of air pollution data in Madrid, Spain we highlight the added value of flexibly modeling extreme nitrogen dioxide (NO$_2$) concentrations and reveal that threshold-relevant pollution levels are driven differently by climatological variables and traffic-related spatial structure. These findings underline the need for advanced statistical models that support short-term decision-making and help local authorities mitigate, or potentially prevent, exceedances of NO$_2$ concentration limits.

stat.ME

Robust Nonparametric Testing Approaches for Spatial Regression

Reliable inference for spatial regression remains challenging because it requires the correct specification of the spatial dependence structure, the mean trend, and the error distribution. Existing parametric testing methods rely on restrictive assumptions that are difficult to verify in practice and can lead to inaccurate conclusions under misspecification. To address this, we develop a robust nonparametric Monte Carlo testing framework for spatial regression based on random shifts. We construct test statistics that measure the dependence between residuals, obtained after removing the effects of nuisance covariates, and the covariate of interest. This allows us to assess the significance of the covariate in the sense of partial correlation. The proposed framework enables robust inference across various models without requiring parametric assumptions or even a closed-form distribution of the test statistics. Furthermore, we establish the asymptotic exactness of the random shift test in the increasing-domain setting when the sample covariance is used as the test statistic. Through extensive numerical experiments, we demonstrate that our method maintains the nominal significance level while achieving competitive power, whereas parametric methods can exhibit inflated type I error rates, even when they are correctly specified.

stat.ME

Effective Sample Size for Functional Spatial Data

The effective sample size quantifies the amount of independent information contained in a dataset, accounting for redundancy due to correlation between observations. While widely used in geostatistics for scalar data, its extension to functional spatial data has remained largely unexplored. In this work, we introduce a novel definition of the effective sample size for functional geostatistical data, employing the trace-covariogram as a measure of correlation, and show that it retains the intuitive properties of the classical scalar ESS. We illustrate the behavior of this measure using a functional autoregressive process, demonstrating how serial dependence and the allocation of variability across eigen-directions influence the resulting functional ESS. Finally, the approach is applied to a real meteorological dataset of geometric vertical velocities over a portion of the Earth, showing how the method can quantify redundancy and determine the effective number of independent curves in functional spatial datasets.

stat.ME

Spatio-temporal Hawkes point processes: statistical inference and simulation strategies

Spatio-temporal Hawkes point processes are a particularly interesting class of stochastic point processes for modeling self-exciting behavior, in which the occurrence of one event increases the probability of other events occurring. These processes are able to handle complex interrelationships between stochastic and deterministic components of spatio-temporal phenomena. However, despite its widespread use in practice, there is no common and unified formalism and every paper proposes different views of these stochastic mechanisms. With this in mind, we implement two simulation techniques and three unified, self-consistent inference techniques, which are widely used in the practical modeling of spatio-temporal Hawkes processes. Furthermore, we provide an evaluation of the practical performance of these methods, while providing useful code for reproducibility.

stat.CO

A concordance coefficient for lattice data: An application to poverty indices in Chile

This paper introduces a novel coefficient for measuring agreement between two lattice sequences observed in the same areal units, motivated by the analysis of different methodologies for measuring poverty rates in Chile. Building on the multivariate concordance coefficient framework, our approach accounts for dependencies in the multivariate lattice process using a non-negative definite matrix of weights, assuming a Multivariate Conditionally Autoregressive (GMCAR) process. We adopt a Bayesian perspective for inference, using summaries from Bayesian estimates. The methodology is illustrated through an analysis of poverty rates in the Metropolitan and Valpara\'iso regions of Chile, with High Posterior Density (HPD) intervals provided for the poverty rates. This work addresses a methodological gap in the understanding of agreement coefficients and enhances the usability of these measures in the context of social variables typically assessed in areal units.

stat.ME

A Spatio-Temporal Dirichlet Process Mixture Model on Linear Networks for Crime Data

Analyzing crime events is crucial to understand crime dynamics and it is largely helpful for constructing prevention policies. Point processes specified on linear networks can provide a more accurate description of crime incidents by considering the geometry of the city. We propose a spatio-temporal Dirichlet process mixture model on a linear network to analyze crime events in Valencia, Spain. We propose a Bayesian hierarchical model with a Dirichlet process prior to automatically detect space-time clusters of the events and adopt a convolution kernel estimator to account for the network structure in the city. From the fitted model, we provide crime hotspot visualizations that can inform social interventions to prevent crime incidents. Furthermore, we study the relationships between the detected cluster centers and the city's amenities, which provides an intuitive explanation of criminal contagion.

stat.AP

Spatio-Temporal-Network Point Processes for Modeling Crime Events with Landmarks

Self-exciting point processes are widely used to model the contagious effects of crime events living within continuous geographic space, using their occurrence time and locations. However, in urban environments, most events are naturally constrained within the city's street network structure, and the contagious effects of crime are governed by such a network geography. Meanwhile, the complex distribution of urban infrastructures also plays an important role in shaping crime patterns across space. We introduce a novel spatio-temporal-network point process framework for crime modeling that integrates these urban environmental characteristics by incorporating self-attention graph neural networks. Our framework incorporates the street network structure as the underlying event space, where crime events can occur at random locations on the network edges. To realistically capture criminal movement patterns, distances between events are measured using street network distances. We then propose a new mark for a crime event by concatenating the event's crime category with the type of its nearby landmark, aiming to capture how the urban design influences the mixing structures of various crime types. A graph attention network architecture is adopted to learn the existence of mark-to-mark interactions. Extensive experiments on crime data from Valencia, Spain, demonstrate the effectiveness of our framework in understanding the crime landscape and forecasting crime risks across regions.

stat.AP

A point process approach for the classification of noisy calcium imaging data

We study noisy calcium imaging data, with a focus on the classification of spike traces. As raw traces obscure the true temporal structure of neuron's activity, we performed a tuned filtering of the calcium concentration using two methods: a biophysical model and a kernel mapping. The former characterizes spike trains related to a particular triggering event, while the latter filters out the signal and refines the selection of the underlying neuronal response. Transitioning from traditional time series analysis to point process theory, the study explores spike-time distance metrics and point pattern prototypes to describe repeated observations. We assume that the analyzed neuron's firing events, i.e. spike occurrences, are temporal point process events. In particular, the study aims to categorize 47 point patterns by depth, assuming the similarity of spike occurrences within specific depth categories. The results highlight the pivotal roles of depth and stimuli in discerning diverse temporal structures of neuron firing events, confirming the point process approach based on prototype analysis is largely useful in the classification of spike traces.

q-bio.NC

Estimating velocities of infectious disease spread through spatio-temporal log-Gaussian Cox point processes

Understanding the spread of infectious diseases such as COVID-19 is crucial for informed decision-making and resource allocation. A critical component of disease behavior is the velocity with which disease spreads, defined as the rate of change between time and space. In this paper, we propose a spatio-temporal modeling approach to determine the velocities of infectious disease spread. Our approach assumes that the locations and times of people infected can be considered as a spatio-temporal point pattern that arises as a realization of a spatio-temporal log-Gaussian Cox process. The intensity of this process is estimated using fast Bayesian inference by employing the integrated nested Laplace approximation (INLA) and the Stochastic Partial Differential Equations (SPDE) approaches. The velocity is then calculated using finite differences that approximate the derivatives of the intensity function. Finally, the directions and magnitudes of the velocities can be mapped at specific times to examine better the spread of the disease throughout the region. We demonstrate our method by analyzing COVID-19 spread in Cali, Colombia, during the 2020-2021 pandemic.

stat.AP

Function-valued marked spatial point processes on linear networks: application to urban cycling profiles

In the literature on spatial point processes, there is an emerging challenge in studying marked point processes with points being labelled by functions. In this paper, we focus on point processes living on linear networks and, from distinct points of view, propose several marked summary characteristics that are of great use in studying the average association and dispersion of the function-valued marks. Through a simulation study, we evaluate the performance of our proposed marked summary characteristics, both when marks are independent and when some sort of spatial dependence is evident among them. Finally, we employ our proposed mark summary characteristics to study the spatial structure of urban cycling profiles in Vancouver, Canada.

stat.ME

Semi-parametric profile pseudolikelihood via local summary statistics for spatial point pattern intensity estimation

Second-order statistics play a crucial role in analysing point processes. Previous research has specifically explored locally weighted second-order statistics for point processes, offering diagnostic tests in various spatial domains. However, there remains a need to improve inference for complex intensity functions, especially when the point process likelihood is intractable and in the presence of interactions among points. This paper addresses this gap by proposing a method that exploits local second-order characteristics to account for local dependencies in the fitting procedure. Our approach utilises the Papangelou conditional intensity function for general Gibbs processes, avoiding explicit assumptions about the degree of interaction and homogeneity. We provide simulation results and an application to real data to assess the proposed method's goodness-of-fit. Overall, this work contributes to advancing statistical techniques for point process analysis in the presence of spatial interactions.

stat.ME

Summary characteristics for multivariate function-valued spatial point process attributes

Prompted by modern technologies in data acquisition, the statistical analysis of spatially distributed function-valued quantities has attracted a lot of attention in recent years. In particular, combinations of functional variables and spatial point processes yield a highly challenging instance of such modern spatial data applications. Indeed, the analysis of spatial random point configurations, where the point attributes themselves are functions rather than scalar-valued quantities, is just in its infancy, and extensions to function-valued quantities still remain limited. In this view, we extend current existing first- and second-order summary characteristics for real-valued point attributes to the case where in addition to every spatial point location a set of distinct function-valued quantities are available. Providing a flexible treatment of more complex point process scenarios, we build a framework to consider points with multivariate function-valued marks, and develop sets of different cross-function (cross-type and also multi-function cross-type) versions of summary characteristics that allow for the analysis of highly demanding modern spatial point process scenarios. We consider estimators of the theoretical tools and analyse their behaviour through a simulation study and two real data applications.

stat.ME

A non-separable first-order spatio-temporal intensity for events on linear networks: an application to ambulance interventions

The algorithms used for the optimal management of an ambulance fleet require an accurate description of the spatio-temporal evolution of the emergency events. In the last years, several authors have proposed sophisticated statistical approaches to forecast ambulance dispatches, typically modelling the data as a point pattern occurring on a planar region. Nevertheless, ambulance interventions can be more appropriately modelled as a realisation of a point process occurring on a linear network. The constrained spatial domain raises specific challenges and unique methodological problems that cannot be ignored when developing a proper statistical approach. Hence, this paper proposes a spatio-temporal model to analyse ambulance dispatches focusing on the interventions that occurred in the road network of Milan (Italy) from 2015 to 2017. We adopt a non-separable first-order intensity function with spatial and temporal terms. The temporal dimension is estimated semi-parametrically using a Poisson regression model, while the spatial dimension is estimated non-parametrically using a network kernel function. A set of weights is included in the spatial term to capture space-time interactions, inducing non-separability in the intensity function. A series of tests show that our approach successfully models the ambulance interventions and captures the space-time patterns more accurately than planar or separable point process models.

stat.AP

Measurement Error Models for Spatial Network Lattice Data: Analysis of Car Crashes in Leeds

Road casualties represent an alarming concern for modern societies. During the last years, several authors proposed sophisticated approaches to help authorities implement new policies. These models were usually developed considering a set of socioeconomic variables and ignoring the measurement error, which can bias the statistical inference. This paper presents a Bayesian model to analyse car crashes occurrences at the network-lattice level, taking into account measurement error in the spatial covariate. The suggested methodology is exemplified by considering the collisions in the road network of Leeds (UK) during 2011-2019. Traffic volumes are approximated using an extensive set of counts obtained from mobile devices and the estimates are adjusted using a spatial measurement error correction.

stat.AP

Non-stationary spatio-temporal point process modeling for high-resolution COVID-19 data

Most COVID-19 studies commonly report figures of the overall infection at a state- or county-level. This aggregation tends to miss out on fine details of virus propagation. In this paper, we analyze a high-resolution COVID-19 dataset in Cali, Colombia, that records the precise time and location of every confirmed case. We develop a non-stationary spatio-temporal point process equipped with a neural network-based kernel to capture the heterogeneous correlations among COVID-19 cases. The kernel is carefully crafted to enhance expressiveness while maintaining model interpretability. We also incorporate some exogenous influences imposed by city landmarks. Our approach outperforms the state-of-the-art in forecasting new COVID-19 cases with the capability to offer vital insights into the spatio-temporal interaction between individuals concerning the disease spread in a metropolis.

stat.AP

Locally weighted minimum contrast estimation for spatio-temporal log-Gaussian Cox processes

We propose a local version of spatio-temporal log-Gaussian Cox processes using Local Indicators of Spatio-Temporal Association (LISTA) functions into the minimum contrast procedure to obtain space as well as time-varying parameters. We resort to the joint minimum contrast fitting method to estimate the set of second-order parameters. This approach has the advantage of being suitable in both separable and non-separable parametric specifications of the correlation function of the underlying Gaussian Random Field. We present simulation studies to assess the performance of the proposed fitting procedure, and show an application to seismic spatio-temporal point pattern data.

stat.ME

Local inhomogeneous weighted summary statistics for marked point processes

We introduce a family of local inhomogeneous mark-weighted summary statistics, of order two and higher, for general marked point processes. Depending on how the involved weight function is specified, these summary statistics capture different kinds of local dependence structures. We first derive some basic properties and show how these new statistical tools can be used to construct most existing summary statistics for (marked) point processes. We then propose a local test of random labelling. This procedure allows us to identify points, and consequently regions, where the random labelling assumption does not hold, e.g.~when the (functional) marks are spatially dependent. Through a simulation study we show that the test is able to detect local deviations from random labelling. We also provide an application to an earthquake point pattern with functional marks given by seismic waveforms.

stat.ME

Clustering constrained on linear networks

An unsupervised classification method for point events occurring on a network of lines is proposed. The idea relies on the distributional flexibility and practicality of random partition models to discover the clustering structure featuring observations from a particular phenomenon taking place on a given set of edges. By incorporating the spatial effect in the random partition distribution, induced by a Dirichlet process, one is able to control the distance between edges and events, thus leading to an appealing clustering method. A Gibbs sampler algorithm is proposed and evaluated with a sensitivity analysis. The proposal is motivated and illustrated by the analysis of crime and violence patterns in Mexico City.

stat.ME