SearcharxivSearch

arXiv subjects

Francesco Finazzi

Publications and source records attributed to Francesco Finazzi.

13 recordsLinked to original sources

Multivariate Low-Rank State-Space Model with SPDE Approach for High-Dimensional Data

This paper proposes a novel low-rank approximation to the multivariate State-Space Model. The Stochastic Partial Differential Equation (SPDE) approach is applied component-wise to the independent-in-time Mat\'ern Gaussian innovation term in the latent equation, assuming component independence. This results in a sparse representation of the latent process on a finite element mesh, allowing for scalable inference through sparse matrix operations. Dependencies among observed components are introduced through a matrix of weights applied to the latent process. Model parameters are estimated using the Expectation-Maximisation algorithm, which features closed-form updates for most parameters and efficient numerical routines for the remaining parameters. We prove theoretical results regarding the accuracy and convergence of the SPDE-based approximation under fixed-domain asymptotics. Simulation studies show our theoretical results. We include an empirical application on air quality to demonstrate the practical usefulness of the proposed model, which maintains computational efficiency in high-dimensional settings. In this application, we reduce computation time by about 93%, with only a 15% increase in the validation error.

stat.ME

Spatial Proportional Hazards Model with Differential Regularization

The Proportional Hazards (PH) model is one of the most widely used models in survival analysis, typically assuming a log-linear relationship between covariates and the hazard function. However, in the context of spatial survival data, where the time-to-event variable is associated with a spatial location within a given domain, this assumption is often unrealistic in capturing spatial effects. Thus, this paper proposes modeling the location effect through a nonparametric function of spatial location. The function is approximated using finite element methods on a triangulated mesh to accommodate irregular domains. Estimation is carried out within the classical partial likelihood framework, with smoothness of the spatial effect enforced through differential penalization. Using sieve methods, we establish the consistency and asymptotic normality of the parametric component. Simulations and two empirical applications demonstrate superior performance compared to existing approaches.

stat.ME

Scenario analysis of livestock-related PM2.5 pollution based on a new heteroskedastic spatiotemporal model

The air in the Lombardy Plain, Italy, is one of the most polluted in Europe due to limited atmosphere circulation and high emission levels. There is broad scientific consensus that ammonia (NH$_3$) emissions have a primary impact on air quality, and, in Lombardy, the agricultural sector and livestock activities are widely recognised as being responsible for approximately 97% of regional ammonia emissions due to the high density of livestock. In this paper, we quantify the relationship between ammonia emissions and PM2.5 concentrations in the Lombardy Plain and evaluate PM2.5 changes due to the reduction of ammonia emissions through a "what-if" scenario analysis. The information in the data is exploited using a spatiotemporal statistical model capable of handling spatial and temporal correlation, as well as missing data. To do this, we propose a new heteroskedastic extension of the well-established Hidden Dynamic Geostatistical Model. Maximum likelihood parameter estimates are obtained by the expectation-maximisation algorithm and implemented in a new version of the D-STEM software. Considering the years between 2016 and 2020, the scenario analysis is carried out on high-resolution PM2.5 maps of the Lombardy Plain. As a result, it is shown that a 26% reduction in NH3 emissions in the wintertime could reduce the PM2.5 average by 1.44 mg/m^3 while a 50% reduction could reduce the PM2.5 average by 2.76 mg / m^3 which corresponds to a reduction close to 3.6% and 7% respectively. Finally, results are detailed by province and land type.

stat.AP

Spatiotemporal modelling of PM$_{2.5}$ concentrations in Lombardy (Italy) -- A comparative study

This study presents a comparative analysis of three predictive models with an increasing degree of flexibility: hidden dynamic geostatistical models (HDGM), generalised additive mixed models (GAMM), and the random forest spatiotemporal kriging models (RFSTK). These models are evaluated for their effectiveness in predicting PM$_{2.5}$ concentrations in Lombardy (North Italy) from 2016 to 2020. Despite differing methodologies, all models demonstrate proficient capture of spatiotemporal patterns within air pollution data with similar out-of-sample performance. Furthermore, the study delves into station-specific analyses, revealing variable model performance contingent on localised conditions. Model interpretation, facilitated by parametric coefficient analysis and partial dependence plots, unveils consistent associations between predictor variables and PM$_{2.5}$ concentrations. Despite nuanced variations in modelling spatiotemporal correlations, all models effectively accounted for the underlying dependence. In summary, this study underscores the efficacy of conventional techniques in modelling correlated spatiotemporal data, concurrently highlighting the complementary potential of Machine Learning and classical statistical approaches.

stat.AP

Survival modelling of smartphone trigger data for earthquake parameter estimation in early warning. With applications to 2023 Turkish-Syrian and 2019 Ridgecrest events

Crowdsourced smartphone-based earthquake early warning systems recently emerged as reliable alternatives to the more expensive solutions based on scientific-grade instruments. For instance, during the 2023 Turkish-Syrian deadly event, the system implemented by the Earthquake Network citizen science initiative provided a forewarning up to 25 seconds. We develop a statistical methodology based on a survival mixture cure model which provides full Bayesian inference on epicentre, depth and origin time, and we design an efficient tempering MCMC algorithm to address multi-modality of the posterior distribution. The methodology is applied to data collected by the Earthquake Network, including the 2023 Turkish-Syrian and 2019 Ridgecrest events.

stat.AP

A simulation framework for statistical inference on the alerting capabilities of smartphone-based earthquake early warning systems. With a case study on the Earthquake Network system in Haiti

Smartphone-based earthquake early warning systems implemented by citizen science initiatives are characterized by a significant variability in their smartphone network geometry. This has an direct impact on the earthquake detection capability and performance of the system. Here, a simulation framework based on the Monte Carlo method is implemented for making inference on relevant quantities of the earthquake detection such as the detection distance from the epicentre, the detection delay and the warning time for people exposed to high ground shaking levels. The framework is applied to Haiti, which has experienced deadly earthquakes in the past decades, and to the network of the Earthquake Network citizen science initiative, which is popular in the country. It is discovered that relatively low penetrations of the initiative among the population allow to offer a robust early warning service, with warning times up to 12 second for people exposed to intensities between 7.5 and 8.5 of the modified Mercalli scale.

stat.AP

A statistical approach for controlling the probability of false alarm and missed detection in smartphone-based earthquake early warning systems

Smartphone-based earthquake early warning systems (EEWS) are emerging as a complementary solution to classic EEWS based on expensive scientific-grade instruments. Smartphone-based systems, however, are characterized by a highly dynamic network geometry and by noisy measurements. Thus the need to control the probability of false alarm and the probability of missed detection. This paper proposes a statistical approach based on the maximum likelihood method to address this challenge and to jointly estimate in near real-time earthquake parameters like epicentre and depth. The approach is tested using data coming from the Earthquake Network citizen science initiative which implements a global smartphone-based EEWS.

stat.AP

Agrimonia: a dataset on livestock, meteorology and air quality in the Lombardy region, Italy

The air in the Lombardy region, Italy, is one of the most polluted in Europe because of limited air circulation and high emission levels. There is a large scientific consensus that the agricultural sector has a significant impact on air quality. To support studies quantifying the role of the agricultural and livestock sectors on the Lombardy air quality, this paper presents a harmonised dataset containing daily values of air quality, weather, emissions, livestock, and land and soil use in the years 2016 - 2021, for the Lombardy region. The pollutant data come from the European Environmental Agency and the Lombardy Regional Environment Protection Agency, weather and emissions data from the European Copernicus programme, livestock data from the Italian zootechnical registry, and land and soil use data from the CORINE Land Cover project. The resulting dataset is designed to be used as is by those using air quality data for research.

stat.AP

Replacing discontinued Big Tech mobility reports: a penetration-based analysis

People mobility data sets played a role during the COVID-19 pandemic in assessing the impact of lockdown measures and correlating mobility with pandemic trends. Two global data sets were Apple's Mobility Trends Reports and Google's Community Mobility Reports. The former is no longer available, while the latter will be discontinued in October 2022. Thus, new products will be required. To establish a lower bound on data set penetration guaranteeing high adherence between new products and the Big Tech products, an independent mobility data set based on 3.8 million smartphone trajectories is analysed to compare its information content with that of the Google data set. This lower bound is determined to be between 10^-4 and 10^-3 (1 trajectory every 10,000 and 1000 people, respectively) suggesting that relatively small data sets are suitable for replacing Big Tech reports.

stat.AP

"Shaking in 5 seconds!" A Voluntary Smartphone-based Earthquake Early Warning System

Public earthquake early warning systems have the potential to reduce individual risk by warning people of an incoming tremor but their development has been hampered by costly infrastructure. Furthermore, users' understanding of such a service and their reactions to warnings remains poorly studied. The smartphone app of the Earthquake Network initiative turns users' smartphones into motion detectors and provides the first example of purely smartphone-based earthquake early warnings, without the need for dedicated seismic station infrastructure and operating in multiple countries. We demonstrate here that early warnings have been emitted in multiple countries even for damaging shaking levels and so this offers an alternative in the many regions unlikely to be covered by conventional early warning systems in the foreseeable future. We also show that although warnings are understood and appreciated by users, notably to get psychologically prepared, only a fraction take protective actions such as "drop, cover and hold".

physics.geo-ph

D-STEM v2: A Software for Modelling Functional Spatio-Temporal Data

Functional spatio-temporal data naturally arise in many environmental and climate applications where data are collected in a three-dimensional space over time. The MATLAB D-STEM v1 software package was first introduced for modelling multivariate space-time data and has been recently extended to D-STEM v2 to handle functional data indexed across space and over time. This paper introduces the new modelling capabilities of D-STEM v2 as well as the complexity reduction techniques required when dealing with large data sets. Model estimation, validation and dynamic kriging are demonstrated in two case studies, one related to ground-level air quality data in Beijing, China, and the other one related to atmospheric profile data collected globally through radio sounding.

stat.ME

Statistical harmonization and uncertainty assessment in the comparison of satellite and radiosonde climate variables

Satellite product validation is key to ensure the delivery of quality products for climate and weather applications. To do this, a fundamental step is the comparison with other instruments, such us radiosonde. This is specially true for Essential Climate Variables such as temperature and humidity. Thanks to a functional data representation, this paper uses a likelihood based approach which exploits the measurement uncertainties in a natural way. In particular the comparison of temperature and humdity radiosonde measurements collected within RAOB network and the corresponding atmospheric profiles derived from IASI interferometers aboard of Metop-A and Metop-B satellites is developed with the aim of understanding the vertical smoothing mismatch uncertainty. Moreover, conventional RAOB functional data representation is assessed by means of a comparison with radiosonde reference measurements given by GRUAN network, which provides high resolution fully traceable radiosouding profiles. In this way the uncertainty related to coarse vertical resolution, or sparseness, of conventional RAOB is assessed. It has been found that the uncertainty of vertical smoothing mismatch averaged along the profile is 0.50 K for temperature and 0.16 g/kg for water vapour mixing ratio. Moreover the uncertainty related to RAOB sparseness, averaged along the profile is 0.29 K for temperature and 0.13 g/kg for water vapour mixing ratio.

stat.AP

Geostatistical modeling in the presence of interaction between the measuring instruments, with an application to the estimation of spatial market potentials

This paper addresses the problem of recovering the spatial market potential of a retail product from spatially distributed sales data. In order to tackle the problem in a general way, the concept of spatial potential is introduced. The potential is concurrently measured at different spatial locations and the measurements are analyzed in order to recover the spatial potential. The measuring instruments used to collect the data interact with each other, that is, the measurement at a given spatial location is affected by the concurrent measurements at other locations. An approach based on a novel geostatistical model is developed. In particular, the model is able to handle both the measuring instrument interaction and the missing data. A model estimation procedure based on the expectation-maximization algorithm is provided as well as standard inferential tools. The model is applied to the estimation of the spatial market potential of a newspaper for the city of Bergamo, Italy. The estimated spatial market potential is eventually analyzed in order to identify the areas with the highest potential, to identify the areas where it is profitable to open additional newsstands and to evaluate the newspaper total market volume of the city.

stat.AP