SearcharxivSearch

arXiv subjects

Paolo Maranzano

Publications and source records attributed to Paolo Maranzano.

15 recordsLinked to original sources

Small Area Estimation under Spatial Regimes: Spatially Clustered Fay-Herriot Models for Agricultural Indicators

Area-level small area estimation (SAE) models, such as the Fay--Herriot (FH) model, borrow strength across domains through covariates and random effects, but they can struggle when the relationship between the covariates and the outcome is spatially heterogeneous, that is, when it changes across the spatial domain of interest. We propose a spatially-clustered FH (SC-FH) framework that simultaneously (i) estimates cluster-specific regression coefficients and random effects variances and (ii) generates spatially coherent partitions of the geographical domain. Estimation maximizes a penalized likelihood that augments the FH likelihood with a Potts-type spatial cohesion term over the areal adjacency graph, through an efficient strategy that alternates between sequential label updates and closed-form FH updates within clusters. In simulation experiments run on the real geography of the application, the method recovers the latent regimes almost exactly whenever they are separated in the covariate--response space and improves prediction accuracy over the standard FH benchmark, with the spatial penalty acting as a stabilizer of both classification and estimation. An empirical application to the average standard output of farms in the Po Valley (Northern Italy) identifies two spatially compact production regimes with significantly different cluster-wise coefficients, and shows that the clusterwise predictor improves on the direct estimates while avoiding the over-shrinkage of the pooled model.

stat.AP

A new framework for non-stationary spatio-temporal data fusion of multi-fidelity models

We propose a new scalable framework for spatio-temporal data fusion with multi-fidelity Gaussian processes (MFGPs) that enables fully likelihood-based inference for both stationary and non-stationary fidelity integration. The framework is designed for environmental applications, where abundant but noisy low-fidelity data (e.g., satellite or reanalysis products) must be fused with sparse yet accurate high-fidelity in-situ observations to obtain high-resolution reconstructions. Our key methodological contribution is a decomposed multi-fidelity covariance formulation that allows the Vecchia approximation to be applied directly to the latent low-fidelity and discrepancy processes. Combined with a Woodbury-based reconstruction, this yields a numerically stable and computationally efficient evaluation of the joint marginal likelihood without ever forming the full multi-fidelity covariance matrix. In addition, we introduce a generalized least squares (GLS) mean-removal strategy with fidelity-specific offsets, preventing systematic biases from being absorbed into cross-fidelity dependence. We validate the proposed approach through extensive experiments on synthetic data and a large-scale real-world application to wind speed reconstruction in the Lombardy region of Italy. The results show that the proposed Vecchia-based MFGP closely matches exact multi-fidelity inference in controlled settings, while substantially outperforming standard single-fidelity spatio-temporal Gaussian processes in terms of predictive accuracy, correlation, and representation of local variability in realistic large-data scenarios.

stat.CO

On the use of satellite information to estimate agricultural carbon footprint in a small area framework

The agricultural sector is undergoing rapid change due to climate pressures, demographic shifts, and uneven economic development, increasing the demand for reliable environmental indicators at fine spatial scales. However, limited data availability often constrains subregional analyses. This study develops a model-based framework for producing reliable small-area estimates for assessing the agricultural carbon footprint in the Po Valley (Northern Italy), a region characterized by intensive livestock farming and high environmental pressure. We integrate survey, census, and satellite-derived emission data into a unified framework and produce estimates at the level of Agrarian Subregions, defined as agriculturally homogeneous municipalities by the Italian National Institute of Statistics. Satellite-based ammonia emission data are incorporated as auxiliary covariates to improve precision and spatial coherence. A key methodological contribution is the treatment of spatial misalignment between gridded satellite data and administrative boundaries. This issue is addressed through a geostatistical upscaling procedure combined with a parametric bootstrap that propagates uncertainty from the covariate construction stage to the final small-area estimates. The results show that satellite-derived information substantially improves the accuracy and stability of carbon footprint estimates while reducing reliance on large, heterogeneous auxiliary datasets, illustrating the potential of Earth observation data in model-based environmental statistics.

stat.AP

SCARFACE: a harmonized spatio-temporal dataset integrating socio-economic, environmental, and agricultural indicators for the Po Valley (Italy), 2011--2024

We present "Sequestering CARbon through Forests, AgriCulture, and land usE (SCARFACE)", a harmonized spatio-temporal dataset that integrates climate, air quality, airborne pollutant emissions, land cover, soil properties, agro-industry dynamics and socio-economic indicators, to jointly investigate interconnected processes linking agricultural systems, atmospheric dynamics, emissions, and socioeconomic conditions in the Po Valley, Northern Italy. The spatial reference unit is the Agrarian Sub Region (ASR), that is, groups of contiguous municipalities that are considered homogeneous with respect to physical geography, agronomic characteristics, and prevailing agricultural production systems. The dataset adopts an annual panel structure from 2011 to 2024 defined over the 256 ASRs partitioning the Po Valley and comprises more than 2,700 indicators sourced from national and international public institutions. Heterogeneous data are harmonized within a processing workflow, tailored to the specific characteristics of each dataset, that guarantee spatial and temporal consistency of the output dataset. The resource supports reuse in applied econometrics, spatio-temporal modeling, clustering, and policy analysis focused on agriculture, air quality, and land use in a major European hotspot.

stat.AP

Spatiotemporal clustering of GHGs emissions in Europe: exploring the role of spatial component

In this study, we propose a novel application of spatiotemporal clustering in the environmental sciences, with a particular focus on regionalised time series of greenhouse gases (GHGs) emissions from a range of economic sectors. Utilising a hierarchical spatiotemporal clustering methodology, we analyse yearly time series of emissions by gases and sectors from 1990 to 2022 for European regions at the NUTS-2 level. While the clustering algorithm inherently incorporates spatial information based on geographical distance, the extent to which space contributes to the definition of groups still requires further exploration. To address this gap in the literature, we propose a novel indicator, namely the Joint Inertia, which quantifies the contribution of spatial distances when integrated with other features. Through a simulation experiment, we explore the relationship between the Joint Inertia and the relevance of geography in exploiting the groups structure under several configurations of spatial and features patterns, providing insights into the behaviour and potential of the proposed indicator. The empirical findings demonstrate the relevance of the spatial component in identifying emission patterns and dynamics, and the results reveal significant heterogeneity across clusters in trends and dynamics by gases and sectors. This reflects the heterogeneous economic and industrial characteristics of European regions. The study highlights the importance of the spatial and temporal dimensions in understanding GHGs emissions, offering baseline insights for future spatiotemporal modelling and supporting more targeted and regionally informed environmental policies.

stat.AP

Mapping climate change awareness through spatial hierarchical clustering

Climate change is a critical issue that will be in the political agenda for the next decades. While it is important for this topic to be discussed at higher levels, it is also of paramount importance that the populations became aware of the problem. As different countries may face more or less severe repercussions, it is also useful to understand the degree of awareness of specific populations. In this paper, we present a geographically-informed hierarchical clustering analysis aimed at identify groups of countries with a similar level of climate change awareness. We employ a Ward-like clustering algorithm that combines information pertaining climate change awareness, socio-economic factors, climate-related characteristics of different countries, and the physical distances between countries. To choose suitable values for the clustering hyperparameters, we propose a customized algorithm that takes into account the within-clusters homogeneity, the between-clusters separation and that explicitly compares the geographically-informed and non-geographical partitioning. The results show that the geographically-informed clustering provides more stability of the partitions and leads to interpretable and geographically-compact aggregations compared to a clustering in which the geographical component is absent. In particular, we identify a clear contrast among Western countries, characterized by high and compact awareness, and Asian, African, and Middle Eastern countries having greater variability but still lower awareness.

stat.AP

Inequality and Concentration in Farmland Production and Size: Regional Analysis for the European Union from 2010 to 2020

According to Eurostat estimates, the overall number of farms in Europe declined of about 3 million units between 2010 and 2020. Parallel, the agricultural standard output increased from 304 billion to nearly 360 billion over the same period. Such evidence, legitimately leads to questions about how the structure (e.g., type of production and average size) of farms has changed and whether this change has been uniform or heterogeneous within Europe. In this paper, we aim at investigating the phenomenon of market concentration in the European agricultural and livestock farming industry from 2010 to 2020 at the regional level by exploiting the spatio-temporal dynamics of the Gini concentration index for the land owned by the European farmers and for their standard output. In particular, we are interested in exploring the variability within-and-between regions with regard to land and production size to assess if the European agricultural market suffered from an increasingly concentration of power in fewer but larger farm holding. The extensive mapping provided by this study may allow a fine spatial-scale socio-economic and political assessment of the European agricultural market integration process, its recent and future trends in the complex and uncertain post-COVID context and the restructuring of international relations due to crises and the green energy transition.

physics.soc-ph

Warped multifidelity Gaussian processes for data fusion of skewed environmental data

Understanding the dynamics of climate variables is paramount for numerous sectors, like energy and environmental monitoring. This study focuses on the critical need for a precise mapping of environmental variables for national or regional monitoring networks, a task notably challenging when dealing with skewed data. To address this issue, we propose a novel data fusion approach, the \textit{warped multifidelity Gaussian process} (WMFGP). The method performs prediction using multiple time-series, accommodating varying reliability and resolutions and effectively handling skewness. In an extended simulation experiment the benefits and the limitations of the methods are explored, while as a case study, we focused on the wind speed monitored by the network of ARPA Lombardia, one of the regional environmental agencies operting in Italy. ARPA grapples with data gaps, and due to the connection between wind speed and air quality, it struggles with an effective air quality management. We illustrate the efficacy of our approach in filling the wind speed data gaps through two extensive simulation experiments. The case study provides more informative wind speed predictions crucial for predicting air pollutant concentrations, enhancing network maintenance, and advancing understanding of relevant meteorological and climatic phenomena.

stat.AP

Spatially-clustered spatial autoregressive models with application to agricultural market concentration in Europe

In this paper, we present an extension of the spatially-clustered linear regression models, namely, the spatially-clustered spatial autoregression (SCSAR) model, to deal with spatial heterogeneity issues in clustering procedures. In particular, we extend classical spatial econometrics models, such as the spatial autoregressive model, the spatial error model, and the spatially-lagged model, by allowing the regression coefficients to be spatially varying according to a cluster-wise structure. Cluster memberships and regression coefficients are jointly estimated through a penalized maximum likelihood algorithm which encourages neighboring units to belong to the same spatial cluster with shared regression coefficients. Motivated by the increase of observed values of the Gini index for the agricultural production in Europe between 2010 and 2020, the proposed methodology is employed to assess the presence of local spatial spillovers on the market concentration index for the European regions in the last decade. Empirical findings support the hypothesis of fragmentation of the European agricultural market, as the regions can be well represented by a clustering structure partitioning the continent into three-groups, roughly approximated by a division among Western, North Central and Southeastern regions. Also, we detect heterogeneous local effects induced by the selected explanatory variables on the regional market concentration. In particular, we find that variables associated with social, territorial and economic relevance of the agricultural sector seem to act differently throughout the spatial dimension, across the clusters and with respect to the pooled model, and temporal dimension.

stat.ME

Multidimensional spatiotemporal clustering -- An application to environmental sustainability scores in Europe

The assessment of corporate sustainability performance is extremely relevant in facilitating the transition to a green and low-carbon intensity economy. However, companies located in different areas may be subject to different sustainability and environmental risks and policies. Henceforth, the main objective of this paper is to investigate the spatial and temporal pattern of the sustainability evaluations of European firms. We leverage on a large dataset containing information about companies' sustainability performances, measured by MSCI ESG ratings, and geographical coordinates of firms in Western Europe between 2013 and 2023. By means of a modified version of the Chavent et al. (2018) hierarchical algorithm, we conduct a spatial clustering analysis, combining sustainability and spatial information, and a spatiotemporal clustering analysis, which combines the time dynamics of multiple sustainability features and spatial dissimilarities, to detect groups of firms with homogeneous sustainability performance. We are able to build cross-national and cross-industry clusters with remarkable differences in terms of sustainability scores. Among other results, in the spatio-temporal analysis, we observe a high degree of geographical overlap among clusters, indicating that the temporal dynamics in sustainability assessment are relevant within a multidimensional approach. Our findings help to capture the diversity of ESG ratings across Western Europe and may assist practitioners and policymakers in evaluating companies facing different sustainability-linked risks in different areas.

stat.AP

A review of regularised estimation methods and cross-validation in spatiotemporal statistics

This review article focuses on regularised estimation procedures applicable to geostatistical and spatial econometric models. These methods are particularly relevant in the case of big geospatial data for dimensionality reduction or model selection. To structure the review, we initially consider the most general case of multivariate spatiotemporal processes (i.e., $g > 1$ dimensions of the spatial domain, a one-dimensional temporal domain, and $q \geq 1$ random variables). Then, the idea of regularised/penalised estimation procedures and different choices of shrinkage targets are discussed. Finally, guided by the elements of a mixed-effects model setup, which allows for a variety of spatiotemporal models, we show different regularisation procedures and how they can be used for the analysis of geo-referenced data, e.g. for selection of relevant regressors, dimensionality reduction of the covariance matrices, detection of conditionally independent locations, or the estimation of a full spatial interaction matrix.

stat.ME

Spatiotemporal modelling of PM$_{2.5}$ concentrations in Lombardy (Italy) -- A comparative study

This study presents a comparative analysis of three predictive models with an increasing degree of flexibility: hidden dynamic geostatistical models (HDGM), generalised additive mixed models (GAMM), and the random forest spatiotemporal kriging models (RFSTK). These models are evaluated for their effectiveness in predicting PM$_{2.5}$ concentrations in Lombardy (North Italy) from 2016 to 2020. Despite differing methodologies, all models demonstrate proficient capture of spatiotemporal patterns within air pollution data with similar out-of-sample performance. Furthermore, the study delves into station-specific analyses, revealing variable model performance contingent on localised conditions. Model interpretation, facilitated by parametric coefficient analysis and partial dependence plots, unveils consistent associations between predictor variables and PM$_{2.5}$ concentrations. Despite nuanced variations in modelling spatiotemporal correlations, all models effectively accounted for the underlying dependence. In summary, this study underscores the efficacy of conventional techniques in modelling correlated spatiotemporal data, concurrently highlighting the complementary potential of Machine Learning and classical statistical approaches.

stat.AP

Spatio-temporal Event Studies for Air Quality Assessment under Cross-sectional Dependence

Event Studies (ES) are statistical tools that assess whether a particular event of interest has caused changes in the level of one or more relevant time series. We are interested in ES applied to multivariate time series characterized by high spatial (cross-sectional) and temporal dependence. We pursue two goals. First, we propose to extend the existing taxonomy on ES, mainly deriving from the financial field, by generalizing the underlying statistical concepts and then adapting them to the time series analysis of airborne pollutant concentrations. Second, we address the spatial cross-sectional dependence by adopting a twofold adjustment. Initially, we use a linear mixed spatio-temporal regression model (HDGM) to estimate the relationship between the response variable and a set of exogenous factors, while accounting for the spatio-temporal dynamics of the observations. Later, we apply a set of sixteen ES test statistics, both parametric and nonparametric, some of which directly adjusted for cross-sectional dependence. We apply ES to evaluate the impact on NO2 concentrations generated by the lockdown restrictions adopted in the Lombardy region (Italy) during the COVID-19 pandemic in 2020. The HDGM model distinctly reveals the level shift caused by the event of interest, while reducing the volatility and isolating the spatial dependence of the data. Moreover, all the test statistics unanimously suggest that the lockdown restrictions generated significant reductions in the average NO2 concentrations.

stat.AP

Agrimonia: a dataset on livestock, meteorology and air quality in the Lombardy region, Italy

The air in the Lombardy region, Italy, is one of the most polluted in Europe because of limited air circulation and high emission levels. There is a large scientific consensus that the agricultural sector has a significant impact on air quality. To support studies quantifying the role of the agricultural and livestock sectors on the Lombardy air quality, this paper presents a harmonised dataset containing daily values of air quality, weather, emissions, livestock, and land and soil use in the years 2016 - 2021, for the Lombardy region. The pollutant data come from the European Environmental Agency and the Lombardy Regional Environment Protection Agency, weather and emissions data from the European Copernicus programme, livestock data from the Italian zootechnical registry, and land and soil use data from the CORINE Land Cover project. The resulting dataset is designed to be used as is by those using air quality data for research.

stat.AP

Adaptive LASSO estimation for functional hidden dynamic geostatistical model

We propose a novel model selection algorithm based on a penalized maximum likelihood estimator (PMLE) for functional hidden dynamic geostatistical models (f-HDGM). These models employ a classic mixed-effect regression structure with embedded spatiotemporal dynamics to model georeferenced data observed in a functional domain. Thus, the parameters of interest are functions across this domain. The algorithm simultaneously selects the relevant spline basis functions and regressors that are used to model the fixed-effects relationship between the response variable and the covariates. In this way, it automatically shrinks to zero irrelevant parts of the functional coefficients or the entire effect of irrelevant regressors. The algorithm is based on iterative optimisation and uses an adaptive least absolute shrinkage and selector operator (LASSO) penalty function, wherein the weights are obtained by the unpenalised f-HDGM maximum-likelihood estimators. The computational burden of maximisation is drastically reduced by a local quadratic approximation of the likelihood. Through a Monte Carlo simulation study, we analysed the performance of the algorithm under different scenarios, including strong correlations among the regressors. We showed that the penalised estimator outperformed the unpenalised estimator in all the cases we considered. We applied the algorithm to a real case study in which the recording of the hourly nitrogen dioxide concentrations in the Lombardy region in Italy was modelled as a functional process with several weather and land cover covariates.

stat.ME