Searcharxiv⌕ Search

arXiv subjects

Giuseppe Arbia

Publications and source records attributed to Giuseppe Arbia.

15 recordsLinked to original sources

Sampled Grid Pairwise Likelihood (SG-PL): An Efficient Approach for Spatial Regression on Large Data

Estimating spatial regression models on large, irregularly structured datasets poses significant computational hurdles. While Pairwise Likelihood (PL) methods offer a pathway to simplify these estimations, the efficient selection of informative observation pairs remains a critical challenge, particularly as data volume and complexity grow. This paper introduces the Sampled Grid Pairwise Likelihood (SG-PL) method, a novel approach that employs a grid-based sampling strategy to strategically select observation pairs. Simulation studies demonstrate SG-PL's principal advantage: a dramatic reduction in computational time -- often by orders of magnitude -- when compared to benchmark methods. This substantial acceleration is achieved with a manageable trade-off in statistical efficiency. An empirical application further validates SG-PL's practical utility. Consequently, SG-PL emerges as a highly scalable and effective tool for spatial analysis on very large datasets, offering a compelling balance where substantial gains in computational feasibility are realized for a limited cost in statistical precision, a trade-off that increasingly favors SG-PL with larger N.

stat.ME↗

Evaluating Large Language Model Capabilities in Assessing Spatial Econometrics Research

This paper investigates Large Language Models (LLMs) ability to assess the economic soundness and theoretical consistency of empirical findings in spatial econometrics. We created original and deliberately altered "counterfactual" summaries from 28 published papers (2005-2024), which were evaluated by a diverse set of LLMs. The LLMs provided qualitative assessments and structured binary classifications on variable choice, coefficient plausibility, and publication suitability. The results indicate that while LLMs can expertly assess the coherence of variable choices (with top models like GPT-4o achieving an overall F1 score of 0.87), their performance varies significantly when evaluating deeper aspects such as coefficient plausibility and overall publication suitability. The results further revealed that the choice of LLM, the specific characteristics of the paper and the interaction between these two factors significantly influence the accuracy of the assessment, particularly for nuanced judgments. These findings highlight LLMs' current strengths in assisting with initial, more surface-level checks and their limitations in performing comprehensive, deep economic reasoning, suggesting a potential assistive role in peer review that still necessitates robust human oversight.

cs.CY↗

On Spatio-Temporal Stochastic Frontier Models

In the literature on stochastic frontier models until the early 2000s, the joint consideration of spatial and temporal dimensions was often inadequately addressed, if not completely neglected. However, from an evolutionary economics perspective, the production process of the decision-making units constantly changes over both dimensions: it is not stable over time due to managerial enhancements and/or internal or external shocks, and is influenced by the nearest territorial neighbours. This paper proposes an extension of the Fusco and Vidoli [2013] SEM-like approach, which globally accounts for spatial and temporal effects in the term of inefficiency. In particular, coherently with the stochastic panel frontier literature, two different versions of the model are proposed: the time-invariant and the time-varying spatial stochastic frontier models. In order to evaluate the inferential properties of the proposed estimators, we first run Monte Carlo experiments and we then present the results of an application to a set of commonly referenced data, demonstrating robustness and stability of estimates across all scenarios.

stat.ME↗

Detecting Spatial Outliers: the Role of the Local Influence Function

In the analysis of large spatial datasets, identifying and treating spatial outliers is essential for accurately interpreting geographical phenomena. While spatial correlation measures, particularly Local Indicators of Spatial Association (LISA), are widely used to detect spatial patterns, the presence of abnormal observations frequently distorts the landscape and conceals critical spatial relationships. These outliers can significantly impact analysis due to the inherent spatial dependencies present in the data. Traditional influence function (IF) methodologies, commonly used in statistical analysis to measure the impact of individual observations, are not directly applicable in the spatial context because the influence of an observation is determined not only by its own value but also by its spatial location, its connections with neighboring regions, and the values of those neighboring observations. In this paper, we introduce a local version of the influence function (LIF) that accounts for these spatial dependencies. Through the analysis of both simulated and real-world datasets, we demonstrate how the LIF provides a more nuanced and accurate detection of spatial outliers compared to traditional LISA measures and local impact assessments, improving our understanding of spatial patterns.

stat.ME↗

On Robust Measures of Spatial Correlation

As a rule statistical measures are often vulnerable to the presence of outliers and spatial correlation coefficients, critical in the assessment of spatial data, remain susceptible to this inherent flaw. In contexts where data originates from a variety of domains (such as, e. g., socio-economic, environmental or epidemiological disciplines) it is quite common to encounter not just anomalous data points, but also non-normal distributions. These irregularities can significantly distort the broader analytical landscape often masking significant spatial attributes. This paper embarks on a mission to enhance the resilience of traditional spatial correlation metrics, specifically the Moran Coefficient (MC), Geary's Contiguity ratio (GC), and the Approximate Profile Likelihood Estimator (APLE) and to propose a series of alternative measures. Drawing inspiration from established analytical paradigms, our research harnesses the power of influence function studies to examine the robustness of the proposed novel measures in the presence of different outlier scenarios.

stat.ME↗

Feasible pairwise pseudo-likelihood inference on spatial regressions in irregular lattice grids: the KD-T PL algorithm

Spatial regression models are central to the field of spatial statistics. Nevertheless, their estimation in case of large and irregular gridded spatial datasets presents considerable computational challenges. To tackle these computational problems, Arbia \citep{arbia_2014_pairwise} introduced a pseudo-likelihood approach (called pairwise likelihood, say PL) which required the identification of pairs of observations that are internally correlated, but mutually conditionally uncorrelated. However, while the PL estimators enjoy optimal theoretical properties, their practical implementation when dealing with data observed on irregular grids suffers from dramatic computational issues (connected with the identification of the pairs of observations) that, in most empirical cases, negatively counter-balance its advantages. In this paper we introduce an algorithm specifically designed to streamline the computation of the PL in large and irregularly gridded spatial datasets, dramatically simplifying the estimation phase. In particular, we focus on the estimation of Spatial Error models (SEM). Our proposed approach, efficiently pairs spatial couples exploiting the KD tree data structure and exploits it to derive the closed-form expressions for fast parameter approximation. To showcase the efficiency of our method, we provide an illustrative example using simulated data, demonstrating the computational advantages if compared to a full likelihood inference are not at the expenses of accuracy.

stat.ME↗

Spatial sampling design to improve the efficiency of the estimation of the critical parameters of the SARS-CoV-2 epidemic

The pandemic linked to COVID-19 infection represents an unprecedented clinical and healthcare challenge for many medical researchers attempting to prevent its worldwide spread. This pandemic also represents a major challenge for statisticians involved in quantifying the phenomenon and in offering timely tools for the monitoring and surveillance of critical pandemic parameters. In a recent paper, Alleva et al. (2020) proposed a two-stage sample design to build a continuous-time surveillance system designed to correctly quantify the number of infected people through an indirect sampling mechanism that could be repeated in several waves over time to capture different target variables in the different stages of epidemic development. The proposed method exploits the indirect sampling (Lavalle, 2007; Kiesl, 2016) method employed in the estimation of rare and elusive populations (Borchers, 2009; Lavallée and Rivest, 2012) and a capture/recapture mechanism (Sudman, 1988; Thompson and Seber, 1996). In this paper, we extend the proposal of Alleva et al. (2020) to include a spatial sampling mechanism (Müller, 1998; Grafström et al., 2012, Jauslin and Tillè, 2020) in the process of data collection to achieve the same level of precision with fewer sample units, thereby facilitating the process of data collection in a situation where timeliness and costs are crucial elements. We present the basic idea of the new sample design, analytically prove the theoretical properties of the associated estimators and show the relative advantages through a systematic simulation study where all the typical elements of an epidemic are accounted for.

stat.ME↗

A sample approach to the estimation of the critical parameters of the SARS-CoV-2 epidemics: an operational design

Given the urgent informational needs connected with the diffusion of infection with regard to the COVID-19 pandemic, in this paper, we propose a sampling design for building a continuous-time surveillance system. Compared with other observational strategies, the proposed method has three important elements of strength and originality: (i) it aims to provide a snapshot of the phenomenon at a single moment in time, and it is designed to be a continuous survey that is repeated in several waves over time, taking different target variables during different stages of the development of the epidemic into account; (ii) the statistical optimality properties of the proposed estimators are formally derived and tested with a Monte Carlo experiment; and (iii) it is rapidly operational as this property is required by the emergency connected with the diffusion of the virus. The sampling design is thought to be designed with the diffusion of SAR-CoV-2 in Italy during the spring of 2020 in mind. However, it is very general, and we are confident that it can be easily extended to other geographical areas and to possible future epidemic outbreaks. Formal proofs and a Monte Carlo exercise highlight that the estimators are unbiased and have higher efficiency than the simple random sampling scheme.

stat.AP↗

On Spatial Lag Models estimated using crowdsourcing, web-scraping or other unconventionally collected data

The Big Data revolution is challenging the state-of-the-art statistical and econometric techniques not only for the computational burden connected with the high volume and speed which data are generated, but even more for the variety of sources through which data are collected (Arbia, 2021). This paper concentrates specifically on this last aspect. Common examples of non traditional Big Data sources are represented by crowdsourcing (data voluntarily collected by individuals) and web scraping (data extracted from websites and reshaped in a structured dataset). A common characteristic to these unconventional data collections is the lack of any precise statistical sample design, a situation described in statistics as 'convenience sampling'. As it is well known, in these conditions no probabilistic inference is possible. To overcome this problem, Arbia et al. (2018) proposed the use of a special form of post-stratification (termed 'post-sampling'), with which data are manipulated prior their use in an inferential context. In this paper we generalize this approach using the same idea to estimate a Spatial Lag Model (SLM). We start showing through a Monte Carlo study that using data collected without a proper design, parameters' estimates can be biased. Secondly, we propose a post sampling strategy to tackle this problem. We show that the proposed strategy indeed achieves a bias-reduction, but at the price of a concomitant increase in the variance of the estimators. We thus suggest an MSE-correction operational strategy. The paper also contains a formal derivation of the increase in variance implied by the post-sampling procedure and concludes with an empirical application of the method in the estimation of a hedonic price model in the city of Milan using web scraped data.

stat.ME↗

Observed and estimated prevalence of Covid-19 in Italy: Is it possible to estimate the total cases from medical swabs data?

During the current Covid-19 pandemic in Italy, official data are collected with medical swabs following a pure convenience criterion which, at least in an early phase, has privileged the exam of patients showing evident symptoms. However, there are evidences of a very high proportion of asymptomatic patients (e. g. Aguilar et al., 2020; Chugthai et al, 2020; Li, et al., 2020; Mizumoto et al., 2020a, 2020b and Yelin et al., 2020). In this situation, in order to estimate the real number of infected (and to estimate the lethality rate), it should be necessary to run a properly designed sample survey through which it would be possible to calculate the probability of inclusion and hence draw sound probabilistic inference. Some researchers proposed estimates of the total prevalence based on various approaches, including epidemiologic models, time series and the analysis of data collected in countries that faced the epidemic in earlier time (Brogi et al., 2020). In this paper, we propose to estimate the prevalence of Covid-19 in Italy by reweighting the available official data published by the Istituto Superiore di Sanità so as to obtain a more representative sample of the Italian population. Reweighting is a procedure commonly used to artificially modify the sample composition so as to obtain a distribution which is more similar to the population (Valliant et al., 2018). In this paper, we will use post-stratification of the official data, in order to derive the weights necessary for reweighting them using age and gender as post-stratification variables thus obtaining more reliable estimation of prevalence and lethality.

q-bio.QM↗

Post-sampling crowdsourced data to allow reliable statistical inference: the case of food price indices in Nigeria

Sound policy and decision making in developing countries is often limited by the lack of timely and reliable data. Crowdsourced data may provide a valuable alternative for data collection and analysis, e. g. in remote and insecure areas or of poor accessibility where traditional methods are difficult or costly. However, crowdsourced data are not directly usable to draw sound statistical inference. Indeed, its use involves statistical problems because data do not obey any formal sampling design and may also suffer from various non-sampling errors. To overcome this, we propose the use of a special form of post-stratification with which crowdsourced data are reweighted prior their use in an inferential context. An example in Nigeria illustrates the applicability of the method.

stat.ME↗

A Note on Early Epidemiological Analysis of Coronavirus Disease 2019 Outbreak using Crowdsourced Data

Crowdsourcing data can prove of paramount importance in monitoring and controlling the spread of infectious diseases. The recent paper by Sun, Chen and Viboud (2020) is important because it contributes to the understanding of the epidemiology and of the spreading of Covid-19 in a period when most of the epidemic characteristics are still unknown. However, the use of crowdsourcing data raises a number of problems from the statistical point of view which run the risk of invalidating the results and of biasing estimation and hypothesis testing. While the work by Sun, Chen and Viboud (2020) has to be commended, given the importance of the topic for worldwide health security, in this paper we deem important to remark the presence of the possible sources of statistical biases and to point out possible solutions to them

stat.AP↗

Reduced-bias estimation of spatial econometric models with incompletely geocoded data

The application of state-of-the-art spatial econometric models requires that the information about the spatial coordinates of statistical units is completely accurate, which is usually the case in the context of areal data. With micro-geographic point-level data, however, such information is inevitably affected by locational errors, that can be generated intentionally by the data producer for privacy protection or can be due to inaccuracy of the geocoding procedures. This unfortunate circumstance can potentially limit the use of the spatial econometric modelling framework for the analysis of micro data. Indeed, some recent contributions (see e.g. Arbia, Espa and Giuliani 2016) have shown that the presence of locational errors may have a non-negligible impact on the results. In particular, wrong spatial coordinates can lead to downward bias and increased variance in the estimation of model parameters. This contribution aims at developing a strategy to reduce the bias and produce more reliable inference for spatial econometrics models with location errors. The validity of the proposed approach is assessed by means of a Monte Carlo simulation study under different real-case scenarios. The study results show that the method is promising and can make the spatial econometric modelling of micro-geographic data possible.

stat.ME↗

Measurement error induced by locational uncertainty when estimating discrete choice models with a distance as a regressor

Spatial microeconometric studies typically suffer from various forms of inaccuracies that are not present when dealing with the classical regional spatial econometrics models. Among those, missing data, locational errors, sampling without a formal sample design, measurement errors and misalignment are the typical sources of inaccuracy that can affects the results in a spatial microeconometric analysis. In this paper, we have examined the effects of measurement error introduced in a logistic model by random geo-masking, when distances are used as predictors. Extending the classical results on the measurement error in a linear regression model, our MC experiment on hospital choices showed that the higher the distortion produced by the geo-masking, the higher is the downward bias in absolute value towards zero of the coefficient associated to the distance in a regression model.

stat.ME↗

A bivariate marginal likelihood specification of spatial econometric modeling of very large datasets

This paper proposes a bivariate marginal likelihood specification of spatial econometrics models that simplifies the derivation of the log-likelihood and leads to a closed form expression for the estimation of the parameters. With respect to the more traditional specifications of spatial autoregressive models, our method avoids the arbitrariness of the specification of a weight matrix, presents analytical and computational advantages and provides interesting interpretative insights. We establish small sample and asymptotic properties of the estimators and we derive the associated Fisher information matrix needed in confidence interval estimation and hypothesis testing.

stat.ME↗