SearcharxivSearch

arXiv subjects

Simone Vantini

Publications and source records attributed to Simone Vantini.

At least 19 recordsLinked to original sources

Generalized propensity score weighting for functional causal inference framework

Estimating causal effects in observational studies requires adjustment for confounding, a task that becomes challenging when the exposure is a function observed over a continuous domain rather than a scalar variable. We develop a functional propensity score weighting framework that achieves covariate balance by removing dependence between time-varying treatments and observed confounders, thereby enabling estimation of marginal causal effects in settings with functional treatments, covariates, and outcomes. We propose a dual formulation of the weight estimation problem that yields a smooth unconstrained optimization and improves computational scalability. The proposed framework extends naturally to settings with time-varying covariates and to longitudinal outcomes via a function-on-function marginal structural model, allowing estimation of causal effect surfaces. The proposed method improves covariate balance, estimation accuracy, and computational efficiency compared to the existing approach and retains these properties when extended to functional covariates and outcomes. We apply the method to data from the UK Biobank to estimate the causal effect of body mass index trajectories on the risk of Type 2 Diabetes and on subsequent glycated hemoglobin trajectories, a functional measure of metabolic status.

stat.ME

Regularized covariance estimation from partially observed interferometric data

The Small BAseline Subset technique provides remote measurements of ground displacement with high spatial resolution, making it a key tool for monitoring geophysical processes in hazard-prone areas. An effective analysis of this type of data requires reliable estimation of their second-order structure, which is difficult to achieve because the measurements are systematically missing over relatively large portions of the investigated areas. We tackle the problem from a functional data analysis perspective and treat the observations as partially observed functional data with two-dimensional domain. To properly characterize the data, we introduce the fragmented regime of partial observation, where parts of the curves are systematically missing across replicates. For this regime, we propose a novel method for covariance estimation, formulating the task as a matrix completion problem with Laplacian regularization. The estimator is nonparametric and free from stationarity or isotropy assumptions. Extensive simulations show that our method achieves consistently low estimation error across a range of covariance structures. Application to ground displacement data relative to the Phlegraean Fields demonstrates its ability to recover meaningful spatial dependence patterns, highlighting its potential for environmental risk assessment and monitoring.

stat.ME

Reconstructing Movement from Sparse Samples: Enhanced Spatio-Temporal Matching Strategies for Low-Frequency Data

This paper explores potential improvements to the Spatial-Temporal Matching algorithm for aligning the GPS trajectories to road networks. While this algorithm is effective, it presents some limitations in computational efficiency and the accuracy of the results, especially in dense environments with relatively high sampling intervals. To address this, the paper proposes four modifications to the original algorithm: a dynamic buffer, an adaptive observation probability, a redesigned temporal scoring function, and a behavioral analysis to account for the historical mobility patterns. The enhancements are assessed using real-world data from the urban area of Milan, and through newly defined evaluation metrics to be applied in the absence of ground truth. The results of the experiment show significant improvements in performance efficiency and path quality across various metrics.

cs.LG

Multi-state Modeling of Delay Evolution in Suburban Rail Transports

Train delays are a persistent issue in railway systems, particularly in suburban networks where operational complexity is heightened by frequent services and high passenger volumes. Traditional delay models often overlook the temporal and structural dynamics of real delay propagation. This work applies continuous-time multi-state models to analyze the temporal evolution of delay on the S5 suburban line in Lombardy, Italy. Using detailed operational, meteorological, and contextual data, the study models delay transitions while accounting for observable heterogeneity. The findings reveal how delay dynamics vary by travel direction, time slot, and route segment. Covariates such as station saturation and passenger load are shown to significantly affect the risk of delay escalation or recovery. The study offers both methodological advancements and practical results for improving the reliability of rail services.

stat.AP

Hyper-spectral Unmixing algorithms for remote compositional surface mapping: a review of the state of the art

This work concerns a detailed review of data analysis methods used for remotely sensed images of large areas of the Earth and of other solid astronomical objects. In detail, it focuses on the problem of inferring the materials that cover the surfaces captured by hyper-spectral images and estimating their abundances and spatial distributions within the region. The most successful and relevant hyper-spectral unmixing methods are reported as well as compared, as an addition to analysing the most recent methodologies. The most important public data-sets in this setting, which are vastly used in the testing and validation of the former, are also systematically explored. Finally, open problems are spotlighted and concrete recommendations for future research are provided.

astro-ph.IM

Noise-Adaptive Conformal Classification with Marginal Coverage

Conformal inference provides a rigorous statistical framework for uncertainty quantification in machine learning, enabling well-calibrated prediction sets with precise coverage guarantees for any classification model. However, its reliance on the idealized assumption of perfect data exchangeability limits its effectiveness in the presence of real-world complications, such as low-quality labels -- a widespread issue in modern large-scale data sets. This work tackles this open problem by introducing an adaptive conformal inference method capable of efficiently handling deviations from exchangeability caused by random label noise, leading to informative prediction sets with tight marginal coverage guarantees even in those challenging scenarios. We validate our method through extensive numerical experiments demonstrating its effectiveness on synthetic and real data sets, including CIFAR-10H and BigEarthNet.

stat.ME

Calibrated quantile prediction for Growth-at-Risk

Accurate computation of robust estimates for extremal quantiles of empirical distributions is an essential task for a wide range of applicative fields, including economic policymaking and the financial industry. Such estimates are particularly critical in calculating risk measures, such as Growth-at-Risk (GaR). % and Value-at-Risk (VaR). This work proposes a conformal framework to estimate calibrated quantiles, and presents an extensive simulation study and a real-world analysis of GaR to examine its benefits with respect to the state of the art. Our findings show that CP methods consistently improve the calibration and robustness of quantile estimates at all levels. The calibration gains are appreciated especially at extremal quantiles, which are critical for risk assessment and where traditional methods tend to fall short. In addition, we introduce a novel property that guarantees coverage under the exchangeability assumption, providing a valuable tool for managing risks by quantifying and controlling the likelihood of future extreme observations.

stat.ME

Generating Synthetic Functional Data for Privacy-Preserving GPS Trajectories

This research presents FDASynthesis, a novel algorithm designed to generate synthetic GPS trajectory data while preserving privacy. After pre-processing the input GPS data, human mobility traces are modeled as multidimensional curves using Functional Data Analysis (FDA). Then, the synthesis process identifies the K-nearest trajectories and averages their Square-Root Velocity Functions (SRVFs) to generate synthetic data. This results in synthetic trajectories that maintain the utility of the original data while ensuring privacy. Although applied for human mobility research, FDASynthesis is highly adaptable to different types of functional data, offering a scalable solution in various application domains.

stat.AP

Spatio-Temporal Analysis of Public Transportation Undercrowding: Leveraging APC Data for a Comprehensive Evaluation of Usage Rates

The analysis of the transportation usage rate provides opportunities for evaluating the efficacy of the transportation service offered by proposing an indicator that integrates actual demand and capacity. This study aims to develop a methodology for analyzing the occupancy rate from large-scale datasets to identify gaps between supply and demand in public transportation. Leveraging the spatio-temporal granularity of data from Automatic People Counting (APC) and relying on the Generalized Linear Mixed Effects Model and the Generalized Mixed-Effect Random Forest, in this study we propose a methodology for analyzing factors determining undercrowding. The results of the model are examined at both the segment and ride levels. Initially, the analysis focuses on identifying segments more likely associated with undercrowding, understanding factors influencing the probability of undercrowding, and exploring their relationships. Subsequently, the analysis extends to the temporal distribution of undercrowding, encompassing its impact on the entire journey. The proposed methodology is applied to analyze APC data, provided by the company responsible for public transport management in Milan, on a radial route of the surface transportation network.

stat.AP

An interpretable and transferable model for shallow landslides detachment combining spatial Poisson point processes and generalized additive models

Less than 10 meters deep, shallow landslides are rapidly moving and strongly dangerous slides. In the present work, the probabilistic distribution of the landslide detachment points within a valley is modelled as a spatial Poisson point process, whose intensity depends on geophysical predictors according to a generalized additive model. Modelling the intensity with a generalized additive model jointly allows to obtain good predictive performance and to preserve the interpretability of the effects of the geophysical predictors on the intensity of the process. We propose a novel workflow, based on Random Forests, to select the geophysical predictors entering the model for the intensity. In this context, the statistically significant effects are interpreted as activating or stabilizing factors for landslide detachment. In order to guarantee the transferability of the resulting model, training, validation, and test of the algorithm are performed on mutually disjoint valleys in the Alps of Lombardy (Italy). Finally, the uncertainty around the estimated intensity of the process is quantified via semiparametric bootstrap.

stat.AP

Monitoring road infrastructures from satellite images in Greater Maputo

The information about pavement surface type is rarely available in road network databases of developing countries although it represents a cornerstone of the design of efficient mobility systems. This research develops an automatic classification pipeline for road pavement which makes use of satellite images to recognize road segments as paved or unpaved. The proposed methodology is based on an object-oriented approach, so that each road is classified by looking at the distribution of its pixels in the RGB space. The proposed approach is proven to be accurate, inexpensive, and readily replicable in other cities.

stat.AP

Urban mobility and learning: analyzing the influence of commuting time on students' GPA at Politecnico di Milano

Despite its crucial role in students' daily lives, commuting time remains an underexplored dimension in higher education research. To address this gap, this study focuses on challenges that students face in urban environments and investigates the impact of commuting time on the academic performance of first-year bachelor students of Politecnico di Milano, Italy. This research employs an innovative two-step methodology. In the initial phase, machine learning algorithms trained on privacy-preserving GPS data from anonymous users are used to construct accessibility maps to the university and to obtain an estimate of students' commuting times. In the subsequent phase, authors utilize polynomial linear mixed-effects models and investigate the factors influencing students' academic performance, with a particular emphasis on commuting time. Notably, this investigation incorporates a causal framework, which enables the establishment of causal relationships between commuting time and academic outcomes. The findings underscore the significant impact of travel time on students' performance and may support policies and implications aiming at improving students' educational experience in metropolitan areas. The study's innovation lies both in its exploration of a relatively uncharted factor and the novel methodologies applied in both phases.

stat.AP

Conformal Prediction Sets for Populations of Graphs

The analysis of data such as graphs has been gaining increasing attention in the past years. This is justified by the numerous applications in which they appear. Several methods are present to predict graphs, but much fewer to quantify the uncertainty of the prediction. The present work proposes an uncertainty quantification methodology for graphs, based on conformal prediction. The method works both for graphs with the same set of nodes (labelled graphs) and graphs with no clear correspondence between the set of nodes across the observed graphs (unlabelled graphs). The unlabelled case is dealt with the creation of prediction sets embedded in a quotient space. The proposed method does not rely on distributional assumptions, it achieves finite-sample validity, and it identifies interpretable prediction sets. To explore the features of this novel forecasting technique, we perform two simulation studies to show the methodology in both the labelled and the unlabelled case. We showcase the applicability of the method in analysing the performance of different teams during the FIFA 2018 football world championship via their player passing networks.

stat.ME

Local inference for functional data on manifold domains using permutation tests

Pini and Vantini (2017) introduced the interval-wise testing procedure which performs local inference for functional data defined on an interval domain, where the output is an adjusted p-value function that controls for type I errors. We extend this idea to a general setting where domain is a Riemannian manifolds. This requires new methodology such as how to define adjustment sets on product manifolds and how to approximate the test statistic when the domain has non-zero curvature. We propose to use permutation tests for inference and apply the procedure in three settings: a simulation on a "chameleon-shaped" manifold and two applications related to climate change where the manifolds are a complex subset of $S^2$ and $S^2 \times S^1$, respectively. We note the tradeoff between type I and type II errors: increasing the adjustment set reduces the type I error but also results in smaller areas of significance. However, some areas still remain significant even at maximal adjustment.

stat.ME

funLOCI: a local clustering algorithm for functional data

Nowadays, more and more problems are dealing with data with one infinite continuous dimension: functional data. In this paper, we introduce the funLOCI algorithm which allows to identify functional local clusters or functional loci, i.e., subsets/groups of functions exhibiting similar behaviour across the same continuous subset of the domain. The definition of functional local clusters leverages ideas from multivariate and functional clustering and biclustering and it is based on an additive model which takes into account the shape of the curves. funLOCI is a three-step algorithm based on divisive hierarchical clustering. The use of dendrograms allows to visualize and to guide the searching procedure and the cutting thresholds selection. To deal with the large quantity of local clusters, an extra step is implemented to reduce the number of results to the minimum.

stat.ME

An Evaluation of Researchers' Migration Patterns in Europe using Digital Trace Data

The comprehension of the mechanisms behind the mobility of skilled workers is of paramount importance for policy making. The lacking nature of official measurements motivates the use of digital trace data extracted from ORCID public records. We use such data to investigate European regions, studied at NUTS2 level, over the time horizon of 2009 to 2020. We present a novel perspective where regions roles are dictated by the overall activity of the research community, contradicting the common brain drain interpretation of the phenomenon. We find that a high mobility is usually correlated with strong university prestige, high magnitude of investments and an overall good schooling level in a region.

stat.AP

Conformal Prediction Bands for Two-Dimensional Functional Time Series

Time evolving surfaces can be modeled as two-dimensional Functional time series, exploiting the tools of Functional data analysis. Leveraging this approach, a forecasting framework for such complex data is developed. The main focus revolves around Conformal Prediction, a versatile nonparametric paradigm used to quantify uncertainty in prediction problems. Building upon recent variations of Conformal Prediction for Functional time series, a probabilistic forecasting scheme for two-dimensional functional time series is presented, while providing an extension of Functional Autoregressive Processes of order one to this setting. Estimation techniques for the latter process are introduced and their performance are compared in terms of the resulting prediction regions. Finally, the proposed forecasting procedure and the uncertainty quantification technique are applied to a real dataset, collecting daily observations of Sea Level Anomalies of the Black Sea

stat.ME

funcharts: Control charts for multivariate functional data in R

Modern statistical process monitoring (SPM) applications focus on profile monitoring, i.e., the monitoring of process quality characteristics that can be modeled as profiles, also known as functional data. Despite the large interest in the profile monitoring literature, there is still a lack of software to facilitate its practical application. This article introduces the funcharts R package that implements recent developments on the SPM of multivariate functional quality characteristics, possibly adjusted by the influence of additional variables, referred to as covariates. The package also implements the real-time version of all control charting procedures to monitor profiles partially observed up to an intermediate domain point. The package is illustrated both through its built-in data generator and a real-case study on the SPM of Ro-Pax ship CO2 emissions during navigation, which is based on the ShipNavigation data provided in the Supplementary Material.

stat.CO