SearcharxivSearch

arXiv subjects

Piercesare Secchi

Publications and source records attributed to Piercesare Secchi.

16 recordsLinked to original sources

Understanding Long-Term Dynamics of Individual Metro Usage: A Hidden Semi-Markov State Framework with Survival Analysis

Understanding how individual metro usage evolves over multi-year horizons is essential for transit planning and passenger retention. However, existing approaches typically characterize mobility patterns as static clusters or short-term variability, leaving the lifecycle dynamics of transit participation underexplored. This study proposes a state-based lifecycle modeling framework that integrates Hidden Semi-Markov Models (HSMM) with discrete-time survival analysis to characterize the evolution of individual metro mobility. The HSMM infers latent mobility states with explicit duration distributions and a transition matrix governing regime changes, while the survival component models exit and re-entry events via state-dependent hazard functions conditioned on mobility-state trajectories and behavioral history. Applied to four years of smart card data from the Shanghai metro system (2021-2024), the framework enables the identification of interpretable mobility states, the characterization of transition dynamics, and the quantification of state-dependent exit and re-entry processes. The analysis reveals five robust mobility states with a directional transition hierarchy centered on an occasional-usage gateway state, and fundamentally different temporal mechanisms governing disengagement and return: exit hazard is state-dependent but duration-independent, whereas re-entry hazard decays sharply with inactivity length. These findings provide a methodological foundation for lifecycle-oriented mobility analysis and practical guidance for transit operators to identify at-risk users and time retention interventions.

stat.AP

A Convolution Process for Sea Surface Temperature Hot-Spot Identification in the Mediterranean Sea

Sea surface temperature (SST) is a fundamental determinant of global climate dynamics and economic activity. Reliable projections of future SST patterns depend critically on a rigorous characterization of the underlying spatial random field. In this study, we introduce a novel convolution-based covariance framework tailored to geostatistical domains constrained by physical barriers and influenced by vector-driven flows. By discretizing the continuous marine domain into a directed linear network that preserves the orientation of ocean currents, we construct a moving-average stochastic process whose dynamic is encoded via a Markovian transition-probability matrix on the network's vertices. The induced covariance structure emerges as a weighted combination of a spatial kernel and flow-dependent weights, giving rise to a complex estimation problem. To stabilize inference, we propose a penalized estimator that regularizes covariance parameters while enforcing consistency with known hydrodynamic properties. We then embed this covariance model into a Monte Carlo simulation framework to refine RCP-based SST projections and to identify thermal 'hot spots' of heightened ecological risk. Our approach delivers a statistically principled framework that prevents physical inconsistencies -- such as correlations across land barriers -- providing a robust basis for quantifying uncertainty in future SST forecasts and for guiding targeted environmental assessments.

stat.ME

Fixed Rank co-Kriging: a model for multivariate spatial prediction

This work develops a multivariate extension of the Fixed Rank Kriging (FRK) framework for spatial prediction in settings where multiple spatial processes may provide complementary information. The goal is to preserve the computational efficiency, the ability to operate without assuming stationarity over the domain, and the spatial support flexibility of FRK, while incorporating cross-process dependence. To this end, we employ a multiresolution coregionalization structure for the latent spatial effects, in which spatial basis functions are combined with Gaussian Markov Random Field coefficients. An estimation procedure based on the expectation-maximization algorithm is developed, designed to exploit the multiresolution latent structure. Through simulation studies, we examine when the proposed joint modeling is beneficial. We consider cases in which one process is observed more sparsely or is entirely unobserved in a subregion and find that the multivariate formulation is able to borrow information from the more densely observed process, producing coherent and accurate predictions even where direct observations are limited or absent. Finally, the model is applied to the analysis of PM10 concentrations in Northern Italy, illustrating its applicability in a real environmental context.

stat.ME

Unraveling time-varying causal effects of multiple exposures: integrating Functional Data Analysis with Multivariable Mendelian Randomization

Mendelian Randomization is a widely used instrumental variable method for assessing causal effects of lifelong exposures on health outcomes. Many exposures, however, have causal effects that vary across the life course and often influence outcomes jointly with other exposures or indirectly through mediating pathways. Existing approaches to multivariable Mendelian Randomization assume constant effects over time and therefore fail to capture these dynamic relationships. We introduce Multivariable Functional Mendelian Randomization (MV-FMR), a new framework that extends functional Mendelian Randomization to simultaneously model multiple time-varying exposures. The method combines functional principal component analysis with a data-driven cross-validation strategy for basis selection and accounts for overlapping instruments and mediation effects. Through extensive simulations, we assessed MV-FMR's ability to recover time-varying causal effects under a range of data-generating scenarios and compared the performance of joint versus separate exposure effect estimation strategies. Across scenarios involving nonlinear effects, horizontal pleiotropy, mediation, and sparse data, MV-FMR consistently recovered the true causal functions and outperformed univariable approaches. To demonstrate its practical value, we applied MV-FMR to UK Biobank data to investigate the time-varying causal effects of systolic blood pressure and body mass index on coronary artery disease. MV-FMR provides a flexible and interpretable framework for disentangling complex time-dependent causal processes and offers new opportunities for identifying life-course critical periods and actionable drivers relevant to disease prevention.

stat.AP

Day-Ahead Electricity Price Forecasting Using Merit-Order Curves Time Series

We introduce a general, simple, and computationally efficient functional data analysis framework for forecasting day-ahead supply and demand merit-order curves, and the resulting electricity price. We conduct a rigorous empirical comparison on data from the Italian (GME), German (EPEX-DE-LU), and French (EPEX-FR) day-ahead markets over the 2023-2024 period, analyzing curve forecasting performance, price forecasting performance, and the relationship between the two. We find that strong curve forecasting performance does not necessarily translate into strong price forecasting performance, with important implications for curve model evaluation and selection when price forecasting is among the objectives. We also show that this functional data representation approach consistently outperforms the original discretization-based approach of Ziel and Steinert (2016) on price forecasting across all three markets. Finally, the proposed curve-based approach is competitive with state-of-the-art price-based models for two out of three markets (GME and EPEX-FR), and substantially improves accuracy during midday hours (when prices frequently drop due to high renewable generation) with MAE reductions of up to 27% in those windows. For EPEX-DE-LU, however, price-based models retain a clear and significant advantage.

stat.AP

A Blind Source Separation Framework to Monitor Sectoral Power Demand from Grid-Scale Load Measurements

As demand-side flexibility becomes increasingly necessary to integrate variable renewable energy, understanding electricity demand composition across different grid levels is essential. However, at regional and national scales, visibility into the relative contributions of different consumer categories remains limited due to the complexity and cost of collecting end-use consumption data. To address this challenge, we propose a blind source separation framework to disaggregate open-access high-voltage grid load measurements into sectoral contributions. The approach relies on a constrained variant of non-negative matrix factorization, termed linearly-constrained non-negative matrix factorization (LCNMF), which allows prior information to be incorporated as linear constraints on the factor matrices, thereby providing weak supervision of the separation process. The framework is evaluated using Italian national load data from 2021 to 2023. Results demonstrate the identifiability of residential, services, and industrial load components and provide monthly sectoral consumption estimates consistent with reported statistics. The proposed method is generalizable and applicable to load disaggregation problems across multiple grid scales where disaggregated measurements are unavailable.

stat.AP

Three Distributional Approaches for PM10 Assessment in Northern Italy

We propose three spatial methods for estimating the full probability distribution of PM10 concentrations, with the ultimate goal of assessing air quality in Northern Italy. Moving beyond spatial averages and simple indicators, we adopt a distributional perspective to capture the complex variability of pollutant concentrations across space. The first proposed approach predicts class-based compositions via Fixed Rank Kriging; the second estimates multiple, non-crossing quantiles through a spatial regression with differential regularization; the third directly reconstructs full probability densities leveraging on both Fixed Rank Kriging and multiple quantiles spatial regression within a Simplicial Principal Component Analysis framework. These approaches are applied to daily PM10 measurements, collected from 2018 to 2022 in Northern Italy, to estimate spatially continuous distributions and to identify regions at risk of regulatory exceedance. The three approaches exhibit localized differences, revealing how modeling assumptions may influence the prediction of fine-scale pollutant concentration patterns. Nevertheless, they consistently agree on the broader spatial patterns of pollution. This general agreement supports the robustness of a distributional approach, which offers a comprehensive and policy-relevant framework for assessing air quality and regulatory exceedance risks.

stat.AP

Spatio-Temporal Analysis of Public Transportation Undercrowding: Leveraging APC Data for a Comprehensive Evaluation of Usage Rates

The analysis of the transportation usage rate provides opportunities for evaluating the efficacy of the transportation service offered by proposing an indicator that integrates actual demand and capacity. This study aims to develop a methodology for analyzing the occupancy rate from large-scale datasets to identify gaps between supply and demand in public transportation. Leveraging the spatio-temporal granularity of data from Automatic People Counting (APC) and relying on the Generalized Linear Mixed Effects Model and the Generalized Mixed-Effect Random Forest, in this study we propose a methodology for analyzing factors determining undercrowding. The results of the model are examined at both the segment and ride levels. Initially, the analysis focuses on identifying segments more likely associated with undercrowding, understanding factors influencing the probability of undercrowding, and exploring their relationships. Subsequently, the analysis extends to the temporal distribution of undercrowding, encompassing its impact on the entire journey. The proposed methodology is applied to analyze APC data, provided by the company responsible for public transport management in Milan, on a radial route of the surface transportation network.

stat.AP

Persistence diagrams for exploring the shape variability of abdominal aortic aneurysms

Abdominal Aortic Aneurysm consists of a permanent dilation in the abodminal portion of the aorta and, along with its associated pathologies like calcifications and intraluminal thrombi, is one of the most important pathologies of the circulatory system. The shape of the aorta is among the primary drivers for these health issues, with particular reference to all the characteristics which affects the hemodynamics. Starting from the computed tomography angiography of a patient, we propose to summarize such information using tools derived from Topological Data Analysis, obtaining persistence diagrams which describe the irregularities of the lumen of the aorta. We showcase the effectiveness of such shape-related descriptors with a series of supervised and unsupervised case studies.

physics.med-ph

Heterogeneous drivers of overnight and same-day visits

This paper aims to explore the factors stimulating different tourism behaviours, with specific reference to same-day visits and overnight stays. To this aim, we employ mobile network data referred to the area of Lombardy. The paper highlights that larger availability of tourism accommodations, cultural and natural endowments are relevant factors explaining overnight stays. Conversely, temporary entertainment and transportation facilities increase municipalities attractiveness for same-day visits. The results also highlight a trade-off in the capability of municipalities of being attractive in connection to both the tourism behaviours, with higher overnight stays in areas with more limited same-day visits. Mobile data offer a spatial and temporal granularity allowing to detect relevant patterns and support the design of tourism precision policies.

econ.GN

Estimation of Dynamic Origin-Destination Matrices in a Railway Transportation Network integrating Ticket Sales and Passenger Count Data

Accurately estimating Origin-Destination (OD) matrices is a topic of increasing interest for efficient transportation network management and sustainable urban planning. Traditionally, travel surveys have supported this process; however, their availability and comprehensiveness can be limited. Moreover, the recent COVID-19 pandemic has triggered unprecedented shifts in mobility patterns, underscoring the urgency of accurate and dynamic mobility data supporting policies and decisions with data-driven evidence. In this study, we tackle these challenges by introducing an innovative pipeline for estimating dynamic OD matrices. The real motivating problem behind this is based on the Trenord railway transportation network in Lombardy, Italy. We apply a novel approach that integrates ticket and subscription sales data with passenger counts obtained from Automated Passenger Counting (APC) systems, making use of the Iterative Proportional Fitting (IPF) algorithm. Our work effectively addresses the complexities posed by incomplete and diverse data sources, showcasing the adaptability of our pipeline across various transportation contexts. Ultimately, this research bridges the gap between available data sources and the escalating need for precise OD matrices. The proposed pipeline fosters a comprehensive grasp of transportation network dynamics, providing a valuable tool for transportation operators, policymakers, and researchers. Indeed, to highlight the potentiality of dynamic OD matrices, we showcase some methods to perform anomaly detection of mobility trends in the network through such matrices and interpret them in light of events that happened in the last months of 2022.

stat.AP

Functional Data Representation with Merge Trees

In this paper we face the problem of representation of functional data with the tools of algebraic topology. We represent functions by means of merge trees, which, like the more commonly used persistence diagrams, are invariant under homeomorphic reparametrizations of the functions they represent, thus allowing for a statistical analysis which is indifferent to functional misalignment. We consider a recently defined metric for merge trees and we prove some theoretical results related to its specific implementation when merge trees represent functions, establishing also a class of consistent estimators with convergence rates. To showcase the good properties of our topological approach to functional data analysis, we test it on the Aneurisk65 dataset replicating, from our different perspective, the supervised classification analysis which contributed to make this dataset a benchmark for methods dealing with misaligned functional data. In the Appendix we provide an extensive comparison between merge trees and persistence diagrams, highlighting similarities and differences, which can guide the analyst in choosing between the two representations.

stat.ME

Social and material vulnerability in the face of seismic hazard: an analysis of the Italian case

The assessment of the vulnerability of a community endangered by seismic hazard is of paramount importance for planning a precision policy aimed at the prevention and reduction of its seismic risk. We aim at measuring the vulnerability of the Italian municipalities exposed to seismic hazard, by analyzing the open data offered by the Mappa dei Rischi dei Comuni Italiani provided by ISTAT, the Italian National Institute of Statistics. Encompassing the Index of Social and Material Vulnerability already computed by ISTAT, we also consider as referents of the latent social and material vulnerability of a community, its demographic dynamics and the age of the building stock where the community resides. Fusing the analyses of different indicators, within the context of seismic risk we offer a tentative ranking of the Italian municipalities in terms of their social and material vulnerability, together with differential profiles of their dominant fragilities which constitute the basis for planning precision policies aimed at seismic risk prevention and reduction.

stat.AP

Kriging Riemannian Data via Random Domain Decompositions

Data taking value on a Riemannian manifold and observed over a complex spatial domain are becoming more frequent in applications, e.g. in environmental sciences and in geoscience. The analysis of these data needs to rely on local models to account for the non stationarity of the generating random process, the non linearity of the manifold and the complex topology of the domain. In this paper, we propose to use a random domain decomposition approach to estimate an ensemble of local models and then to aggregate the predictions of the local models through Fréchet averaging. The algorithm is introduced in complete generality and is valid for data belonging to any smooth Riemannian manifold but it is then described in details for the case of the manifold of positive definite matrices, the hypersphere and the Cholesky manifold. The predictive performance of the method are explored via simulation studies for covariance matrices and correlation matrices, where the Cholesky manifold geometry is used. Finally, the method is illustrated on an environmental dataset observed over the Chesapeake Bay (USA).

stat.ME

Permutation tests for the equality of covariance operators of functional data with applications to evolutionary biology

In this paper, we generalize the metric-based permutation test for the equality of covariance operators proposed by Pigoli et al. (2014) to the case of multiple samples of functional data. To this end, the non-parametric combination methodology of Pesarin and Salmaso (2010) is used to combine all the pairwise comparisons between samples into a global test. Different combining functions and permutation strategies are reviewed and analyzed in detail. The resulting test allows to make inference on the equality of the covariance operators of multiple groups and, if there is evidence to reject the null hypothesis, to identify the pairs of groups having different covariances. It is shown that, for some combining functions, step-down adjusting procedures are available to control for the multiple testing problem in this setting. The empirical power of this new test is then explored via simulations and compared with those of existing alternative approaches in different scenarios. Finally, the proposed methodology is applied to data from wheel running activity experiments, that used selective breeding to study the evolution of locomotor behavior in mice.

stat.ME

A functional equation whose unknown is P([0,1]) valued

We study a functional equation whose unknown maps a Euclidean space into the space of probability distributions on [0,1]. We prove existence and uniqueness of its solution under suitable regularity and boundary conditions, we show that it depends continuously on the boundary datum, and we characterize solutions that are diffuse on [0,1]. A canonical solution is obtained by means of a Randomly Reinforced Urn with different reinforcement distributions having equal means. The general solution to the functional equation defines a new parametric collection of distributions on [0,1] generalizing the Beta family.

math.PR