SearcharxivSearch

arXiv subjects

Alessandra Menafoglio

Publications and source records attributed to Alessandra Menafoglio.

16 recordsLinked to original sources

Random mixtures in Bayes Hilbert spaces

We present a framework for the analysis and unmixing of random density mixtures in the Bayes Hilbert space. General identifiability results for mixtures in Hilbert spaces are established and applied to the Bayes Hilbert space setting. Building on these results, we propose a penalised maximum likelihood approach for the unmixing of Bayes Hilbert mixtures aimed at recovering the statistically space-efficient representation, together with a computationally efficient coordinate-wise maximisation algorithm for its implementation. The methodology is illustrated through a hyperspectral data application, where observations can be naturally embedded in the Bayes Hilbert space and analyzed in terms of distributional shape rather than amplitude. A complementary simulation study demonstrates the interpretability and practical performance of the proposed approach.

stat.ME

Regularized covariance estimation from partially observed interferometric data

The Small BAseline Subset technique provides remote measurements of ground displacement with high spatial resolution, making it a key tool for monitoring geophysical processes in hazard-prone areas. An effective analysis of this type of data requires reliable estimation of their second-order structure, which is difficult to achieve because the measurements are systematically missing over relatively large portions of the investigated areas. We tackle the problem from a functional data analysis perspective and treat the observations as partially observed functional data with two-dimensional domain. To properly characterize the data, we introduce the fragmented regime of partial observation, where parts of the curves are systematically missing across replicates. For this regime, we propose a novel method for covariance estimation, formulating the task as a matrix completion problem with Laplacian regularization. The estimator is nonparametric and free from stationarity or isotropy assumptions. Extensive simulations show that our method achieves consistently low estimation error across a range of covariance structures. Application to ground displacement data relative to the Phlegraean Fields demonstrates its ability to recover meaningful spatial dependence patterns, highlighting its potential for environmental risk assessment and monitoring.

stat.ME

K-Models: a Flexible and Interpretable Method for Ordinal Clustering with Application to Antigen-Antibody Interaction Profiles

Existing clustering methods for functional data often prioritize partitioning accuracy over interpretability, making it challenging to extract meaningful insights when the data-generating process follows a specific underlying structure and an ordinal relationship among clusters is suspected. This work introduces K-Models, a novel framework that integrates ordinal constraints and estimates key underlying elements of the random process generating the observed functional profiles, improving both interpretability and structure identification. The proposed method is evaluated through simulations and real-world applications. In particular, it is tested on Region of Interest (ROI) curves, which represent reaction profiles from a reflectometric sensor monitoring biomolecular interactions, such as antigen-antibody binding. These curves represent changes in reflected light intensity over time at multiple measurement spots with immobilized antigens during analyte exposure, capturing the binding dynamics of the system. The goal is to identify intrinsic signal patterns solely from the observed dynamics, making this dataset an ideal benchmark for assessing the added interpretability of the proposed approach. By incorporating structural assumptions into the clustering process, K-Models enhances interpretability while maintaining performance comparable to state-of-the-art techniques, providing a valuable tool for analyzing functional data with an underlying ordinal structure.

stat.ML

A Convolution Process for Sea Surface Temperature Hot-Spot Identification in the Mediterranean Sea

Sea surface temperature (SST) is a fundamental determinant of global climate dynamics and economic activity. Reliable projections of future SST patterns depend critically on a rigorous characterization of the underlying spatial random field. In this study, we introduce a novel convolution-based covariance framework tailored to geostatistical domains constrained by physical barriers and influenced by vector-driven flows. By discretizing the continuous marine domain into a directed linear network that preserves the orientation of ocean currents, we construct a moving-average stochastic process whose dynamic is encoded via a Markovian transition-probability matrix on the network's vertices. The induced covariance structure emerges as a weighted combination of a spatial kernel and flow-dependent weights, giving rise to a complex estimation problem. To stabilize inference, we propose a penalized estimator that regularizes covariance parameters while enforcing consistency with known hydrodynamic properties. We then embed this covariance model into a Monte Carlo simulation framework to refine RCP-based SST projections and to identify thermal 'hot spots' of heightened ecological risk. Our approach delivers a statistically principled framework that prevents physical inconsistencies -- such as correlations across land barriers -- providing a robust basis for quantifying uncertainty in future SST forecasts and for guiding targeted environmental assessments.

stat.ME

Hyper-spectral Unmixing algorithms for remote compositional surface mapping: a review of the state of the art

This work concerns a detailed review of data analysis methods used for remotely sensed images of large areas of the Earth and of other solid astronomical objects. In detail, it focuses on the problem of inferring the materials that cover the surfaces captured by hyper-spectral images and estimating their abundances and spatial distributions within the region. The most successful and relevant hyper-spectral unmixing methods are reported as well as compared, as an addition to analysing the most recent methodologies. The most important public data-sets in this setting, which are vastly used in the testing and validation of the former, are also systematically explored. Finally, open problems are spotlighted and concrete recommendations for future research are provided.

astro-ph.IM

Functional-Ordinal Canonical Correlation Analysis With Application to Data from Optical Sensors

We address the problem of predicting a target ordinal variable based on observable features consisting of functional profiles. This problem is crucial, especially in decision-making driven by sensor systems, when the goal is to assess an ordinal variable such as the degree of deterioration, quality level, or risk stage of a process, starting from functional data observed via sensors. We purposely introduce a novel approach called functional-ordinal Canonical Correlation Analysis (foCCA), which is based on a functional data analysis approach. FoCCA allows for dimensionality reduction of observable features while maximizing their ability to differentiate between consecutive levels of an ordinal target variable. Unlike existing methods for supervised learning from functional data, foCCA fully incorporates the ordinal nature of the target variable. This enables the model to capture and represent the relative dissimilarities between consecutive levels of the ordinal target, while also explaining these differences through the functional features. Extensive simulations demonstrate that foCCA outperforms current state-of-the-art methods in terms of prediction accuracy in the reduced feature space. A case study involving the prediction of antigen concentration levels from optical biosensor signals further confirms the superior performance of foCCA, showcasing both improved predictive power and enhanced interpretability compared to competing approaches.

stat.ME

Noise-Adaptive Conformal Classification with Marginal Coverage

Conformal inference provides a rigorous statistical framework for uncertainty quantification in machine learning, enabling well-calibrated prediction sets with precise coverage guarantees for any classification model. However, its reliance on the idealized assumption of perfect data exchangeability limits its effectiveness in the presence of real-world complications, such as low-quality labels -- a widespread issue in modern large-scale data sets. This work tackles this open problem by introducing an adaptive conformal inference method capable of efficiently handling deviations from exchangeability caused by random label noise, leading to informative prediction sets with tight marginal coverage guarantees even in those challenging scenarios. We validate our method through extensive numerical experiments demonstrating its effectiveness on synthetic and real data sets, including CIFAR-10H and BigEarthNet.

stat.ME

Robust functional PCA for relative data

This paper introduces a robust approach to functional principal component analysis (FPCA) for relative data, particularly density functions. While recent papers have studied density data within the Bayes space framework, there has been limited focus on developing robust methods to effectively handle anomalous observations and large noise. To address this, we extend the Mahalanobis distance concept to Bayes spaces, proposing its regularized version that accounts for the constraints inherent in density data. Based on this extension, we introduce a new method, robust density principal component analysis (RDPCA), for more accurate estimation of functional principal components in the presence of outliers. The method's performance is validated through simulations and real-world applications, showing its ability to improve covariance estimation and principal component analysis compared to traditional methods.

stat.ME

funcharts: Control charts for multivariate functional data in R

Modern statistical process monitoring (SPM) applications focus on profile monitoring, i.e., the monitoring of process quality characteristics that can be modeled as profiles, also known as functional data. Despite the large interest in the profile monitoring literature, there is still a lack of software to facilitate its practical application. This article introduces the funcharts R package that implements recent developments on the SPM of multivariate functional quality characteristics, possibly adjusted by the influence of additional variables, referred to as covariates. The package also implements the real-time version of all control charting procedures to monitor profiles partially observed up to an intermediate domain point. The package is illustrated both through its built-in data generator and a real-case study on the SPM of Ro-Pax ship CO2 emissions during navigation, which is based on the ShipNavigation data provided in the Supplementary Material.

stat.CO

Robust Functional ANOVA with Application to Additive Manufacturing

The development of data acquisition systems is facilitating the collection of data that are apt to be modelled as functional data. In some applications, the interest lies in the identification of significant differences in group functional means defined by varying experimental conditions, which is known as functional analysis of variance (FANOVA). With real data, it is common that the sample under study is contaminated by some outliers, which can strongly bias the analysis. In this paper, we propose a new robust nonparametric functional ANOVA method (RoFANOVA) that reduces the weights of outlying functional data on the results of the analysis. It is implemented through a permutation test based on a test statistic obtained via a functional extension of the classical robust $ M $-estimator. By means of an extensive Monte Carlo simulation study, the proposed test is compared with some alternatives already presented in the literature, in both one-way and two-way designs. The performance of the RoFANOVA is demonstrated in the framework of a motivating real-case study in the field of additive manufacturing that deals with the analysis of spatter ejections. The RoFANOVA method is implemented in the R package rofanova, available online at https://github.com/unina-sfere/rofanova.

stat.AP

A new class of $α$-transformations for the spatial analysis of Compositional Data

Georeferenced compositional data are prominent in many scientific fields and in spatial statistics. This work addresses the problem of proposing models and methods to analyze and predict, through kriging, this type of data. To this purpose, a novel class of transformations, named the Isometric $α$-transformation ($α$-IT), is proposed, which encompasses the traditional Isometric Log-Ratio (ILR) transformation. It is shown that the ILR is the limit case of the $α$-IT as $α$ tends to 0 and that $α=1$ corresponds to a linear transformation of the data. Unlike the ILR, the proposed transformation accepts 0s in the compositions when $α>0$. Maximum likelihood estimation of the parameter $α$ is established. Prediction using kriging on $α$-IT transformed data is validated on synthetic spatial compositional data, using prediction scores computed either in the geometry induced by the $α$-IT, or in the simplex. Application to land cover data shows that the relative superiority of the various approaches w.r.t. a prediction objective depends on whether the compositions contained any zero component. When all components are positive, the limit cases (ILR or linear transformations) are optimal for none of the considered metrics. An intermediate geometry, corresponding to the $α$-IT with maximum likelihood estimate, better describes the dataset in a geostatistical setting. When the amount of compositions with 0s is not negligible, some side-effects of the transformation gets amplified as $α$ decreases, entailing poor kriging performances both within the $α$-IT geometry and for metrics in the simplex.

stat.ME

Social and material vulnerability in the face of seismic hazard: an analysis of the Italian case

The assessment of the vulnerability of a community endangered by seismic hazard is of paramount importance for planning a precision policy aimed at the prevention and reduction of its seismic risk. We aim at measuring the vulnerability of the Italian municipalities exposed to seismic hazard, by analyzing the open data offered by the Mappa dei Rischi dei Comuni Italiani provided by ISTAT, the Italian National Institute of Statistics. Encompassing the Index of Social and Material Vulnerability already computed by ISTAT, we also consider as referents of the latent social and material vulnerability of a community, its demographic dynamics and the age of the building stock where the community resides. Fusing the analyses of different indicators, within the context of seismic risk we offer a tentative ranking of the Italian municipalities in terms of their social and material vulnerability, together with differential profiles of their dominant fragilities which constitute the basis for planning precision policies aimed at seismic risk prevention and reduction.

stat.AP

Bivariate Densities in Bayes Spaces: Orthogonal Decomposition and Spline Representation

A new orthogonal decomposition for bivariate probability densities embedded in Bayes Hilbert spaces is derived. It allows one to represent a density into independent and interactive parts, the former being built as the product of revised definitions of marginal densities and the latter capturing the dependence between the two random variables being studied. The developed framework opens new perspectives for dependence modelling (which is commonly performed through copulas), and allows for the analysis of dataset of bivariate densities, in a Functional Data Analysis perspective. A spline representation for bivariate densities is also proposed, providing a computational cornerstone for the developed theory.

math.ST

Adaptive Smoothing Spline Estimator for the Function-on-Function Linear Regression Model

In this paper, we propose an adaptive smoothing spline (AdaSS) estimator for the function-on-function linear regression model where each value of the response, at any domain point, depends on the full trajectory of the predictor. The AdaSS estimator is obtained by the optimization of an objective function with two spatially adaptive penalties, based on initial estimates of the partial derivatives of the regression coefficient function. This allows the proposed estimator to adapt more easily to the true coefficient function over regions of large curvature and not to be undersmoothed over the remaining part of the domain. A novel evolutionary algorithm is developed ad hoc to obtain the optimization tuning parameters. Extensive Monte Carlo simulations have been carried out to compare the AdaSS estimator with competitors that have already appeared in the literature before. The results show that our proposal mostly outperforms the competitor in terms of estimation and prediction accuracy. Lastly, those advantages are illustrated also on two real-data benchmark examples.

stat.ME

A novel dowscaling procedure for compositional data in the Aitchison geometry with application to soil texture data

In this work, we present a novel downscaling procedure for compositional quantities based on the Aitchison geometry. The method is able to naturally consider compositional constraints, i.e. unit-sum and positivity. We show that the method can be used in a block sequential Gaussian simulation framework in order to assess the variability of downscaled quantities. Finally, to validate the method, we test it first in an idealized scenario and then apply it for the downscaling of digital soil maps on a more realistic case study. The digital soil maps for the realistic case study are obtained from SoilGrids, a system for automated soil mapping based on state-of-the-art spatial predictions methods.

stat.ME

Kriging Riemannian Data via Random Domain Decompositions

Data taking value on a Riemannian manifold and observed over a complex spatial domain are becoming more frequent in applications, e.g. in environmental sciences and in geoscience. The analysis of these data needs to rely on local models to account for the non stationarity of the generating random process, the non linearity of the manifold and the complex topology of the domain. In this paper, we propose to use a random domain decomposition approach to estimate an ensemble of local models and then to aggregate the predictions of the local models through Fréchet averaging. The algorithm is introduced in complete generality and is valid for data belonging to any smooth Riemannian manifold but it is then described in details for the case of the manifold of positive definite matrices, the hypersphere and the Cholesky manifold. The predictive performance of the method are explored via simulation studies for covariance matrices and correlation matrices, where the Cholesky manifold geometry is used. Finally, the method is illustrated on an environmental dataset observed over the Chesapeake Bay (USA).

stat.ME