SearcharxivSearch

arXiv subjects

Matteo Pegoraro

Publications and source records attributed to Matteo Pegoraro.

16 recordsLinked to original sources

Weighted persistence intensity regression

Persistence diagrams summarize the multiscale topological structure of data, and in applications they often arrive paired with covariates. We develop nonparametric methodology and theory for estimating the expected weighted persistence diagram conditional on a Euclidean covariate. Representing each weighted diagram as a finite random measure on a compact window, we take the density of its conditional expectation as the regression target, the conditional weighted persistence intensity. For a conditional double-kernel estimator we establish finite-sample sup-norm rates with a matching minimax lower bound, uniform rates in partial optimal transport, and an unbiased-risk cross-validation criterion for bandwidth selection. Simulations with analytically known intensities corroborate the theory and show that cross-validation selects the oracle candidate bandwidth in the exact-intensity design. The method is illustrated by studying how radial geometry in cerebral artery trees varies with age.

math.ST

Building confidence regions for Reeb graphs using the interleaving distance

We develop confidence regions for Reeb graphs from finite samples using the interleaving distance. Given a point cloud equipped with a filter function, we construct a finite proximity graph, extend the filter linearly, and use the Reeb cosheaf of the resulting filtered graph as the primary estimator. Mapper graphs are then treated as controlled cover-based coarsenings of this estimator, separating the statistical approximation problem from the visualization problem. We prove stability bounds for the Reeb estimators obtained both using intrinsic and extrinsic metrics, the latter under positive-reach assumptions, and derive interleaving-distance confidence regions from either \((a,b)\)-standard sampling assumptions or subsampling-based Hausdorff scale estimates. We also compare this object-level metric viewpoint with persistence-based guarantees by showing that the extended-persistence pseudometric is bounded by twice the interleaving distance, with sharp constant \(1\) for the \(H_0\)-related components. Numerical experiments illustrate how statistically significant features can be identified and then projected to Mapper graphs for interpretation.

math.ST

Persistence Spheres: a Bi-continuous Linear Representation of Measures for Partial Optimal Transport

We improve and extend persistence spheres, introduced in~\cite{pegoraro2025persistence}. Persistence spheres map an integrable measure $\mu$ on the upper half-plane, including persistence diagrams (PDs) as counting measures, to a function $S(\mu)\in C(\mathbb{S}^2)$, and the map is stable with respect to 1-Wasserstein partial transport distance $\mathrm{POT}_1$. Moreover, to the best of our knowledge, persistence spheres are the first explicit representation used in topological machine learning for which continuity of the inverse on the image is established at every compactly supported target. Recent bounded-cardinality bi-Lipschitz embedding results in partial transport spaces, despite being powerful, are not given by the kind of explicit summary map considered here. Our construction is rooted in convex geometry: for positive measures, the defining ReLU integral is the support function of the lift zonoid. Building on~\cite{pegoraro2025persistence}, we refine the definition to better match the $\mathrm{POT}_1$ deletion mechanism, encoding partial transport via a signed diagonal augmentation. In particular, for integrable $\mu$, the uniform norm between $S(0)$ and $S(\mu)$ depends only on the persistence of $\mu$, without any need of ad-hoc re-weightings, reflecting optimal transport to the diagonal at persistence cost. This yields a parameter-free representation at the level of measures (up to numerical discretization), while accommodating future extensions where $\mu$ is a smoothed measure derived from PDs (e.g., persistence intensity functions~\citep{wu2024estimation}). Across clustering, regression, and classification tasks involving functional data, time series, graphs, meshes, and point clouds, the updated persistence spheres are competitive and often improve upon persistence images, persistence landscapes, persistence splines, and sliced Wasserstein kernel baselines.

stat.ML

Persistence Spheres: Bi-continuous Representations of Persistence Diagrams

We introduce persistence spheres, a novel functional representation of persistence diagrams. Unlike existing embeddings (such as persistence images, landscapes, or kernel methods), persistence spheres provide a bi-continuous mapping: they are Lipschitz continuous with respect to the 1-Wasserstein distance and admit a continuous inverse on their image. This ensures, in a theoretically optimal way, both stability and geometric fidelity, making persistence spheres the representation that most closely mirrors the Wasserstein geometry of PDs in linear space. We derive explicit formulas for persistence spheres, showing that they can be computed efficiently and parallelized with minimal overhead. Empirically, we evaluate them on diverse regression and classification tasks involving functional data, time series, graphs, meshes, and point clouds. Across these benchmarks, persistence spheres consistently deliver state-of-the-art or competitive performance compared to persistence images, persistence landscapes, and the sliced Wasserstein kernel.

cs.LG

Persistence diagrams for exploring the shape variability of abdominal aortic aneurysms

Abdominal Aortic Aneurysm consists of a permanent dilation in the abodminal portion of the aorta and, along with its associated pathologies like calcifications and intraluminal thrombi, is one of the most important pathologies of the circulatory system. The shape of the aorta is among the primary drivers for these health issues, with particular reference to all the characteristics which affects the hemodynamics. Starting from the computed tomography angiography of a patient, we propose to summarize such information using tools derived from Topological Data Analysis, obtaining persistence diagrams which describe the irregularities of the lumen of the aorta. We showcase the effectiveness of such shape-related descriptors with a series of supervised and unsupervised case studies.

physics.med-ph

Circular Max-Flow for Periodic Data via Reeb Graphs

We introduce a max-flow framework for data with periodic boundary conditions, motivated by the analysis of transport in atomistic materials. Starting from a space X equipped with a map into the circle encoding a chosen periodic direction, we use the associated Reeb graph to reduce the geometry of X to a directed one-dimensional tunnel network. We then augment this graph with capacity constraints derived from cross-sectional integrals with respect to Hausdorff measure, so that edge capacities represent bottlenecks in the corresponding level-set components. To obtain a scalar transport descriptor from this capacity-augmented directed Reeb graph, we define circular max-flow for directed graphs mapped to the circle. Unlike classical source-target max-flow, this formulation does not require choosing an inlet and an outlet, and is therefore intrinsic to the periodic setting. We show that circular max-flow can be computed through a linear optimization problem related to minimum-cost circulations, and we prove that its value agrees with the flow obtained on the periodically unrolled graph. We also prove the continuity results needed to justify the capacity construction and verify that the assumptions cover void spaces arising from finite thickened backbones in the torus. The appendix illustrates the framework on simulated periodic point clouds and reports results from a separate materials-science application to self-diffusion in glasses.

math.AT

A Persistence-Driven Edit Distance for Trees with Abstract Weights

In this work we define a novel edit distance for trees considered with some abstract weights on the edges. The metric is driven by the idea of considering trees as topological summaries in the context of persistence and topological data analysis. Several examples related to persistent sets are presented. The metric can be computed with a dynamical binary linear programming approach. This framework is applied and further studied in other works focused on merge trees, where the problems of stability and merge trees estimation are also assessed.

math.CO

Wasserstein Principal Component Analysis for Circular Measures

We consider the 2-Wasserstein space of probability measures supported on the unit-circle, and propose a framework for Principal Component Analysis (PCA) for data living in such a space. We build on a detailed investigation of the optimal transportation problem for measures on the unit-circle which might be of independent interest. In particular, we derive an expression for optimal transport maps in (almost) closed form and propose an alternative definition of the tangent space at an absolutely continuous probability measure, together with the associated exponential and logarithmic maps. PCA is performed by mapping data on the tangent space at the Wasserstein barycentre, which we approximate via an iterative scheme, and for which we establish a sufficient a posteriori condition to assess its convergence. Our methodology is illustrated on several simulated scenarios and a real data analysis of measurements of optical nerve thickness.

stat.ME

Imaging-based representation and stratification of intra-tumor Heterogeneity via tree-edit distance

Personalized medicine is the future of medical practice. In oncology, tumor heterogeneity assessment represents a pivotal step for effective treatment planning and prognosis prediction. Despite new procedures for DNA sequencing and analysis, non-invasive methods for tumor characterization are needed to impact on daily routine. On purpose, imaging texture analysis is rapidly scaling, holding the promise to surrogate histopathological assessment of tumor lesions. In this work, we propose a tree-based representation strategy for describing intra-tumor heterogeneity of patients affected by metastatic cancer. We leverage radiomics information extracted from PET/CT imaging and we provide an exhaustive and easily readable summary of the disease spreading. We exploit this novel patient representation to perform cancer subtyping according to hierarchical clustering technique. To this purpose, a new heterogeneity-based distance between trees is defined and applied to a case study of prostate cancer. Clusters interpretation is explored in terms of concordance with severity status, tumor burden and biological characteristics. Results are promising, as the proposed method outperforms current literature approaches. Ultimately, the proposed method draws a general analysis framework that would allow to extract knowledge from daily acquired imaging data of patients and provide insights for effective treatment planning.

stat.ME

A Graph-Matching Formulation of the Interleaving Distance between Merge Trees

In this work we study the interleaving distance between merge trees from a combinatorial point of view. We use a particular type of matching between trees to obtain a novel formulation of the distance. With such formulation, we tackle the problem of approximating the interleaving distance by solving linear binary optimization problems in a recursive and dynamical fashion, obtaining lower and upper bounds. We implement those algorithms to compare the outputs with another approximation procedure presented by other authors. We believe that further research in this direction could lead to polynomial time algorithms to approximate the distance and novel theoretical developments on the topic.

math.CO

Projected Statistical Methods for Distributional Data on the Real Line with the Wasserstein Metric

We present a novel class of projected methods, to perform statistical analysis on a data set of probability distributions on the real line, with the 2-Wasserstein metric. We focus in particular on Principal Component Analysis (PCA) and regression. To define these models, we exploit a representation of the Wasserstein space closely related to its weak Riemannian structure, by mapping the data to a suitable linear space and using a metric projection operator to constrain the results in the Wasserstein space. By carefully choosing the tangent point, we are able to derive fast empirical methods, exploiting a constrained B-spline approximation. As a byproduct of our approach, we are also able to derive faster routines for previous work on PCA for distributions. By means of simulation studies, we compare our approaches to previously proposed methods, showing that our projected PCA has similar performance for a fraction of the computational cost and that the projected regression is extremely flexible even under misspecification. Several theoretical properties of the models are investigated and asymptotic consistency is proven. Two real world applications to Covid-19 mortality in the US and wind speed forecasting are discussed.

stat.ME

A Finitely Stable Edit Distance for Merge Trees

In this paper we define a novel edit distance for merge trees, which we argue to be suitable for a good range of applications. Relying also on some technical results contained in other works, we investigate its stability properties, which end up being analogous to the ones of the 1-Wasserstein distance between persistence diagrams. In the appendix, we extensively compare our metric in relationship with other metrics appearing in the literature, with both theoretic and practical considerations and a simulation.

math.MG

A Finitely Stable Edit Distance for Functions Defined on Merge Trees

In this work we define a metric structure to compare functions defined on different merge trees. The metric introduced possesses some stability properties, which we illustrate within a standard topological data analysis (TDA) framework, and can be computed with a dynamical binary linear programming approach. We showcase the effectiveness of the whole framework with simulated data sets. Using functions defined on merge trees proves to be very effective in situations where other topological data analysis tools, like persistence diagrams, cannot be used meaningfully.

math.CO

Functional Data Representation with Merge Trees

In this paper we face the problem of representation of functional data with the tools of algebraic topology. We represent functions by means of merge trees, which, like the more commonly used persistence diagrams, are invariant under homeomorphic reparametrizations of the functions they represent, thus allowing for a statistical analysis which is indifferent to functional misalignment. We consider a recently defined metric for merge trees and we prove some theoretical results related to its specific implementation when merge trees represent functions, establishing also a class of consistent estimators with convergence rates. To showcase the good properties of our topological approach to functional data analysis, we test it on the Aneurisk65 dataset replicating, from our different perspective, the supervised classification analysis which contributed to make this dataset a benchmark for methods dealing with misaligned functional data. In the Appendix we provide an extensive comparison between merge trees and persistence diagrams, highlighting similarities and differences, which can guide the analyst in choosing between the two representations.

stat.ME

Spatially dependent mixture models via the Logistic Multivariate CAR prior

We consider the problem of spatially dependent areal data, where for each area independent observations are available, and propose to model the density of each area through a finite mixture of Gaussian distributions. The spatial dependence is introduced via a novel joint distribution for a collection of vectors in the simplex, that we term logisticMCAR. We show that salient features of the logisticMCAR distribution can be described analytically, and that a suitable augmentation scheme based on the Pólya-Gamma identity allows to derive an efficient Markov Chain Monte Carlo algorithm. When compared to competitors, our model has proved to better estimate densities in different (disconnected) areal locations when they have different characteristics. We discuss an application on a real dataset of Airbnb listings in the city of Amsterdam, also showing how to easily incorporate for additional covariate information in the model.

stat.ME

Effects of breaking vibrational energy equipartition on measurements of temperature in macroscopic oscillators subject to heat flux

When the energy content of a resonant mode of a crystalline solid in thermodynamic equilibrium is directly measured, assuming that quantum effects can be neglected it coincides with temperature except for a proportionality factor. This is due to the principle of energy equipartition and the equilibrium hypothesis. However, most natural systems found in nature are not in thermodynamic equilibrium and thus the principle cannot be granted. We measured the extent to which the low-frequency modes of vibration of a solid can defy energy equipartition, in presence of a steady state heat flux, even close to equilibrium. We found, experimentally and numerically, that the energy separately associated with low frequency normal modes strongly depends on the heat flux, and decouples sensibly from temperature. A 4% in the relative temperature difference across the object around room temperature suffices to excite two modes of a macroscopic oscillator, as if they were at equilibrium, respectively, at temperatures about 20% and a factor 3.5 higher. We interpret the result in terms of new flux-mediated correlations between modes in the nonequilibrium state, which are absent at equilibrium.

cond-mat.stat-mech