SearcharxivSearch

arXiv subjects

V. Masarotto

Publications and source records attributed to V. Masarotto.

2 recordsLinked to original sources

fdWasserstein: Optimal Transport Methods for Covariance Operators of Functional Data

Data increasingly arrive as collections of curves - a voice recording, a growth trajectory, a day of sensor readings - where each observation is a whole function rather than a single number. The usual question asked of such data is how the average curve differs from one group to the next. But the average is only half the picture: two populations of curves can share almost the same mean and still differ profoundly in how they fluctuate around it, and it is often this variability - the pattern of covariation within a curve - that carries the scientific signal. Comparing populations at this level means comparing their covariance operators, and statistics on covariances is impaired by their non-linearity. The fdWasserstein package equips R users with functions to make such comparisons. It is centered on the geometry of optimal transport, under which covariance operators can be meaningfully averaged, contrasted, and interpolated. It provides the Procrustes-Wasserstein distance between covariance operators, their Frechet mean (barycenter), an ANOVA-type permutation test for the equality of several covariances, principal component analysis of covariance variation, and an entropy-regularized soft clustering of curves by their covariance structure. We outline the underlying ideas, discuss the implementation and demonstrate the complete workflow on the phoneme data shipped with the package.

stat.CO

Covariance-based soft clustering of functional data based on the Wasserstein-Procrustes metric

We consider the problem of clustering functional data according to their covariance structure. We contribute a soft clustering methodology based on the Wasserstein-Procrustes distance, where the in-between cluster variability is penalised by a term proportional to the entropy of the partition matrix. In this way, each covariance operator can be partially classified into more than one group. Such soft classification allows for clusters to overlap, and arises naturally in situations where the separation between all or some of the clusters is not well-defined. We also discuss how to estimate the number of groups and to test for the presence of any cluster structure. The algorithm is illustrated using simulated and real data. An R implementation is available in the Supplementary materials.

stat.ME