SearcharxivSearch

arXiv subjects

Xavier Pennec

Publications and source records attributed to Xavier Pennec.

At least 19 recordsLinked to original sources

Log-Euclidean Lie Groups

We develop a self-contained theory of log-Euclidean Lie groups: smooth manifolds diffeomorphic to finite-dimensional vector spaces, equipped with the pullback of a constant Euclidean metric. This framework encompasses symmetric positive-definite (SPD) matrices S+(n) and full-rank correlation matrices Cor+(n), and explains why many seemingly different log-Euclidean constructions yield the same Riemannian geometry. We provide explicit Riemannian isometries (and Lie group isomorphisms) linking several log-Euclidean metrics on SPD and correlation matrix manifolds, and we characterize quotient log-Euclidean metrics in a principal-bundle setting. Finally, using the diagonal correction map underlying the off-log parametrization, we construct an explicit log-Euclidean metric on S+(n) for which the standard inclusion i\,: Cor+(n) $\rightarrow$ S+(n) becomes an isometric (indeed, totally geodesic) embedding, yielding closed-form formulas for geodesics and orthogonal decompositions in adapted coordinates. The nested isometric embeddings constructed here also provide a simple solution to the comparison of matrices of different dimensions in the log-Euclidean setting: SPD or correlation matrices may be transported to a common dimension via explicit maps while preserving all intrinsic distances.

math.DG

Barycentric subspace analysis of network-valued data

Certain data are naturally modeled by networks or weighted graphs, be they arterial networks or mobility networks. When there is no canonical labeling of the nodes across the dataset, we talk about unlabeled networks. In this paper, we focus on the question of dimensionality reduction for this type of data. More specifically, we address the issue of interpreting the feature subspace constructed by dimensionality reduction methods. Most existing methods for network-valued data are derived from principal component analysis (PCA) and therefore rely on subspaces generated by a set of vectors, which we identify as a major limitation in terms of interpretability. Instead, we propose to implement the method called barycentric subspace analysis (BSA), which relies on subspaces generated by a set of points. In order to provide a computationally feasible framework for BSA, we introduce a novel embedding for unlabeled networks where we replace their usual representation by equivalence classes of isomorphic networks with that by equivalence classes of cospectral networks. We then illustrate BSA on simulated and real-world datasets, and compare it to tangent PCA.

math.DG

Log-Euclidean Frameworks for Smooth Brain Connectivity Trajectories

The brain is often studied from a network perspective, where functional activity is assessed using functional Magnetic Resonance Imaging (fMRI) to estimate connectivity between predefined neuronal regions. Functional connectivity can be represented by correlation matrices computed over time, where each matrix captures the Pearson correlation between the mean fMRI signals of different regions within a sliding window. We introduce several Log-Euclidean Riemannian framework for constructing smooth approximations of functional brain connectivity trajectories. Representing dynamic functional connectivity as time series of full-rank correlation matrices, we leverage recent theoretical Log-Euclidean diffeomorphisms to map these trajectories in practice into Euclidean spaces where polynomial regression becomes feasible. Pulling back the regressed curve ensures that each estimated point remains a valid correlation matrix, enabling a smooth, interpretable, and geometrically consistent approximation of the original brain connectivity dynamics. Experiments on fMRI-derived connectivity trajectories demonstrate the geometric consistency and computational efficiency of our approach.

stat.AP

Parsimonious Gaussian mixture models with piecewise-constant eigenvalue profiles

Gaussian mixture models (GMMs) are ubiquitous in statistical learning, particularly for unsupervised problems. While full GMMs suffer from the overparameterization of their covariance matrices in high-dimensional spaces, spherical GMMs (with isotropic covariance matrices) certainly lack flexibility to fit certain anisotropic distributions. Connecting these two extremes, we introduce a new family of parsimonious GMMs with piecewise-constant covariance eigenvalue profiles. These extend several low-rank models like the celebrated mixtures of probabilistic principal component analyzers (MPPCA), by enabling any possible sequence of eigenvalue multiplicities. If the latter are prespecified, then we can naturally derive an expectation-maximization (EM) algorithm to learn the mixture parameters. Otherwise, to address the notoriously-challenging issue of jointly learning the mixture parameters and hyperparameters, we propose a componentwise penalized EM algorithm, whose monotonicity is proven. We show the superior likelihood-parsimony tradeoffs achieved by our models on a variety of unsupervised experiments: density fitting, clustering and single-image denoising.

stat.ML

Eigengap Sparsity for Covariance Parsimony

Covariance estimation is a central problem in statistics. An important issue is that there are rarely enough samples $n$ to accurately estimate the $p (p+1) / 2$ coefficients in dimension $p$. Parsimonious covariance models are therefore preferred, but the discrete nature of model selection makes inference computationally challenging. In this paper, we propose a relaxation of covariance parsimony termed "eigengap sparsity" and motivated by the good accuracy-parsimony tradeoffs of eigenvalue-equalization in covariance matrices. This penalty can be included in a penalized-likelihood framework that we propose to solve with a projected gradient descent on a monotone cone. The algorithm turns out to resemble an isotonic regression of mutually-attracted sample eigenvalues, drawing an interesting link between covariance parsimony and shrinkage.

stat.ME

Nested subspace learning with flags

Many machine learning methods look for low-dimensional representations of the data. The underlying subspace can be estimated by first choosing a dimension $q$ and then optimizing a certain objective function over the space of $q$-dimensional subspaces (the Grassmannian). Trying different $q$ yields in general non-nested subspaces, which raises an important issue of consistency between the data representations. In this paper, we propose a simple and easily implementable principle to enforce nestedness in subspace learning methods. It consists in lifting Grassmannian optimization criteria to flag manifolds (the space of nested subspaces of increasing dimension) via nested projectors. We apply the flag trick to several classical machine learning methods and show that it successfully addresses the nestedness issue.

stat.ML

Beyond Euclid: An Illustrated Guide to Modern Machine Learning with Geometric, Topological, and Algebraic Structures

The enduring legacy of Euclidean geometry underpins classical machine learning, which, for decades, has been primarily developed for data lying in Euclidean space. Yet, modern machine learning increasingly encounters richly structured data that is inherently nonEuclidean. This data can exhibit intricate geometric, topological and algebraic structure: from the geometry of the curvature of space-time, to topologically complex interactions between neurons in the brain, to the algebraic transformations describing symmetries of physical systems. Extracting knowledge from such non-Euclidean data necessitates a broader mathematical perspective. Echoing the 19th-century revolutions that gave rise to non-Euclidean geometry, an emerging line of research is redefining modern machine learning with non-Euclidean structures. Its goal: generalizing classical methods to unconventional data types with geometry, topology, and algebra. In this review, we provide an accessible gateway to this fast-growing field and propose a graphical taxonomy that integrates recent advances into an intuitive unified framework. We subsequently extract insights into current challenges and highlight exciting opportunities for future development in this field.

cs.LG

The curse of isotropy: from principal components to principal subspaces

Principal component analysis is a ubiquitous tool in exploratory data analysis. It is widely used by applied scientists for visualization and interpretability purposes. We raise an important issue (the curse of isotropy) about the interpretation of principal components with close eigenvalues. They may indeed suffer from an important rotational variability, which is a pitfall for interpretation. Through the lens of a probabilistic covariance model parameterized with flags of subspaces, we show that the curse of isotropy cannot be overlooked in practice. In this context, we propose to transition from ill-defined principal components to more-interpretable principal subspaces. The final methodology (principal subspace analysis) is extremely simple and shows promising results on a variety of datasets from different fields.

stat.ME

Principal subbundles for dimension reduction

In this paper we demonstrate how sub-Riemannian geometry can be used for manifold learning and surface reconstruction by combining local linear approximations of a point cloud to obtain lower dimensional bundles. Local approximations obtained by local PCAs are collected into a rank $k$ tangent subbundle on $\mathbb{R}^d$, $k<d$, which we call a principal subbundle. This determines a sub-Riemannian metric on $\mathbb{R}^d$. We show that sub-Riemannian geodesics with respect to this metric can successfully be applied to a number of important problems, such as: explicit construction of an approximating submanifold $M$, construction of a representation of the point-cloud in $\mathbb{R}^k$, and computation of distances between observations, taking the learned geometry into account. The reconstruction is guaranteed to equal the true submanifold in the limit case where tangent spaces are estimated exactly. Via simulations, we show that the framework is robust when applied to noisy data. Furthermore, the framework generalizes to observations on an a priori known Riemannian manifold.

stat.ME

Geometry of Sample Spaces

In statistics, independent, identically distributed random samples do not carry a natural ordering, and their statistics are typically invariant with respect to permutations of their order. Thus, an $n$-sample in a space $M$ can be considered as an element of the quotient space of $M^n$ modulo the permutation group. The present paper takes this definition of sample space and the related concept of orbit types as a starting point for developing a geometric perspective on statistics. We aim at deriving a general mathematical setting for studying the behavior of empirical and population means in spaces ranging from smooth Riemannian manifolds to general stratified spaces. We fully describe the orbifold and path-metric structure of the sample space when $M$ is a manifold or path-metric space, respectively. These results are non-trivial even when $M$ is Euclidean. We show that the infinite sample space exists in a Gromov-Hausdorff type sense and coincides with the Wasserstein space of probability distributions on $M$. We exhibit Fréchet means and $k$-means as metric projections onto 1-skeleta or $k$-skeleta in Wasserstein space, and we define a new and more general notion of polymeans. This geometric characterization via metric projections applies equally to sample and population means, and we use it to establish asymptotic properties of polymeans such as consistency and asymptotic normality.

math.ST

Flagfolds: an approach to multi-dimensional varifolds

By interpreting the product of the Principal Component Analysis, that is the covariance matrix, as a sequence of nested subspaces naturally coming with weights according to the level of approximation they provide, we are able to embed all $d$--dimensional Grassmannians into a stratified space of covariance matrices. We observe that Grassmannians constitute the lowest dimensional skeleton of the stratification while it is possible to define a Riemaniann metric on the highest dimensional and dense stratum, such a metric being compatible with the global stratification. With such a Riemaniann metric at hand, it is possible to look for geodesics between two linear subspaces of different dimensions that do not go through higher dimensional linear subspaces as would euclidean geodesics. Building upon the proposed embedding of Grassmannians into the stratified space of covariance matrices, we generalize the concept of varifolds to what we call flagfolds in order to model multi-dimensional shapes.

math.CA

Effective formulas for the geometry of normal homogeneous spaces. Application to flag manifolds

Consider a smooth manifold and an action on it of a compact connected Lie group with a bi-invariant metric. Then, any orbit is an embedded submanifold that is isometric to a normal homogeneous space for the group. In this paper, we establish new explicit and intrinsic formulas for the geometry of any such orbit. We derive our formula of the Levi-Civita connection from an existing generic formula for normal homogeneous spaces, i.e. which determines, a priori only theoretically, the connection. We say that our formulas are effective: they are directly usable, notably in numerical analysis, provided that the ambient manifold is convenient for computations. Then, we deduce new effective formulas for flag manifolds, since we prove that they are orbits under a suitable action of the special orthogonal group on a product of Grassmannians. This representation of them is quite useful, notably for studying flags of eigenspaces of symmetric matrices. Thus, we develop from it a geometric solution to the problem of perturbation of eigenvectors which is explicit and optimal in a certain sense, improving the classical analytic solution.

math.DG

A geometric framework for asymptotic inference of principal subspaces in PCA

In this article, we develop an asymptotic method for constructing confidence regions for the set of all linear subspaces arising from PCA, from which we derive hypothesis tests on this set. Our method is based on the geometry of Riemannian manifolds with which some sets of linear subspaces are endowed.

math.ST

Tangent phylogenetic PCA

Phylogenetic PCA (p-PCA) is a version of PCA for observations that are leaf nodes of a phylogenetic tree. P-PCA accounts for the fact that such observations are not independent, due to shared evolutionary history. The method works on Euclidean data, but in evolutionary biology there is a need for applying it to data on manifolds, particularly shapes. We provide a generalization of p-PCA to data lying on Riemannian manifolds, called Tangent p-PCA. Tangent p-PCA thus makes it possible to perform dimension reduction on a data set of shapes, taking into account both the non-linear structure of the shape space as well as phylogenetic covariance. We show simulation results on the sphere, demonstrating well-behaved error distributions and fast convergence of estimators. Furthermore, we apply the method to a data set of mammal jaws, represented as points on a landmark manifold equipped with the LDDMM metric.

stat.ME

Bures-Wasserstein minimizing geodesics between covariance matrices of different ranks

The set of covariance matrices equipped with the Bures-Wasserstein distance is the orbit space of the smooth, proper and isometric action of the orthogonal group on the Euclidean space of square matrices. This construction induces a natural orbit stratification on covariance matrices, which is exactly the stratification by the rank. Thus, the strata are the manifolds of symmetric positive semi-definite (PSD) matrices of fixed rank endowed with the Bures-Wasserstein Riemannian metric. In this work, we study the geodesics of the Bures-Wasserstein distance. Firstly, we complete the literature on geodesics in each stratum by clarifying the set of preimages of the exponential map and by specifying the injection domain. We also give explicit formulae of the horizontal lift, the exponential map and the Riemannian logarithms that were kept implicit in previous works. Secondly, we give the expression of all the minimizing geodesic segments joining two covariance matrices of any rank. More precisely, we show that the set of all minimizing geodesics between two covariance matrices $Σ$ and $Λ$ is parametrized by the closed unit ball of $\mathbb{R}^{(k-r)\times(l-r)}$ for the spectral norm, where $k, l, r$ are the respective ranks of $Σ$, $Λ$, $Σ$$Λ$. In particular, the minimizing geodesic is unique if and only if $r = \min(k, l)$. Otherwise, there are infinitely many.

math.DG

Theoretically and computationally convenient geometries on full-rank correlation matrices

In contrast to SPD matrices, few tools exist to perform Riemannian statistics on the open elliptope of full-rank correlation matrices. The quotient-affine metric was recently built as the quotient of the affine-invariant metric by the congruence action of positive diagonal matrices. The space of SPD matrices had always been thought of as a Riemannian homogeneous space. In contrast, we view in this work SPD matrices as a Lie group and the affine-invariant metric as a left-invariant metric. This unexpected new viewpoint allows us to generalize the construction of the quotient-affine metric and to show that the main Riemannian operations can be computed numerically. However, the uniqueness of the Riemannian logarithm or the Fr{é}chet mean are not ensured, which is bad for computing on the elliptope. Hence, we define three new families of Riemannian metrics on full-rank correlation matrices which provide Hadamard structures, including two flat. Thus the Riemannian logarithm and the Fr{é}chet mean are unique. We also define a nilpotent group structure for which the affine logarithm and the group mean are unique. We provide the main Riemannian/group operations of these four structures in closed form.

math.DG

Geodesic squared exponential kernel for non-rigid shape registration

This work addresses the problem of non-rigid registration of 3D scans, which is at the core of shape modeling techniques. Firstly, we propose a new kernel based on geodesic distances for the Gaussian Process Morphable Models (GPMMs) framework. The use of geodesic distances into the kernel makes it more adapted to the topological and geometric characteristics of the surface and leads to more realistic deformations around holes and curved areas. Since the kernel possesses hyperparameters we have optimized them for the task of face registration on the FaceWarehouse dataset. We show that the Geodesic squared exponential kernel performs significantly better than state of the art kernels for the task of face registration on all the 20 expressions of the FaceWarehouse dataset. Secondly, we propose a modification of the loss function used in the non-rigid ICP registration algorithm, that allows to weight the correspondences according to the confidence given to them. As a use case, we show that we can make the registration more robust to outliers in the 3D scans, such as non-skin parts.

cs.CV

The geometry of mixed-Euclidean metrics on symmetric positive definite matrices

Several Riemannian metrics and families of Riemannian metrics were defined on the manifold of Symmetric Positive Definite (SPD) matrices. Firstly, we formalize a common general process to define families of metrics: the principle of deformed metrics. We relate the recently introduced family of alpha-Procrustes metrics to the general class of mean kernel metrics by providing a sufficient condition under which elements of the former belongs to the latter. Secondly, we focus on the principle of balanced bilinear forms that we recently introduced. We give a new sufficient condition under which the balanced bilinear form is a metric. It allows us to introduce the Mixed-Euclidean (ME) metrics which generalize the Mixed-Power-Euclidean (MPE) metrics. We unveal their link with the (u, v)-divergences and the ($α$, $β$)-divergences of information geometry and we provide an explicit formula of the Riemann curvature tensor. We show that the sectional curvature of all ME metrics can take negative values and we show experimentally that the sectional curvature of all MPE metrics but the log-Euclidean, power-Euclidean and power-affine metrics can take positive values.

math.DG