SearcharxivSearch

arXiv subjects

Ian L. Dryden

Publications and source records attributed to Ian L. Dryden.

18 recordsLinked to original sources

Compositional regression using principal nested spheres

Regression with compositional responses is challenging due to the nonlinear geometry of the simplex and the limitations of Euclidean methods. We propose a regression framework for manifold-valued data based on mappings to statistically tractable intermediate spaces. For compositional data, responses are embedded in the positive orthant of the sphere and analysed using Principal Nested Spheres (PNS), yielding a cylindrical intermediate space with a circular leading score and Euclidean higher-order scores. Regression is performed in this intermediate space and fitted values are mapped back to the simplex. A simulation study demonstrates good performance of PNS-based regression. An application to environmental chemical exposure data illustrates the interpretability and practical utility of the method.

stat.ME

Principal Nested Cones

In many applications, the data lie on a type of cone, where there is a distinction between an overall scale variable and the remaining scale-free structure. For example, the joint size and shape of objects are points on a cone, where size represents scale, and shape is the scale-free structure. Dimension reduction is central in such applications, as shape data are often high-dimensional. Interactions between shape and size are widespread and of significant interest in real-world applications. However, most existing methods either lack a single notion of size or focus solely on shape, effectively removing size information. We propose Principal Nested Cones (PNC), a nonlinear dimension reduction framework that preserves both shape and size. PNC represents data through a sequence of nested hypercones and progressively projects observations onto lower-dimensional cone spaces. The resulting PNC scores provide low-dimensional representations that jointly capture size-shape variation in an interpretable manner. To enable scalable computation in ultra-high-dimensional settings, we develop a fast approximation combining PCA-based transformation with standard PNC. Simulation studies and real data applications demonstrate that PNC captures nonlinear size-shape structure, improves representation and reconstruction, and yields interpretable insights across morphometric, developmental, and molecular datasets.

stat.ME

Principal nested spheres for high-dimensional data

The method of Principal Nested Spheres (PNS) is a non-linear dimension reduction technique for spherical data. The method is a backwards fitting procedure, starting with fitting a high-dimensional sphere and then successively reducing dimension at each stage. After reviewing the PNS method in detail, we introduce some new methods for model selection at each stage between great and small subspheres, based on the Kolmogorov-Smirnov test, a variance test and a likelihood ratio test. The current PNS fitting method is slow for high-dimensional spherical data, and so we introduce a fast PNS method which involves an initial principal components analysis decomposition to select a basis for lower dimensional PNS. A new visual method called the PNS biplot is introduced for examining the effects of the original variables on the PNS, and this involves procedures for back-fitting from the PNS scores back to the original variables. The methodology is illustrated with two high-dimensional datasets from cancer research: Melanoma proteomics data with 500 variables and 205 patients, and a Pan Cancer dataset with 12,478 genes and 300 patients. In both applications the PNS biplot is used to select variables for effective classification.

stat.ME

Urban mapping in Dar es Salaam using AJIVE

Mapping deprivation in urban areas is important, for example for identifying areas of greatest need and planning interventions. Traditional ways of obtaining deprivation estimates are based on either census or household survey data, which in many areas is unavailable or difficult to collect. However, there has been a huge rise in the amount of new, non-traditional forms of data, such as satellite imagery and cell-phone call-record data, which may contain information useful for identifying deprivation. We use Angle-Based Joint and Individual Variation Explained (AJIVE) to jointly model satellite imagery data, cell-phone data, and survey data for the city of Dar es Salaam, Tanzania. We first identify interpretable low-dimensional structure from the imagery and cell-phone data, and find that we can use these to identify deprivation. We then consider what is gained from further incorporating the more traditional and costly survey data. We also introduce a scalar measure of deprivation as a response variable to be predicted, and consider various approaches to multiview regression, including using AJIVE scores as predictors.

stat.AP

Object oriented data analysis of surface motion time series in peatland landscapes

Peatlands account for 10% of UK land area, 80% of which are degraded to some degree, emitting carbon at a similar magnitude to oil refineries or landfill sites. A lack of tools for rapid and reliable assessment of peatland condition has limited monitoring of vast areas of peatland and prevented targeting areas urgently needing action to halt further degradation. Measured using interferometric synthetic aperture radar (InSAR), peatland surface motion is highly indicative of peatland condition, largely driven by the eco-hydrological change in the peatland causing swelling and shrinking of the peat substrate. The computational intensity of recent methods using InSAR time series to capture the annual functional structure of peatland surface motion becomes increasingly challenging as the sample size increases. Instead, we utilize the behavior of the entire peatland surface motion time series using object oriented data analysis to assess peatland condition. In a Gibbs sampling scheme, our cluster analysis based on the functional behavior of the surface motion time series finds features representative of soft/wet peatlands, drier/shrubby peatlands and thin/modified peatlands align with the clusters. The posterior distribution of the assigned peatland types enables the scale of peatland degradation to be assessed, which will guide future cost-effective decisions for peatland restoration.

stat.AP

Non-parametric regression for networks

Network data are becoming increasingly available, and so there is a need to develop suitable methodology for statistical analysis. Networks can be represented as graph Laplacian matrices, which are a type of manifold-valued data. Our main objective is to estimate a regression curve from a sample of graph Laplacian matrices conditional on a set of Euclidean covariates, for example in dynamic networks where the covariate is time. We develop an adapted Nadaraya-Watson estimator which has uniform weak consistency for estimation using Euclidean and power Euclidean metrics. We apply the methodology to the Enron email corpus to model smooth trends in monthly networks and highlight anomalous networks. Another motivating application is given in corpus linguistics, which explores trends in an author's writing style over time based on word co-occurrence networks.

stat.ME

Smoothing splines on Riemannian manifolds, with applications to 3D shape space

There has been increasing interest in statistical analysis of data lying in manifolds. This paper generalizes a smoothing spline fitting method to Riemannian manifold data based on the technique of unrolling and unwrapping originally proposed in Jupp and Kent (1987) for spherical data. In particular we develop such a fitting procedure for shapes of configurations in general $m$-dimensional Euclidean space, extending our previous work for two dimensional shapes. We show that parallel transport along a geodesic on Kendall shape space is linked to the solution of a homogeneous first-order differential equation, some of whose coefficients are implicitly defined functions. This finding enables us to approximate the procedure of unrolling and unwrapping by simultaneously solving such equations numerically, and so to find numerical solutions for smoothing splines fitted to higher dimensional shape data. This fitting method is applied to the analysis of some dynamic 3D peptide data.

stat.ME

Manifold valued data analysis of samples of networks, with applications in corpus linguistics

Networks arise in many applications, such as in the analysis of text documents, social interactions and brain activity. We develop a general framework for extrinsic statistical analysis of samples of networks, motivated by networks representing text documents in corpus linguistics. We identify networks with their graph Laplacian matrices, for which we define metrics, embeddings, tangent spaces, and a projection from Euclidean space to the space of graph Laplacians. This framework provides a way of computing means, performing principal component analysis and regression, and carrying out hypothesis tests, such as for testing for equality of means between two samples of networks. We apply the methodology to the set of novels by Jane Austen and Charles Dickens.

stat.ME

Principal nested shape space analysis of molecular dynamics data

Molecular dynamics simulations produce huge datasets of temporal sequences of molecules. It is of interest to summarize the shape evolution of the molecules in a succinct, low-dimensional representation. However, Euclidean techniques such as principal components analysis (PCA) can be problematic as the data may lie far from in a flat manifold. Principal nested spheres gives a fundamentally different decomposition of data from the usual Euclidean sub-space based PCA (Jung, Dryden and Marron, 2012, Biometrika). Sub-spaces of successively lower dimension are fitted to the data in a backwards manner, with the aim of retaining signal and dispensing with noise at each stage. We adapt the methodology to 3D sub-shape spaces and provide some practical fitting algorithms. The methodology is applied to cluster analysis of peptides, where different states of the molecules can be identified. Also, the temporal transitions between cluster states are explored.

stat.ME

Mixed Effect Modelling of Single Trial Variability in Ultra-High Field fMRI

Neuronal brain activity in response to repeated stimuli can be perceived using functional magnetic resonance imaging (fMRI). In this paper, we develop a statistical model for fMRI data that estimates both the associated haemodynamic response function and the within and between trial variability through maximum likelihood estimation. We discuss our results in the context of other model-driven approaches, extending models already popular in the literature, while removing the need for some of their assumptions. We consider an application to the motor cortex activity caused by a subject pressing a button and observe that the response changes significantly with task and through time.

stat.AP

Bayesian registration of functions and curves

Bayesian analysis of functions and curves is considered, where warping and other geometrical transformations are often required for meaningful comparisons. We focus on two applications involving the classification of mouse vertebrae shape outlines and the alignment of mass spectrometry data in proteomics. The functions and curves of interest are represented using the recently introduced square root velocity function, which enables a warping invariant elastic distance to be calculated in a straightforward manner. We distinguish between various spaces of interest: the original space, the ambient space after standardizing, and the quotient space after removing a group of transformations. Using Gaussian process models in the ambient space and Dirichlet priors for the warping functions, we explore Bayesian inference for curves and functions. Markov chain Monte Carlo algorithms are introduced for simulating from the posterior, including simulated tempering for multimodal posteriors. We also compare ambient and quotient space estimators for mean shape, and explain their frequent similarity in many practical problems using a Laplace approximation. A simulation study is carried out, as well as shape classification of the mouse vertebra outlines and practical alignment of the mass spectrometry functions.

stat.ME

Bayesian matching of unlabeled marked point sets using random fields, with an application to molecular alignment

Statistical methodology is proposed for comparing unlabeled marked point sets, with an application to aligning steroid molecules in chemoinformatics. Methods from statistical shape analysis are combined with techniques for predicting random fields in spatial statistics in order to define a suitable measure of similarity between two marked point sets. Bayesian modeling of the predicted field overlap between pairs of point sets is proposed, and posterior inference of the alignment is carried out using Markov chain Monte Carlo simulation. By representing the fields in reproducing kernel Hilbert spaces, the degree of overlap can be computed without expensive numerical integration. Superimposing entire fields rather than the configuration matrices of point coordinates thereby avoids the problem that there is usually no clear one-to-one correspondence between the points. In addition, mask parameters are introduced in the model, so that partial matching of the marked point sets can be carried out. We also propose an adaptation of the generalized Procrustes analysis algorithm for the simultaneous alignment of multiple point sets. The methodology is illustrated with a simulation study and then applied to a data set of 31 steroid molecules, where the relationship between shape and binding activity to the corticosteroid binding globulin receptor is explored.

stat.AP

Non-Euclidean statistical analysis of covariance matrices and diffusion tensors

The statistical analysis of covariance matrices occurs in many important applications, e.g. in diffusion tensor imaging and longitudinal data analysis. We consider the situation where it is of interest to estimate an average covariance matrix, describe its anisotropy, to carry out principal geodesic analysis and to interpolate between covariance matrices. There are many choices of metric available, each with its advantages. The particular choice of what is best will depend on the particular application. The use of the Procrustes size-and-shape metric is particularly appropriate when the covariance matrices are close to being deficient in rank. We discuss the use of different metrics for diffusion tensor analysis, and we also introduce certain types of regularization for tensors.

stat.ME

Bayesian matching of unlabelled point sets using Procrustes and configuration models

The problem of matching unlabelled point sets using Bayesian inference is considered. Two recently proposed models for the likelihood are compared, based on the Procrustes size-and-shape and the full configuration. Bayesian inference is carried out for matching point sets using Markov chain Monte Carlo simulation. An improvement to the existing Procrustes algorithm is proposed which improves convergence rates, using occasional large jumps in the burn-in period. The Procrustes and configuration methods are compared in a simulation study and using real data, where it is of interest to estimate the strengths of matches between protein binding sites. The performance of both methods is generally quite similar, and a connection between the two models is made using a Laplace approximation.

stat.CO

Power Euclidean metrics for covariance matrices with application to diffusion tensor imaging

Various metrics for comparing diffusion tensors have been recently proposed in the literature. We consider a broad family of metrics which is indexed by a single power parameter. A likelihood-based procedure is developed for choosing the most appropriate metric from the family for a given dataset at hand. The approach is analogous to using the Box-Cox transformation that is frequently investigated in regression analysis. The methodology is illustrated with a simulation study and an application to a real dataset of diffusion tensor images of canine hearts.

stat.ME

Non-Euclidean statistics for covariance matrices, with applications to diffusion tensor imaging

The statistical analysis of covariance matrix data is considered and, in particular, methodology is discussed which takes into account the non-Euclidean nature of the space of positive semi-definite symmetric matrices. The main motivation for the work is the analysis of diffusion tensors in medical image analysis. The primary focus is on estimation of a mean covariance matrix and, in particular, on the use of Procrustes size-and-shape space. Comparisons are made with other estimation techniques, including using the matrix logarithm, matrix square root and Cholesky decomposition. Applications to diffusion tensor imaging are considered and, in particular, a new measure of fractional anisotropy called Procrustes Anisotropy is discussed.

stat.AP

Statistical analysis on high-dimensional spheres and shape spaces

We consider the statistical analysis of data on high-dimensional spheres and shape spaces. The work is of particular relevance to applications where high-dimensional data are available--a commonly encountered situation in many disciplines. First the uniform measure on the infinite-dimensional sphere is reviewed, together with connections with Wiener measure. We then discuss densities of Gaussian measures with respect to Wiener measure. Some nonuniform distributions on infinite-dimensional spheres and shape spaces are introduced, and special cases which have important practical consequences are considered. We focus on the high-dimensional real and complex Bingham, uniform, von Mises-Fisher, Fisher-Bingham and the real and complex Watson distributions. Asymptotic distributions in the cases where dimension and sample size are large are discussed. Approximations for practical maximum likelihood based inference are considered, and in particular we discuss an application to brain shape modeling.

math.ST