SearcharxivSearch

arXiv subjects

Xing Qiu

Publications and source records attributed to Xing Qiu.

9 recordsLinked to original sources

Alignment of Continuous Brain Connectivity

Brain networks are typically represented by adjacency matrices, where each node corresponds to a brain region. In traditional brain network analysis, nodes are assumed to be matched across individuals, but the methods used for node matching often overlook the underlying connectivity information. This oversight can result in inaccurate node alignment, leading to inflated edge variability and reduced statistical power in downstream connectivity analyses. To overcome this challenge, we propose a novel framework for registering high resolution continuous connectivity (ConCon), defined as a continuous function on a product manifold space specifically, the cortical surface capturing structural connectivity between all pairs of cortical points. Leveraging ConCon, we formulate an optimal diffeomorphism problem to align both connectivity profiles and cortical surfaces simultaneously. We introduce an efficient algorithm to solve this problem and validate our approach using data from the Human Connectome Project (HCP). Results demonstrate that our method substantially improves the accuracy and robustness of connectome-based analyses compared to existing techniques.

stat.ME

Continuous and Atlas-free Analysis of Brain Structural Connectivity

Brain structural networks are often represented as discrete adjacency matrices with elements summarizing the connectivity between pairs of regions of interest (ROIs). These ROIs are typically determined a-priori using a brain atlas. The choice of atlas is often arbitrary and can lead to a loss of important connectivity information at the sub-ROI level. This work introduces an atlas-free framework that overcomes these issues by modeling brain connectivity using smooth random functions. In particular, we assume that the observed pattern of white matter fiber tract endpoints is driven by a latent random function defined over a product manifold domain. To facilitate statistical analysis of these high dimensional functional data objects, we develop a novel algorithm to construct a data-driven reduced-rank function space that offers a desirable trade-off between computational complexity and flexibility. Using real data from the Human Connectome Project, we show that our method outperforms state-of-the-art approaches that use the traditional atlas-based structural connectivity representation on a variety of connectivity analysis tasks. We further demonstrate how our method can be used to detect localized regions and connectivity patterns associated with group differences.

stat.CO

Tighter Bound Estimation for Efficient Biquadratic Optimization Over Unit Spheres

Bi-quadratic programming over unit spheres is a fundamental problem in quantum mechanics introduced by pioneer work of Einstein, Schrödinger, and others. It has been shown to be NP-hard; so it must be solve by efficient heuristic algorithms such as the block improvement method (BIM). This paper focuses on the maximization of bi-quadratic forms, which leads to a rank-one approximation problem that is equivalent to computing the M-spectral radius and its corresponding eigenvectors. Specifically, we provide a tight upper bound of the M-spectral radius for nonnegative fourth-order partially symmetric (PS) tensors, which can be considered as an approximation of the M-spectral radius. Furthermore, we showed that the proposed upper bound can be obtained more efficiently, if the nonnegative fourth-order PS-tensors is a member of certain monoid semigroups. Furthermore, as an extension of the proposed upper bound, we derive the exact solutions of the M-spectral radius and its corresponding M-eigenvectors for certain classes of fourth-order PS-tensors. Lastly, as an application of the proposed bound, we obtain a practically testable sufficient condition for nonsingular elasticity M-tensors with strong ellipticity condition. We conduct several numerical experiments to demonstrate the utility of the proposed results. The results show that: (a) our proposed method can attain a tight upper bound of the M-spectral radius with little computational burden, and (b) such tight and efficient upper bounds greatly enhance the convergence speed of the BIM-algorithm, allowing it to be applicable for large-scale problems in applications.

math.NA

Efficient Multidimensional Functional Data Analysis Using Marginal Product Basis Systems

Many modern datasets, from areas such as neuroimaging and geostatistics, come in the form of a random sample of tensor-valued data which can be understood as noisy observations of a smooth multidimensional random function. Most of the traditional techniques from functional data analysis are plagued by the curse of dimensionality and quickly become intractable as the dimension of the domain increases. In this paper, we propose a framework for learning continuous representations from a sample of multidimensional functional data that is immune to several manifestations of the curse. These representations are constructed using a set of separable basis functions that are defined to be optimally adapted to the data. We show that the resulting estimation problem can be solved efficiently by the tensor decomposition of a carefully defined reduction transformation of the observed data. Roughness-based regularization is incorporated using a class of differential operator-based penalties. Relevant theoretical properties are also established. The advantages of our method over competing methods are demonstrated in a simulation study. We conclude with a real data application in neuroimaging.

stat.ME

Hypothesis Testing for Two Sample Comparison of Network Data

Network data is a major object data type that has been widely collected or derived from common sources such as brain imaging. Such data contains numeric, topological, and geometrical information, and may be necessarily considered in certain non-Euclidean space for appropriate statistical analysis. The development of statistical methodologies for network data is challenging and currently at its infancy; for instance, the non-Euclidean counterpart of basic two-sample tests for network data is scarce in literature. In this study, a novel framework is presented for two independent sample comparison of networks. Specifically, an approximation distance metric to quotient Euclidean distance is proposed, and then combined with network spectral distance to quantify the local and global dissimilarity of networks simultaneously. A permutational non-Euclidean analysis of variance is adapted to the proposed distance metric for the comparison of two independent groups of networks. Comprehensive simulation studies and real applications are conducted to demonstrate the superior performance of our method over other alternatives. The asymptotic properties of the proposed test are investigated and its high-dimensional extension is discussed as well.

stat.ME

Identifiability Analysis of Linear Ordinary Differential Equation Systems with a Single Trajectory

Ordinary differential equations (ODEs) are widely used to model dynamical behavior of systems. It is important to perform identifiability analysis prior to estimating unknown parameters in ODEs (a.k.a. inverse problem), because if a system is unidentifiable, the estimation procedure may fail or produce erroneous and misleading results. Although several qualitative identifiability measures have been proposed, much less effort has been given to developing \emph{quantitative} (continuous) scores that are robust to uncertainties in the data, especially for those cases in which the data are presented as a single trajectory beginning with one initial value. In this paper, we first derived a closed-form representation of linear ODE systems that are not identifiable based on a single trajectory. This representation helps researchers design practical systems and choose the right prior structural information in practice. Next, we proposed several quantitative scores for identifiability analysis in practice. In simulation studies, the proposed measures outperformed the main competing method significantly, especially when noise was presented in the data. We also discussed the asymptotic properties of practical identifiability for high-dimensional ODE systems and conclude that, without additional prior information, many random ODE systems are practically unidentifiable when the dimension approaches infinity.

math.OC

Distributional Properties of Nearest-Site Angular Distances on the Sphere

Nearest-site distances arise in many applications involving spherical or directional domains, including global geospatial analysis, wireless communications, spherical clustering, and cosine-similarity-based data analysis. In this paper, we study the distributional and computational properties of $L_2$, the minimal angular great-circle distance from a uniformly distributed random point on a sphere to a set of prespecified sites on the same sphere. We first derive the cumulative distribution function (CDF) and probability density function (PDF) of $L_0$, the angular great-circle distance from a fixed vertex of a spherical triangle to a random point uniformly distributed within that triangle. We then extend these triangle-level results to convex spherical polygons and use spherical Voronoi diagrams, triangulations of Voronoi cells, and numerical integration to obtain computable distributional and moment formulas for $L_2$. In addition, we derive explicit formulas for selected moments of $\cos(L_2)$, which are relevant to cosine similarity and spherical data analysis. Extensive Monte Carlo simulations validate the proposed CDF, PDF, and moment formulas and demonstrate computational efficiency of our method relative to generic numerical integration and simulation-based alternatives.

stat.CO

Discussion of: Treelets--An adaptive multi-scale basis for sparse unordered data

This is a discussion of paper "Treelets--An adaptive multi-scale basis for sparse unordered data" [arXiv:0707.0481] by Ann B. Lee, Boaz Nadler and Larry Wasserman. In this paper the authors defined a new type of dimension reduction algorithm, namely, the treelet algorithm. The treelet method has the merit of being completely data driven, and its decomposition is easier to interpret as compared to PCR. It is suitable in some certain situations, but it also has its own limitations. I will discuss both the strength and the weakness of this method when applied to microarray data analysis.

stat.AP

Control of the mean number of false discoveries, Bonferroni and stability of multiple testing

The Bonferroni multiple testing procedure is commonly perceived as being overly conservative in large-scale simultaneous testing situations such as those that arise in microarray data analysis. The objective of the present study is to show that this popular belief is due to overly stringent requirements that are typically imposed on the procedure rather than to its conservative nature. To get over its notorious conservatism, we advocate using the Bonferroni selection rule as a procedure that controls the per family error rate (PFER). The present paper reports the first study of stability properties of the Bonferroni and Benjamini--Hochberg procedures. The Bonferroni procedure shows a superior stability in terms of the variance of both the number of true discoveries and the total number of discoveries, a property that is especially important in the presence of correlations between individual $p$-values. Its stability and the ability to provide strong control of the PFER make the Bonferroni procedure an attractive choice in microarray studies.

stat.AP