SearcharxivSearch

arXiv subjects

Zhigang Yao

Publications and source records attributed to Zhigang Yao.

At least 19 recordsLinked to original sources

Manifold Fitting: A Review of Methods and Applications

With data growing in scale and complexity, traditional linear dimension reduction techniques are becoming inadequate in some settings. Manifold fitting offers an important alternative by capturing low-dimensional latent geometric structures within high-dimensional spaces. This capability allows it to support downstream analysis in complex data settings. In this review, we explore the development and applications of manifold fitting. First, we introduce the basic concepts of manifold fitting and distinguish it from related techniques such as manifold embedding and denoising. We review the development of manifold fitting with three distinct stages: early nonparametric statistical methods, insights from mathematical analysis, and contemporary practical statistical approaches. Furthermore, we present diverse applications of manifold fitting, particularly in neural networks and bioinformatics, which illustrate its utility in complex data scenarios. Despite considerable progress, manifold fitting remains a fertile area for research. Many theoretical and practical questions remain unanswered, and ongoing investigations will further clarify its role in modern data science as a geometric tool for a wide range of data analysis challenges.

stat.ME

Multidimensional physical fitness is associated with reduced dementia risk through proteomic and neuroimaging pathways: a prospective cohort study of the UK Biobank

Dementia affects over 55 million people worldwide, yet whether distinct domains of physical fitness independently protect against neurodegeneration through shared or divergent biological mechanisms remains unknown. Using the UK Biobank (n = 51,517; 12-year follow-up), we integrated epidemiological, proteomic, and neuroimaging analyses to systematically characterize the multidimensional fitness-dementia relationship. Higher handgrip strength, cardiorespiratory fitness, and pulmonary function were each independently associated with reduced dementia risk (HRs 0.50, 0.62, and 0.73, respectively, for highest vs. lowest tertiles), with stronger associations in women and younger individuals. Plasma proteomic profiling revealed domain-specific molecular signatures--neurofilament light chain predominating for muscular and cardiorespiratory fitness, and inflammatory mediators including GDF15 for pulmonary function--with 22-40 proteins per domain independently predicting dementia, converging on neuroinflammatory and neurovascular pathways. Brain MRI analyses identified hippocampal volume as a significant structural mediator (proportion mediated: 3.7-10.1%), indicating structural preservation as one of multiple mechanistic pathways. Population attributable fraction analyses estimated that suboptimal fitness may account for approximately 26% of dementia cases. These findings reveal that multidimensional physical fitness shapes dementia risk through distinct yet converging neuroinflammatory, neurovascular, and structural brain mechanisms, with implications for life-course prevention.

stat.AP

Routine Blood Biomarkers Reveal a Preclinical Continuum of Multiple Myeloma Risk

Multiple myeloma (MM) is preceded by a long preclinical phase spanning decades, yet scalable, non-specialist tools to identify individuals at elevated risk before end-organ damage are lacking. In a prospective analysis of 299,035 cancer-free UK Biobank participants followed for a median of 12.4 years, during which 768 developed incident MM, we conducted a biomarker-wide association scan across 61 routinely measured blood analytes spanning hematological, protein metabolism, renal, and immune categories. Markers of protein dysregulation-elevated total protein, depressed albumin, and a low albumin-to-globulin (A/G) ratio-showed the strongest preclinical associations (hazard ratios 0.61-1.54 per SD), consistent with progressive monoclonal immunoglobulin accumulation and suppression of normal polyclonal synthesis years before diagnosis. These signals were accompanied by indicators of erythropoietic suppression, morphological red cell dysregulation, and a shift toward lower neutrophil and higher lymphocyte fractions, reflecting coordinated perturbations across hematopoietic and immune compartments. Longitudinal trajectory analyses showed that these multi-system deviations emerge more than a decade before diagnosis and intensify as clinical onset approaches. Dose-response modelling revealed pronounced nonlinear associations for protein and erythrocytic markers, with risk concentrated among individuals with extreme values. Incorporating significant biomarkers into a clinical risk model improved 10-year MM discrimination from a C-index of 0.684 to 0.744, with the high-risk decile accumulating 0.79% cumulative incidence versus 0.47% under the clinical model alone. These findings provide a practical framework for biomarker-guided MM risk stratification and targeted surveillance using routinely available clinical tests.

stat.AP

Differentially Private Manifold Denoising

We introduce a differentially private manifold denoising framework that allows users to exploit sensitive reference datasets to correct noisy, non-private query points without compromising privacy. The method follows an iterative procedure that (i) privately estimates local means and tangent geometry using the reference data under calibrated sensitivity, (ii) projects query points along the privately estimated subspace toward the local mean via corrective steps at each iteration, and (iii) performs rigorous privacy accounting across iterations and queries using $(\varepsilon,δ)$-differential privacy (DP). Conceptually, this framework brings differential privacy to manifold methods, retaining sufficient geometric signal for downstream tasks such as embedding, clustering, and visualization, while providing formal DP guarantees for the reference data. Practically, the procedure is modular and scalable, separating DP-protected local geometry (means and tangents) from budgeted query-point updates, with a simple scheduler allocating privacy budget across iterations and queries. Under standard assumptions on manifold regularity, sampling density, and measurement noise, we establish high-probability utility guarantees showing that corrected queries converge toward the manifold at a non-asymptotic rate governed by sample size, noise level, bandwidth, and the privacy budget. Simulations and case studies demonstrate accurate signal recovery under moderate privacy budgets, illustrating clear utility-privacy trade-offs and providing a deployable DP component for manifold-based workflows in regulated environments without reengineering privacy systems.

cs.LG

Estimation of Riemannian Quantities from Noisy Data via Density Derivatives

We study the recovery of geometric structure from data generated by convolving the uniform measure on a smooth compact submanifold $M\subset\mathbb{R}^D$ with ambient Gaussian noise. Our main result is that several fundamental Riemannian quantities of $M$, including tangent spaces, the intrinsic dimension, and the second fundamental form, are identifiable from derivatives of the noisy density. We first derive uniform small-noise expansions of the data density and its derivatives in a tubular neighborhood of $M$. These expansions show that, at the population level, tangent spaces can be recovered from the density Hessian with $O(σ^2)$ error, while the intrinsic dimension can be estimated consistently. We further construct estimators for the second fundamental form from density derivatives, obtaining $O(d(y,M)+σ)$ and $O(d(y,M)+σ^2)$ errors for hypersurfaces and submanifolds with arbitrary codimension. At the sample level, we estimate the density and its derivatives by kernel methods in the ambient space and plug them into the population constructions, yielding uniform nonparametric rates in the ambient dimension. Finally, we show that these density-based constructions admit a geometric interpretation through density-induced ambient metrics, linking the geometry of $M$ to ambient geodesic structure.

math.ST

Principal Decomposition with Nested Submanifolds

Over the past decades, the increasing dimensionality of data has increased the need for effective data decomposition methods. Existing approaches, however, often rely on linear models or lack sufficient interpretability or flexibility. To address this issue, we introduce a novel nonlinear decomposition technique called the principal nested submanifolds, which builds on the foundational concepts of principal component analysis. This method exploits the local geometric information of data sets by projecting samples onto a series of nested principal submanifolds with progressively decreasing dimensions. It effectively isolates complex information within the data in a backward stepwise manner by targeting variations associated with smaller eigenvalues in local covariance matrices. Unlike previous methods, the resulting subspaces are smooth manifolds, not merely linear spaces or special shape spaces. Validated through extensive simulation studies and applied to real-world RNA sequencing data, our approach surpasses existing models in delineating intricate nonlinear structures. It provides more flexible subspace constraints that improve the extraction of significant data components and facilitate noise reduction. This innovative approach not only advances the non-Euclidean statistical analysis of data with low-dimensional intrinsic structure within Euclidean spaces, but also offers new perspectives for dealing with high-dimensional noisy data sets in fields such as bioinformatics and machine learning.

stat.ME

A Principal Submanifold-based Approach for Clustering and Multiscale RNA Correction

RNA structure determination is essential for understanding its biological functions. However, the reconstruction process often faces challenges, such as atomic clashes, which can lead to inaccurate models. To address these challenges, we introduce the principal submanifold (PSM) approach for analyzing RNA data on a torus. This method provides an accurate, low-dimensional feature representation, overcoming the limitations of previous torus-based methods. By combining PSM with DBSCAN, we propose a novel clustering technique, the principal submanifold-based DBSCAN (PSM-DBSCAN). Our approach achieves superior clustering accuracy and increased robustness to noise. Additionally, we apply this new method for multiscale corrections, effectively resolving RNA backbone clashes at both microscopic and mesoscopic scales. Extensive simulations and comparative studies highlight the enhanced precision and scalability of our method, demonstrating significant improvements over existing approaches. The proposed methodology offers a robust foundation for correcting complex RNA structures and has broad implications for applications in structural biology and bioinformatics.

q-bio.BM

Curvature-driven manifold fitting under unbounded isotropic noise

Manifold fitting aims to reconstruct a low-dimensional manifold from high-dimensional data, whose framework is established by Fefferman et al. \cite{fefferman2020reconstruction,fefferman2021reconstruction}. This paper studies the recovery of a compact $C^3$ submanifold $\mathcal{M} \subset \mathbb{R}^D$ with dimension $d 0$ and achieves a state-of-the-art Hausdorff distance of $O(σ^2)$ to $\mathcal{M}$. Numerical experiments confirm the quadratic decay of the reconstruction error and demonstrate the computational efficiency of the estimator $F$. Our work provides a curvature-driven framework for denoising and reconstructing manifolds with second-order accuracy.

math.ST

Beyond Missing Data: Questionnaire Uncertainty Responses as Early Digital Biomarkers of Cognitive Decline and Neurodegenerative Diseases

Identifying preclinical biomarkers of neurodegenerative diseases remains a major challenge in aging research. In this study, we demonstrate that frequent "Don't know/can't remember" (DK) responses, often treated as missing data in touchscreen questionnaires, serve as a novel digital behavioral biomarker of early cognitive vulnerability and neurodegenerative disease risk. Using data from 502,234 UK Biobank participants, we stratified individuals based on DK response frequency (0-1, 2-4, 5-7, >7) and observed a robust, dose-dependent association with an increased risk of Alzheimer's disease (HR = 1.64, 95% CI: 1.26-2.14) and vascular dementia (HR = 1.93, 95% CI: 1.37-2.72), independent of established risk factors. As DK response frequency increased, participants exhibited higher BMI, reduced physical activity, higher smoking rates, and a higher prevalence of chronic diseases, particularly hypertension, diabetes, and depression. Further analysis revealed a dose-dependent relationship between DK response frequency and the risk of Alzheimer's disease and vascular dementia, with high DK responders showing early neurodegenerative changes, marked by elevated levels of Abeta40, Abeta42, NFL, and pTau-181. Metabolomic analysis also revealed lipid metabolism abnormalities, which may mediate this relationship. Together, these findings reframe DK response patterns as clinically meaningful signals of multidimensional neurobiological alterations, offering a scalable, low-cost, non-invasive tool for early risk identification and prevention at the population level.

stat.AP

Enhanced High-Dimensional Data Visualization through Adaptive Multi-Scale Manifold Embedding

To address the dual challenges of the curse of dimensionality and the difficulty in separating intra-cluster and inter-cluster structures in high-dimensional manifold embedding, we proposes an Adaptive Multi-Scale Manifold Embedding (AMSME) algorithm. By introducing ordinal distance to replace traditional Euclidean distances, we theoretically demonstrate that ordinal distance overcomes the constraints of the curse of dimensionality in high-dimensional spaces, effectively distinguishing heterogeneous samples. We design an adaptive neighborhood adjustment method to construct similarity graphs that simultaneously balance intra-cluster compactness and inter-cluster separability. Furthermore, we develop a two-stage embedding framework: the first stage achieves preliminary cluster separation while preserving connectivity between structurally similar clusters via the similarity graph, and the second stage enhances inter-cluster separation through a label-driven distance reweighting. Experimental results demonstrate that AMSME significantly preserves intra-cluster topological structures and improves inter-cluster separation on real-world datasets. Additionally, leveraging its multi-resolution analysis capability, AMSME discovers novel neuronal subtypes in the mouse lumbar dorsal root ganglion scRNA-seq dataset, with marker gene analysis revealing their distinct biological roles.

cs.LG

Ridge Estimation with Nonlinear Transformations

Ridge estimation is an important manifold learning technique. The goal of this paper is to examine the effects of nonlinear transformations on the ridge sets. The main result proves the inclusion relationship between ridges: $\cR(f\circ p)\subseteq \cR(p)$, provided that the transformation $f$ is strictly increasing and concave on the range of the function $p$. Additionally, given an underlying true manifold $\cM$, we show that the Hausdorff distance between $\cR(f\circ p)$ and its projection onto $\cM$ is smaller than the Hausdorff distance between $\cR(p)$ and the corresponding projection. This motivates us to apply an increasing and concave transformation before the ridge estimation. In specific, we show that the power transformations $f^{q}(y)=y^q/q,-\infty<q\leq 1$ are increasing and concave on $\RR_+$, and thus we can use such power transformations when $p$ is strictly positive. Numerical experiments demonstrate the advantages of the proposed methods.

cs.LG

Manifold Fitting under Unbounded Noise

There has been an emerging trend in non-Euclidean statistical analysis of aiming to recover a low dimensional structure, namely a manifold, underlying the high dimensional data. Recovering the manifold requires the noise to be of certain concentration. Existing methods address this problem by constructing an approximated manifold based on the tangent space estimation at each sample point. Although theoretical convergence for these methods is guaranteed, either the samples are noiseless or the noise is bounded. However, if the noise is unbounded, which is a common scenario, the tangent space estimation at the noisy samples will be blurred. Fitting a manifold from the blurred tangent space might increase the inaccuracy. In this paper, we introduce a new manifold-fitting method, by which the output manifold is constructed by directly estimating the tangent spaces at the projected points on the underlying manifold, rather than at the sample points, to decrease the error caused by the noise. Assuming the noise is unbounded, our new method provides theoretical convergence in high probability, in terms of the upper bound of the distance between the estimated and underlying manifold. The smoothness of the estimated manifold is also evaluated by bounding the supremum of twice difference above. Numerical simulations are provided to validate our theoretical findings and demonstrate the advantages of our method over other relevant manifold fitting methods. Finally, our method is applied to real data examples.

stat.ML

Principal Sub-manifolds

We propose a novel method of finding principal components in multivariate data sets that lie on an embedded nonlinear Riemannian manifold within a higher-dimensional space. Our aim is to extend the geometric interpretation of PCA, while being able to capture non-geodesic modes of variation in the data. We introduce the concept of a principal sub-manifold, a manifold passing through a reference point, and at any point on the manifold extending in the direction of highest variation in the space spanned by the eigenvectors of the local tangent space PCA. Compared to recent work for the case where the sub-manifold is of dimension one Panaretos et al. (2014)$-$essentially a curve lying on the manifold attempting to capture one-dimensional variation$-$the current setting is much more general. The principal sub-manifold is therefore an extension of the principal flow, accommodating to capture higher dimensional variation in the data. We show the principal sub-manifold yields the ball spanned by the usual principal components in Euclidean space. By means of examples, we illustrate how to find, use and interpret a principal sub-manifold and we present an application in shape analysis.

stat.ME

Manifold Fitting

While classical data analysis has addressed observations that are real numbers or elements of a real vector space, at present many statistical problems of high interest in the sciences address the analysis of data that consist of more complex objects, taking values in spaces that are naturally not (Euclidean) vector spaces but which still feature some geometric structure. Manifold fitting is a long-standing problem, and has finally been addressed in recent years by Fefferman et al. (2020, 2021a). We develop a method with a theory guarantee that fits a $d$-dimensional underlying manifold from noisy observations sampled in the ambient space $\mathbb{R}^D$. The new approach uses geometric structures to obtain the manifold estimator in the form of image sets via a two-step mapping approach. We prove that, under certain mild assumptions and with a sample size $N=\mathcal{O}(σ^{(-d+3)})$, these estimators are true $d$-dimensional smooth manifolds whose estimation error, as measured by the Hausdorff distance, is bounded by $\mathcal{O}(σ^2\log(1/σ))$ with high probability. Compared with the existing approaches proposed in Fefferman et al. (2018, 2021b); Genovese et al. (2014); Yao and Xia (2019), our method exhibits superior efficiency while attaining very low error rates with a significantly reduced sample size, which scales polynomially in $σ^{-1}$ and exponentially in $d$. Extensive simulations are performed to validate our theoretical results. Our findings are relevant to various fields involving high-dimensional data in machine learning. Furthermore, our method opens up new avenues for existing non-Euclidean statistical methods in the sense that it has the potential to unify them to analyze data on manifolds in the ambience space domain.

math.ST

Random Fixed Boundary Flows

We consider fixed boundary flow with canonical interpretability as principal components extended on non-linear Riemannian manifolds. We aim to find a flow with fixed starting and ending points for noisy multivariate data sets lying on an embedded non-linear Riemannian manifold. In geometric term, the fixed boundary flow is defined as an optimal curve that moves in the data cloud with two fixed end points. At any point on the flow, we maximize the inner product of the vector field, which is calculated locally, and the tangent vector of the flow. The rigorous definition derives from an optimization problem using the intrinsic metric on the manifolds. For random data sets, we name the fixed boundary flow the random fixed boundary flow and analyze its limiting behavior under noisy observed samples. We construct a high level algorithm to compute the random fixed boundary flow and the convergence of the algorithm is provided. We show that the fixed boundary flow yields a concatenate of three segments, of which one coincides with the usual principal flow when the manifold is reduced to the Euclidean space. We further prove that the random fixed boundary flow converges largely to the population fixed boundary flow with high probability. We illustrate how the random fixed boundary flow can be used and interpreted, and showcase its application in real data sets.

math.OC

High Dimensional Quadratic Discriminant Analysis: Optimality and Phase Transitions

Consider a two-class classification problem where we observe samples $(X_i, Y_i)$ for i = 1, ..., n, $X_i \in R^p$ and $Y_i$ in {0, 1}. Given $Y_i = k$, $X_i$ is assumed to follow a multivariate normal distribution with mean $μ_k \in R^k$ and covariance matrix $Σ_k$, k=0,1. Supposing a new sample X from the same mixture is observed, our goal is to estimate its class label Y. Such a high-dimensional classification problem has been studied thoroughly when Sigma_0 = Sigma_1. However, the discussions over the case $Σ_0 \neq Σ_1$ are much less over the years. This paper presents the quadratic discriminant analysis (QDA) for the weak signals (QDAw) algorithm, and the QDA with feature selection (QDAfs) algorithm. QDAfs applies Partial Correlation Screening to estimate $\hatΩ_0$ and $\hatΩ_1$, and then applies a hard-thresholding on the diagonals of $\hatΩ_0 - \hatΩ_1$. QDAfs further includes the linear term $d^T X$, where d is achieved by a hard-thresholding on $\hatΩ_1\hatμ_1 - \hatΩ_0\hatμ_0$. We further propose the rare and weak model to model the signals in $Ω_0 - Ω_1$ and $μ_0 - μ_1$. Based on the signal weakness and sparsity in $μ_0 - μ_1$, we propose two ways to estimate labels: 1) QDAw for weak but dense signals; 2) QDAfs for relatively strong but sparse signals. We figure out the classification boundary on the 4-dim parameter space: 1) Region of possibility, where either QDAw or QDAfs will achieve a mis-classification error rate of 0; 2) Region of impossibility, where all classifiers will have a constant error rate. The numerical results from real datasets support our theories and demonstrate the necessity and superiority of using QDA over LDA for classification.

math.ST

A Statistical Approach to Estimating Adsorption-Isotherm Parameters in Gradient-Elution Preparative Liquid Chromatography

Determining the adsorption isotherms is an issue of significant importance in preparative chromatography. A modern technique for estimating adsorption isotherms is to solve an inverse problem so that the simulated batch separation coincides with actual experimental results. However, due to the ill-posedness, the high non-linearity, and the uncertainty quantification of the corresponding physical model, the existing deterministic inversion methods are usually inefficient in real-world applications. To overcome these difficulties and study the uncertainties of the adsorption-isotherm parameters, in this work, based on the Bayesian sampling framework, we propose a statistical approach for estimating the adsorption isotherms in various chromatography systems. Two modified Markov chain Monte Carlo algorithms are developed for a numerical realization of our statistical approach. Numerical experiments with both synthetic and real data are conducted and described to show the efficiency of the proposed new method.

stat.AP

Manifold Fitting in Ambient Space

Modern sample points in many applications no longer comprise real vectors in a real vector space but sample points of much more complex structures, which may be represented as points in a space with a certain underlying geometric structure, namely a manifold. Manifold learning is an emerging field for learning the underlying structure. The study of manifold learning can be split into two main branches: dimension reduction and manifold fitting. With the aim of combining statistics and geometry, we address the problem of manifold fitting in the ambient space. Inspired by the relation between the eigenvalues of the Laplace-Beltrami operator and the geometry of a manifold, we aim to find a small set of points that preserve the geometry of the underlying manifold. From this relationship, we extend the idea of subsampling to sample points in high-dimensional space and employ the Moving Least Squares (MLS) approach to approximate the underlying manifold. We analyze the two core steps in our proposed method theoretically and also provide the bounds for the MLS approach. Our simulation results and theoretical analysis demonstrate the superiority of our method in estimating the underlying manifold.

stat.ML