SearcharxivSearch

arXiv subjects

Yuguan Wang

Publications and source records attributed to Yuguan Wang.

6 recordsLinked to original sources

Permutation Recovery on Manifold Data via Spectral Seriation

Data points in many scientific experiments originate from an ordered structure, yet this ordering is often unavailable.We consider noisy data points with the correct ordering to be recovered. The underlying structure naturally places the data on a 1-dimensional manifold. Because eigenfunctions of 1-dimensional manifold Laplacian are trigonometric functions, and the manifold Laplacian can be approximated by the graph data Laplacian, the data ordering can be recovered by inverting the data Laplacian eigenvectors.We propose two spectral algorithms, one for the periodic structure (closed loop) and one for the non-periodic structure (open curve). We have derived the uniform error bound for the algorithms, which is composed of two parts: the discretization error between the manifold eigenfunctions and the noiseless graph Laplacian eigenvectors, and the eigenvectors error caused by data noise. In numerical studies, our spectral seriation algorithms outperform other manifold learning methods. The superior performance of our algorithms is demonstrated further on a biomolecule data example.

stat.ME

Fast operator learning for mapping correlations

We propose a fast, optimization-free method for learning the transition operators of high-dimensional Markov processes. The central idea is to perform a Galerkin projection of the transition operator to a suitable set of low-order bases that capture the correlations between the dimensions. Such a discretized operator can be obtained from moments corresponding to our choice of basis without curse of dimensionality. Furthermore, by exploiting its low-rank structure and the spatial decay of correlations, we can obtain a compressed representation with computational complexity of order $\mathcal{O}(dN)$, where $d$ is the dimensionality and $N$ is the sample size. We further theoretically analyze the approximation error of the proposed compressed representation. We numerically demonstrate that the learned operator allows efficient prediction of future events and solving high-dimensional boundary value problems. This gives rise to a simple linear algebraic method for high-dimensional rare-events simulations.

math.NA

Subspace method of moments for ab initio 3-D single-particle cryo-EM reconstruction

Cryo-electron microscopy (cryo-EM) is a widely used technique for recovering the 3-D structure of biological molecules from a large number of experimentally generated noisy 2-D tomographic projection images of the 3-D structure, taken from unknown viewing angles. Through computationally intensive algorithms, these observed images are processed to reconstruct the 3-D structures. Many popular computational methods rely on estimating the unknown angles as part of the reconstruction process, which becomes particularly challenging at low signal-to-noise ratios. The method of moments (MoM) offers an alternative approach that circumvents the estimation of viewing orientations of individual projection images by instead estimating the underlying distribution of the viewing angles, and is robust to noise given sufficiently many images. However, the method of moments typically entails computing higher-order moments of the projection images, incurring significant computational and memory costs. To mitigate this, we propose a new approach called the subspace method of moments (SubspaceMoM), which compresses the first three moments using data-driven low-rank tensor techniques as well as expansion into a suitable function basis. The compressed moments can be efficiently computed from the set of projection images using numerical quadrature and can be employed to jointly reconstruct the 3-D structure and the distribution of viewing orientations. We illustrate the practical applicability of SubspaceMoM through numerical experiments using up to the third-order moment on synthetic datasets with a simplified cryo-EM image formation model, which significantly improves the reconstruction resolution compared to previous MoM approaches.

math.NA

Fast Multipole Method with Complex Coordinates

In this work we present a variant of the fast multipole method (FMM) for efficiently evaluating standard layer potentials on geometries with complex coordinates in two and three dimensions. The complex scaled boundary integral method for the efficient solution of scattering problems on unbounded domains results in complex point locations upon discretization. Classical real-coordinate FMMs are no longer applicable, hindering the use of this approach for large-scale problems. Here we develop the complex-coordinate FMM based on the analytic continuation of certain special function identities used in the construction of the classical FMM. To achieve the same linear time complexity as the classical FMM, we construct a hierarchical tree based solely on the real parts of the complex point locations, and derive convergence rates for truncated expansions when the imaginary parts of the locations are a Lipschitz function of the corresponding real parts. We demonstrate the efficiency of our approach through several numerical examples and illustrate its application for solving large-scale time-harmonic water wave problems and Helmholtz transmission problems.

math.NA

Type-I Superconductors in the Limit as the London Penetration Depth Goes to 0

This paper provides an explicit formula for the approximate solution of the static London equations. These equations describe the currents and magnetic fields in a Type-I superconductor. We represent the magnetic field as a 2-form and the current as a 1-form, and assume that the superconducting material is contained in a bounded, connected set, $Ω,$ with smooth boundary. The London penetration depth gives an estimate for the thickness of the layer near $\partialΩ$ where the current is largely carried. In an earlier paper, we introduced a system of Fredholm integral equations of second kind, on $\partialΩ,$ for solving the physically relevant scattering problems in this context. In real Type-I superconductors the penetration depth is very small, typically about $100$nm, which often renders the integral equation approach computationally intractable. In this paper we provide an explicit formula for approximate solutions, with essentially optimal error estimates, as the penetration depth tends to zero. Our work makes extensive use of the Hodge decomposition of differential forms on manifolds with boundary, and thus evokes Kohn's work on the tangential Cauchy-Riemann equations.

math.AP

Uniform error bound for PCA matrix denoising

Principal component analysis (PCA) is a simple and popular tool for processing high-dimensional data. We investigate its effectiveness for matrix denoising. We consider the clean data are generated from a low-dimensional subspace, but masked by independent high-dimensional sub-Gaussian noises with standard deviation $σ$. Under the low-rank assumption on the clean data with a mild spectral gap assumption, we prove that the distance between each pair of PCA-denoised data point and the clean data point is uniformly bounded by $O(σ\log n)$. To illustrate the spectral gap assumption, we show it can be satisfied when the clean data are independently generated with a non-degenerate covariance matrix. We then provide a general lower bound for the error of the denoised data matrix, which indicates PCA denoising gives a uniform error bound that is rate-optimal. Furthermore, we examine how the error bound impacts downstream applications such as clustering and manifold learning. Numerical results validate our theoretical findings and reveal the importance of the uniform error.

math.ST