SearcharxivSearch

arXiv subjects

Ruoxu Tan

Publications and source records attributed to Ruoxu Tan.

5 recordsLinked to original sources

High-dimensional Semi-supervised Classification via the Fermat Distance

Semi-supervised classification, where unlabeled data are massive but labeled data are limited, often arises in machine learning applications. We address this challenge under high-dimensional data by leveraging the manifold and cluster assumptions. Based on the Fermat distance, a density-sensitive metric that naturally encodes the cluster assumption, we propose the weighted $k$-nearest neighbors (NN) classifier and multidimensional scaling (MDS)-induced classifiers. The use of MDS with a large target dimension allows the effective application of linear classifiers to complex manifold data. Theoretically, we derive a sharp lower bound for the expected excess risk within clusters and prove that the weighted $k$-NN classifier utilizing the true Fermat distance is minimax optimal. Furthermore, we explicitly quantify the utility of unlabeled data by showing that the error arising from estimating the Fermat distance decays exponentially with the pooled sample size. Such a rate is much faster than the related rates in the literature. Extensive experiments on synthetic and real datasets demonstrate competitive or superior performance of our approaches compared to state-of-the-art graph-based semi-supervised classifiers.

stat.ML

Semi-supervised Classification for Noisy Functional Data with Application to Astronomical Spectra

Despite its extensive development for multivariate data, semi-supervised learning remains underdeveloped for functional data, especially under discrete and noisy observations. We develop a density-sensitive semi-supervised framework for functional data supported on a low-dimensional manifold by adapting the Fermat distance to reconstructed trajectories. The resulting pairwise distances are used to construct a weighted $k$-nearest-neighbor classifier and multidimensional-scaling-based classifiers. To accommodate massive datasets commonly seen in semi-supervised applications, we design a computationally efficient estimation procedure tailored for discrete and noisy functional observations. Theoretically, we establish exponentially decaying convergence rates of the $k$-NN classifier and the consistency of the estimated Fermat distance. Crucially, our results reveal that incorporating unlabeled data may not lead to improved classification accuracy without a sufficiently fast-growing individual sampling rate, precisely due to discrete and noisy observations. In most simulation settings satisfying the manifold and cluster assumptions, the proposed classifiers outperform the supervised benchmarks considered; in the Gaia spectra analysis, they attain higher agreement with high-confidence proxy labels.

stat.ME

Supervised Manifold Learning for Functional Data

Classification is a core topic in functional data analysis. A large number of functional classifiers have been proposed in the literature, most of which are based on functional principal component analysis or functional regression. In contrast, we investigate this topic from the perspective of manifold learning. It is assumed that functional data lie on an unknown low-dimensional manifold, and we expect that superior classifiers can be developed based on the manifold structure. To this end, we propose a novel proximity measure that takes the label information into account to learn the low-dimensional representations, also known as the supervised manifold learning outcomes. When the outcomes are coupled with multivariate classifiers, the procedure induces a new family of functional classifiers. In theory, we prove that our functional classifier induced by the $k$-NN classifier is asymptotically optimal. In practice, we show that our method, coupled with several classical multivariate classifiers, achieves highly competitive classification performance compared to existing functional classifiers across both synthetic and real data examples. Supplementary materials are available online.

stat.ME

Delaunay Weighted Two-sample Test for High-dimensional Data by Incorporating Geometric Information

Two-sample hypothesis testing is a fundamental problem with various applications, which faces new challenges in the high-dimensional context. To mitigate the issue of the curse of dimensionality, high-dimensional data are typically assumed to lie on a low-dimensional manifold. To incorporate geometric information in the data, we propose to apply the Delaunay triangulation and develop the Delaunay weight to measure the geometric proximity among data points. In contrast to existing similarity measures that only utilize pairwise distances, the Delaunay weight can take both the distance and direction information into account. A detailed computation procedure is developed to learn the unknown manifold and approximate the Delaunay weight. We further propose a novel nonparametric test statistic using the Delaunay weight matrix. Asymptotic normality under the null and consistency under the alternative of the test statistic are developed. Applied on simulated data, the new test shows robustness to the learning of the unknown manifold and exhibits substantial power gain if the distributions differ directions. The proposed test also shows great power on a real dataset of mice protein expression levels.

stat.ME

Causal Effect of Functional Treatment

We study the causal effect with a functional treatment variable, where practical applications often arise in neuroscience, biomedical sciences, etc. Previous research concerning the effect of a functional variable on an outcome is typically restricted to exploring correlation rather than causality. The generalized propensity score, which is often used to calibrate the selection bias, is not directly applicable to a functional treatment variable due to a lack of definition of probability density function for functional data. We propose three estimators for the average dose-response functional based on the functional linear model, namely, the functional stabilized weight estimator, the outcome regression estimator and the doubly robust estimator, each of which has its own merits. We study their theoretical properties, which are corroborated through extensive numerical experiments. A real data application on electroencephalography data and disease severity demonstrates the practical value of our methods.

stat.ME