SearcharxivSearch

arXiv subjects

Xiaomeng Yan

Publications and source records attributed to Xiaomeng Yan.

4 recordsLinked to original sources

The M33 Synoptic Stellar Survey. III. Miras and LPVs in griJHKs

We present the results of a search for Miras and long-period variables (LPVs) in M33 using griJHKs archival observations from the Canada-France-Hawai'i Telescope. We use multiband information and machine learning techniques to identify and characterize these variables. We recover ~1,300 previously-discovered Mira candidates and identify ~13,000 new Miras and LPVs. We detect for the first time a clear first-overtone pulsation sequence among Mira candidates in this galaxy. We use O-rich, fundamental-mode Miras in the LMC and M33 to derive a distance modulus for the latter of 24.629 +/- 0.046 mag.

astro-ph.GA

Robust joint modeling of sparsely observed paired functional data

A reduced-rank mixed effects model is developed for robust modeling of sparsely observed paired functional data. In this model, the curves for each functional variable are summarized using a few functional principal components, and the association of the two functional variables is modeled through the association of the principal component scores. Multivariate scale mixture of normal distributions is used to model the principal component scores and the measurement errors in order to handle outlying observations and achieve robust inference. The mean functions and principal component functions are modeled using splines and roughness penalties are applied to avoid overfitting. An EM algorithm is developed for computation of model fitting and prediction. A simulation study shows that the proposed method outperforms an existing method which is not designed for robust estimation. The effectiveness of the proposed method is illustrated in an application of fitting multi-band light curves of Type Ia supernovae.

stat.ME

Optimal subsampling for large scale Elastic-net regression

Datasets with sheer volume have been generated from fields including computer vision, medical imageology, and astronomy whose large-scale and high-dimensional properties hamper the implementation of classical statistical models. To tackle the computational challenges, one of the efficient approaches is subsampling which draws subsamples from the original large datasets according to a carefully-design task-specific probability distribution to form an informative sketch. The computation cost is reduced by applying the original algorithm to the substantially smaller sketch. Previous studies associated with subsampling focused on non-regularized regression from the computational efficiency and theoretical guarantee perspectives, such as ordinary least square regression and logistic regression. In this article, we introduce a randomized algorithm under the subsampling scheme for the Elastic-net regression which gives novel insights into L1-norm regularized regression problem. To effectively conduct consistency analysis, a smooth approximation technique based on alpha absolute function is firstly employed and theoretically verified. The concentration bounds and asymptotic normality for the proposed randomized algorithm are then established under mild conditions. Moreover, an optimal subsampling probability is constructed according to A-optimality. The effectiveness of the proposed algorithm is demonstrated upon synthetic and real data datasets.

math.ST

Functional Principal Subspace Sampling for Large Scale Functional Data Analysis

Functional data analysis (FDA) methods have computational and theoretical appeals for some high dimensional data, but lack the scalability to modern large sample datasets. To tackle the challenge, we develop randomized algorithms for two important FDA methods: functional principal component analysis (FPCA) and functional linear regression (FLR) with scalar response. The two methods are connected as they both rely on the accurate estimation of functional principal subspace. The proposed algorithms draw subsamples from the large dataset at hand and apply FPCA or FLR over the subsamples to reduce the computational cost. To effectively preserve subspace information in the subsamples, we propose a functional principal subspace sampling probability, which removes the eigenvalue scale effect inside the functional principal subspace and properly weights the residual. Based on the operator perturbation analysis, we show the proposed probability has precise control over the first order error of the subspace projection operator and can be interpreted as an importance sampling for functional subspace estimation. Moreover, concentration bounds for the proposed algorithms are established to reflect the low intrinsic dimension nature of functional data in an infinite dimensional space. The effectiveness of the proposed algorithms is demonstrated upon synthetic and real datasets.

stat.CO