Searcharxiv⌕ Search

arXiv subjects

Weichen Wang

Publications and source records attributed to Weichen Wang.

63 records · Page 4Linked to original sources

UVI colour gradients of 0.4<z<1.4 star-forming main sequence galaxies in CANDELS: dust extinction and star formation profiles

This paper uses radial colour profiles to infer the distributions of dust, gas and star formation in z=0.4-1.4 star-forming main sequence galaxies. We start with the standard UVJ-based method to estimate dust extinction and specific star formation rate (sSFR). By replacing J with I band, a new calibration method suitable for use with ACS+WFC3 data is created (i.e. UVI diagram). Using a multi-wavelength multi-aperture photometry catalogue based on CANDELS, UVI colour profiles of 1328 galaxies are stacked in stellar mass and redshift bins. The resulting colour gradients, covering a radial range of 0.2--2.0 effective radii, increase strongly with galaxy mass and with global $A_V$. Colour gradient directions are nearly parallel to the Calzetti extinction vector, indicating that dust plays a more important role than stellar population variations. With our calibration, the resulting $A_V$ profiles fall much more slowly than stellar mass profiles over the measured radial range. sSFR gradients are nearly flat without central quenching signatures, except for $M_*>10^{10.5} M_{\odot}$, where central declines of 20--25 per cent are observed. Both sets of profiles agree well with previous radial sSFR and (continuum) $A_V$ measurements. They are also consistent with the sSFR profiles and, if assuming a radially constant gas-to-dust ratio, gas profiles in recent hydrodynamic models. We finally discuss the striking findings that SFR scales with stellar mass density in the inner parts of galaxies, and that dust content is high in the outer parts despite low stellar-mass surface densities there.

astro-ph.GA↗

A Shrinkage Principle for Heavy-Tailed Data: High-Dimensional Robust Low-Rank Matrix Recovery

This paper introduces a simple principle for robust high-dimensional statistical inference via an appropriate shrinkage on the data. This widens the scope of high-dimensional techniques, reducing the moment conditions from sub-exponential or sub-Gaussian distributions to merely bounded second or fourth moment. As an illustration of this principle, we focus on robust estimation of the low-rank matrix $Θ^*$ from the trace regression model $Y=Tr (Θ^{*T}X) +ε$. It encompasses four popular problems: sparse linear models, compressed sensing, matrix completion and multi-task regression. We propose to apply penalized least-squares approach to appropriately truncated or shrunk data. Under only bounded $2+δ$ moment condition on the response, the proposed robust methodology yields an estimator that possesses the same statistical error rates as previous literature with sub-Gaussian errors. For sparse linear models and multi-tasking regression, we further allow the design to have only bounded fourth moment and obtain the same statistical rates, again, by appropriate shrinkage of the design matrix. As a byproduct, we give a robust covariance matrix estimator and establish its concentration inequality in terms of the spectral norm when the random samples have only bounded fourth moment. Extensive simulations have been carried out to support our theories.

math.ST↗

The UV-optical Color Gradients in Star-Forming Galaxies at 0.5<z<1.5: Origins and Link to Galaxy Assembly

The rest-frame UV-optical (i.e., NUV-B) color index is sensitive to the low-level recent star formation and dust extinction, but it is insensitive to the metallicity. In this Letter, we have measured the rest-frame NUV-B color gradients in ~1400 large ($\rm r_e>0.18^{\prime\prime}$), nearly face-on (b/a>0.5) main-sequence star-forming galaxies (SFGs) between redshift 0.5 and 1.5 in the CANDELS/GOODS-S and UDS fields. With this sample, we study the origin of UV-optical color gradients in the SFGs at z~1 and discuss their link with the buildup of stellar mass. We find that the more massive, centrally compact, and more dust extinguished SFGs tend to have statistically more negative raw color gradients (redder centers) than the less massive, centrally diffuse, and less dusty SFGs. After correcting for dust reddening based on optical-SED fitting, the color gradients in the low-mass ($M_{\ast} <10^{10}M_{\odot}$) SFGs generally become quite flat, while most of the high-mass ($M_{\ast} > 10^{10.5}M_{\odot}$) SFGs still retain shallow negative color gradients. These findings imply that dust reddening is likely the principal cause of negative color gradients in the low-mass SFGs, while both increased central dust reddening and buildup of compact old bulges are likely the origins of negative color gradients in the high-mass SFGs. These findings also imply that at these redshifts the low-mass SFGs buildup their stellar masses in a self-similar way, while the high-mass SFGs grow inside out.

astro-ph.GA↗

Heterogeneity Adjustment with Applications to Graphical Model Inference

Heterogeneity is an unwanted variation when analyzing aggregated datasets from multiple sources. Though different methods have been proposed for heterogeneity adjustment, no systematic theory exists to justify these methods. In this work, we propose a generic framework named ALPHA (short for Adaptive Low-rank Principal Heterogeneity Adjustment) to model, estimate, and adjust heterogeneity from the original data. Once the heterogeneity is adjusted, we are able to remove the biases of batch effects and to enhance the inferential power by aggregating the homogeneous residuals from multiple sources. Under a pervasive assumption that the latent heterogeneity factors simultaneously affect a large fraction of observed variables, we provide a rigorous theory to justify the proposed framework. Our framework also allows the incorporation of informative covariates and appeals to the "Bless of Dimensionality". As an illustrative application of this generic framework, we consider a problem of estimating high-dimensional precision matrix for graphical model inference based on multiple datasets. We also provide thorough numerical studies on both synthetic datasets and a brain imaging dataset to demonstrate the efficacy of the developed theory and methods.

stat.ME↗

Robust Covariance Estimation for Approximate Factor Models

In this paper, we study robust covariance estimation under the approximate factor model with observed factors. We propose a novel framework to first estimate the initial joint covariance matrix of the observed data and the factors, and then use it to recover the covariance matrix of the observed data. We prove that once the initial matrix estimator is good enough to maintain the element-wise optimal rate, the whole procedure will generate an estimated covariance with desired properties. For data with only bounded fourth moments, we propose to use Huber loss minimization to give the initial joint covariance estimation. This approach is applicable to a much wider range of distributions, including sub-Gaussian and elliptical distributions. We also present an asymptotic result for Huber's M-estimator with a diverging parameter. The conclusions are demonstrated by extensive simulations and real data analysis.

stat.ME↗

Projected principal component analysis in factor models

This paper introduces a Projected Principal Component Analysis (Projected-PCA), which employs principal component analysis to the projected (smoothed) data matrix onto a given linear space spanned by covariates. When it applies to high-dimensional factor analysis, the projection removes noise components. We show that the unobserved latent factors can be more accurately estimated than the conventional PCA if the projection is genuine, or more precisely, when the factor loading matrices are related to the projected linear space. When the dimensionality is large, the factors can be estimated accurately even when the sample size is finite. We propose a flexible semiparametric factor model, which decomposes the factor loading matrix into the component that can be explained by subject-specific covariates and the orthogonal residual component. The covariates' effects on the factor loadings are further modeled by the additive model via sieve approximations. By using the newly proposed Projected-PCA, the rates of convergence of the smooth factor loading matrices are obtained, which are much faster than those of the conventional factor analysis. The convergence is achieved even when the sample size is finite and is particularly appealing in the high-dimension-low-sample-size situation. This leads us to developing nonparametric tests on whether observed covariates have explaining powers on the loadings and whether they fully explain the loadings. The proposed method is illustrated by both simulated data and the returns of the components of the S&P 500 index.

stat.ME↗

Estimation of functionals of sparse covariance matrices

High-dimensional statistical tests often ignore correlations to gain simplicity and stability leading to null distributions that depend on functionals of correlation matrices such as their Frobenius norm and other $\ell_r$ norms. Motivated by the computation of critical values of such tests, we investigate the difficulty of estimation the functionals of sparse correlation matrices. Specifically, we show that simple plug-in procedures based on thresholded estimators of correlation matrices are sparsity-adaptive and minimax optimal over a large class of correlation matrices. Akin to previous results on functional estimation, the minimax rates exhibit an elbow phenomenon. Our results are further illustrated in simulated data as well as an empirical study of data arising in financial econometrics.

math.ST↗

Asymptotics of Empirical Eigen-structure for Ultra-high Dimensional Spiked Covariance Model

We derive the asymptotic distributions of the spiked eigenvalues and eigenvectors under a generalized and unified asymptotic regime, which takes into account the spike magnitude of leading eigenvalues, sample size, and dimensionality. This new regime allows high dimensionality and diverging eigenvalue spikes and provides new insights into the roles the leading eigenvalues, sample size, and dimensionality play in principal component analysis. The results are proven by a technical device, which swaps the role of rows and columns and converts the high-dimensional problems into low-dimensional ones. Our results are a natural extension of those in Paul (2007) to more general setting with new insights and solve the rates of convergence problems in Shen et al. (2013). They also reveal the biases of the estimation of leading eigenvalues and eigenvectors by using principal component analysis, and lead to a new covariance estimator for the approximate factor model, called shrinkage principal orthogonal complement thresholding (S-POET), that corrects the biases. Our results are successfully applied to outstanding problems in estimation of risks of large portfolios and false discovery proportions for dependent test statistics and are illustrated by simulation studies.

math.ST↗

Large Covariance Estimation through Elliptical Factor Models

We proposed a general Principal Orthogonal complEment Thresholding (POET) framework for large-scale covariance matrix estimation based on an approximate factor model. A set of high level sufficient conditions for the procedure to achieve optimal rates of convergence under different matrix norms were brought up to better understand how POET works. Such a framework allows us to recover the results for sub-Gaussian in a more transparent way that only depends on the concentration properties of the sample covariance matrix. As a new theoretical contribution, for the first time, such a framework allows us to exploit conditional sparsity covariance structure for the heavy-tailed data. In particular, for the elliptical data, we proposed a robust estimator based on marginal and multivariate Kendall's tau to satisfy these conditions. In addition, conditional graphical model was also studied under the same framework. The technical tools developed in this paper are of general interest to high dimensional principal component analysis. Thorough numerical results were also provided to back up the developed theory.

stat.ME↗