Searcharxiv⌕ Search

arXiv subjects

Lu Yan

Publications and source records attributed to Lu Yan.

25 records · Page 2Linked to original sources

Distributed estimation of spiked eigenvalues in spiked population models

The proliferation of science and technology has led to the prevalence of voluminous data sets that are distributed across multiple machines. It is an established fact that conventional statistical methodologies may be unfeasible in the analysis of such massive data sets due to prohibitively long computing durations, memory constraints, communication overheads, and confidentiality considerations. In this paper, we propose distributed estimators of the spiked eigenvalues in spiked population models. The consistency and asymptotic normality of the distributed estimators are derived, and the statistical error analysis of the distributed estimators is provided as well. Compared to the estimation from the full sample, the proposed distributed estimation shares the same order of convergence. Simulation study and real data analysis indicate that the proposed distributed estimation and testing procedures have excellent properties in terms of estimation accuracy and stability as well as transmission efficiency.

math.ST↗

Generating functions of multiple $t$-star values

In this paper, we study the generating functions of multiple $t$-star values with an arbitrary number of blocks of twos, which are based on the results of the corresponding generating functions of multiple $t$-harmonic star sums. These generating functions can deduce an explicit expression of multiple $t$-star values. As applications, we obtain some evaluations of multiple $t$-star values with one-two-three or more general indices.

math.NT↗

Parametric Euler $T$-sums of odd harmonic numbers

In this paper, we define a parametric variant of generalized Euler sums and call them the (alternating) parametric Euler $T$-sums. By using the contour integration method and residue theorem, we establish several explicit formulae for the linear parametric Euler $T$-sums. Furthermore, by applying the results, we obtain explicit formulae for the Hoffman's (alternating) double $t$-values and Kaneko-Tsumura's (alternating) double $T$-values.

math.NT↗

CrossFix: Collaborative bug fixing by recommending similar bugs

Many automated program repair techniques have been proposed for fixing bugs. Some of these techniques use the information beyond the given buggy program and test suite to improve the quality of generated patches. However, there are several limitations that hinder the wide adoption of these techniques, including (1) they rely on a fixed set of repair templates for patch generation or reference implementation, (2) searching for the suitable reference implementation is challenging, (3) generated patches are not explainable. Meanwhile, a recent approach shows that similar bugs exist across different projects and one could use the GitHub issue from a different project for finding new bugs for a related project. We propose collaborative bug fixing, a novelapproach that suggests bug reports that describe a similar bug. Our studyredefines similar bugs as bugs that share the (1) same libraries, (2) same functionalities, (3) same reproduction steps, (4) same configurations, (5) sameoutcomes, or (6) same errors. Moreover, our study revealed the usefulness of similar bugs in helping developers in finding more context about the bug and fixing. Based on our study, we design CrossFix, a tool that automatically suggests relevant GitHub issues based on an open GitHub issue. Our evaluation on 249 open issues from Java and Android projects shows that CrossFix could suggest similar bugs to help developers in debugging and fixing.

cs.SE↗

MTFuzz: Fuzzing with a Multi-Task Neural Network

Fuzzing is a widely used technique for detecting software bugs and vulnerabilities. Most popular fuzzers generate new inputs using an evolutionary search to maximize code coverage. Essentially, these fuzzers start with a set of seed inputs, mutate them to generate new inputs, and identify the promising inputs using an evolutionary fitness function for further mutation. Despite their success, evolutionary fuzzers tend to get stuck in long sequences of unproductive mutations. In recent years, machine learning (ML) based mutation strategies have reported promising results. However, the existing ML-based fuzzers are limited by the lack of quality and diversity of the training data. As the input space of the target programs is high dimensional and sparse, it is prohibitively expensive to collect many diverse samples demonstrating successful and unsuccessful mutations to train the model. In this paper, we address these issues by using a Multi-Task Neural Network that can learn a compact embedding of the input space based on diverse training samples for multiple related tasks (i.e., predicting for different types of coverage). The compact embedding can guide the mutation process by focusing most of the mutations on the parts of the embedding where the gradient is high. \tool uncovers $11$ previously unseen bugs and achieves an average of $2\times$ more edge coverage compared with 5 state-of-the-art fuzzer on 10 real-world programs.

cs.SE↗

Analysis of Stellar Spectra from LAMOST DR5 with Generative Spectrum Networks

In this study, the fundamental stellar atmospheric parameters (Teff, log g, [Fe/H] and [α/Fe]) were derived for low-resolution spectroscopy from LAMOST DR5 with Generative Spectrum Networks (GSN). This follows the same scheme as a normal artificial neural network with stellar parameters as the input and spectra as the output. The GSN model was effective in producing synthetic spectra after training on the PHOENIX theoretical spectra. In combination with Bayes framework, the application for analysis of LAMOST observed spectra exhibited improved efficiency on the distributed computing platform, Spark. In addition, the results were examined and validated by a comparison with reference parameters from high-resolution surveys and asteroseismic results. Our results show good consistency with the results from other survey and catalogs. Our proposed method is reliable with a precision of 80 K for Teff, 0.14 dex for log g, 0.07 dex for [Fe/H] and 0.168 dex for [α/Fe], for spectra with a signal-to-noise in g bands (SNRg) higher than 50. The parameters estimated as a part of this work are available at http://paperdata.china-vo.org/GSN_parameters/GSN_parameters.csv.

astro-ph.IM↗