SearcharxivSearch

arXiv subjects

Hai-Ling Lu

Publications and source records attributed to Hai-Ling Lu.

6 recordsLinked to original sources

Spectra as Language: Large Language Models for Scalable Stellar Parameter and Abundance Inference

Stellar spectra encode key information on the physical properties and chemical compositions of stars. Accurate stellar parameter determination is essential for addressing major questions such as galaxy and stellar evolution. Large-scale spectroscopic surveys have accumulated unprecedented spectral data. Traditional feature extraction or model-fitting approaches struggle with high-dimensional, massive datasets, limited generalization, and computational inefficiency. Recent advances in large language models demonstrate strong generalization and feature-learning in tasks like natural language processing, DNA/RNA sequence analysis, and protein/chemical parsing. Stellar spectra are continuous sequential signals, enabling the transfer of language models to stellar spectroscopy. Here, we propose a two-stage large language model framework for stellar parameter inference, achieving accurate estimation of effective temperature, surface gravity, metallicity, and abundances of ~20 chemical elements. Scaling-law analyses show systematic performance improvements with increasing data, providing a scalable framework for forthcoming large-scale surveys.

astro-ph.IM

PISP: Projected-Space Inference of Stellar Parameters

To improve the accuracy and efficiency of high-dimensional stellar parameter inference in large spectroscopic datasets, we propose a projection-assisted parameter-inference framework -- Projected-Space Inference of Stellar Parameters (PISP). PISP constructs an orthonormal basis and optimizes in the projected space, reducing the impact of parameter correlations on inference. The basis is constructed using either principal component analysis (PCA) or the active-subspace (AS) method and is combined with two inference strategies -- Non-L1, which optimizes the projection coefficients for a user-specified projected dimensionality, and L1, which introduces L1 regularization in the full projected space to adaptively select projection directions -- yielding four strategies: PCA-Non-L1, AS-Non-L1, PCA-L1, and AS-L1. For different computational scenarios, we implement two versions: PISP-CurveFit for fast single-spectrum inference and PISP-Adam for large-scale GPU-parallel inference. Using a fully connected neural network and a residual network as spectral emulators, we evaluate PISP on Kurucz synthetic spectra and on $722{,}896$ APOGEE DR$17$ observed spectra. Compared to the baseline strategy, PISP improves inference accuracy for multiple parameters across all emulator-optimizer combinations. In synthetic data, PCA-L1 performs best, reducing the standard deviation of differences ($σ(Δ)$) by at least $0.01$ dex for $12$ of $20$ elemental abundances, with [N/H], [O/H], [Na/H], [Co/H], [P/H], [V/H], [Cu/H] showing $0.05$--$0.72$ dex reductions. In observed data, PCA-Non-L1 reduces $σ(Δ)$ by $>30$ K for effective temperature and by at least $0.01$ dex for $9$ of $17$ elemental abundances, with [O/H], [Na/H], [V/H] showing $0.05$--$0.20$ dex reductions, while achieving a $\sim$$4\times$ efficiency gain, slightly outperforming PCA-L1.

astro-ph.SR

Scalable Stellar Parameter Inference Using Python-based LASP: From CPU Optimization to GPU Acceleration

To enhance the efficiency, scalability, and cross-survey applicability of stellar parameter inference in large spectroscopic datasets, we present a modular, parallelized Python framework with automated error estimation, built on the LAMOST Atmospheric Parameter Pipeline (LASP) originally implemented in IDL. Rather than a direct code translation, this framework refactors LASP with two complementary modules: LASP-CurveFit, a new implementation of the LASP fitting procedure that runs on a CPU, preserving legacy logic while improving data I/O and multithreaded execution efficiency; and LASP-Adam-GPU, a GPU-accelerated method that introduces grouped optimization by constructing a joint residual function over multiple observed and model spectra, enabling high-throughput parameter inference across tens of millions of spectra. Applied to 10 million LAMOST spectra, the framework reduces runtime from 84 to 48 hr on the same CPU platform and to 7 hr on an NVIDIA A100 GPU, while producing results consistent with those from the original pipeline. The inferred errors agree well with the parameter variations from repeat observations of the same target (excluding radial velocities), while the official empirical errors used in LASP are more conservative. When applied to DESI DR1, our effective temperatures and surface gravities agree better with APOGEE than those from the DESI pipeline, particularly for cool giants, while the latter performs slightly better in radial velocity and metallicity. These results suggest that the framework delivers reliable accuracy, efficiency, and transferability, offering a practical approach to parameter inference in large spectroscopic surveys. The code and DESI-based catalog are available via \dataset[DOI: 10.12149/101679]{https://doi.org/10.12149/101679} and \dataset[DOI: 10.12149/101675]{https://doi.org/10.12149/101675}, respectively.

astro-ph.GA

Estimating stellar atmospheric parameters and elemental abundances using fully connected residual network

Stellar atmospheric parameters and elemental abundances are traditionally determined using template matching techniques based on high-resolution spectra. However, these methods are sensitive to noise and unsuitable for ultra-low-resolution data. Given that the Chinese Space Station Telescope (CSST) will acquire large volumes of ultra-low-resolution spectra, developing effective methods for ultra-low-resolution spectral analysis is crucial. In this work, we investigated the Fully Connected Residual Network (FCResNet) for simultaneously estimating atmospheric parameters ($T_\text{eff}$, $\log g$, [Fe/H]) and elemental abundances ([C/Fe], [N/Fe], [Mg/Fe]). We trained and evaluated FCResNet using CSST-like spectra (\textit{R} $\sim$ 200) generated by degrading LAMOST spectra (\textit{R} $\sim$ 1,800), with reference labels from APOGEE. FCResNet significantly outperforms traditional machine learning methods (KNN, XGBoost, SVR) and CNN in prediction precision. For spectra with g-band signal-to-noise ratio greater than 20, FCResNet achieves precisions of 78 K, 0.15 dex, 0.08 dex, 0.05 dex, 0.10 dex, and 0.05 dex for $T_\text{eff}$, $\log g$, [Fe/H], [C/Fe], [N/Fe] and [Mg/Fe], respectively, on the test set. FCResNet processes one million spectra in only 42 seconds while maintaining a simple architecture with just 348 KB model size. These results suggest that FCResNet is a practical and promising tool for processing the large volume of ultra-low-resolution spectra that will be obtained by CSST in the future.

astro-ph.IM

Refined M-type Star Catalog from LAMOST DR10: Measurements of Radial Velocities, $T_\text{eff}$, log $g$, [M/H] and [$α$/M]

Precise stellar parameters for M-type stars, the Galaxy's most common stellar type, are crucial for numerous studies. In this work, we refined the LAMOST DR10 M-type star catalog through a two-stage process. First, we purified the catalog using techniques including deep learning and color-magnitude diagrams to remove 22,496 non-M spectra, correct 2,078 dwarf/giant classifications, and update 12,900 radial velocities. This resulted in a cleaner catalog containing 870,518 M-type spectra (820,493 dwarfs, 50,025 giants). Second, applying a label transfer strategy using values from APOGEE DR16 for parameter prediction with a ten-fold cross-validated CNN ensemble architecture, we predicted $T_\text{eff}$, $\log g$, [M/H], and [$α$/M] separately for M dwarfs and giants. The average internal errors for M dwarfs/giants are respectively: $T_\text{eff}$ 30/17 K, log $g$ 0.07/0.07 dex, [M/H] 0.07/0.05 dex, and [$α$/M] 0.02/0.02 dex. Comparison with APOGEE demonstrates external precisions of 34/14 K, 0.12/0.07 dex, 0.09/0.04 dex, and 0.03/0.02 dex for M dwarfs/giants, which represents precision improvements of over 20\% for M dwarfs and over 50\% for M giants compared to previous literature results. The catalog is available at https://nadc.china-vo.org/res/r101668/.

astro-ph.SR

Estimating Stellar Atmospheric Parameters and [α/Fe] for LAMOST O-M type Stars Using a Spectral Emulator

In this paper, we developed a spectral emulator based on the Mapping Nearby Galaxies at Apache Point Observatory Stellar Library (MaStar) and a grouping optimization strategy to estimate effective temperature (T_eff), surface gravity (log g), metallicity ([Fe/H]) and the abundance of alpha elements with respect to iron ([alpha/Fe]) for O-M-type stars within the Large Sky Area Multi-Object Fiber Spectroscopic Telescope (LAMOST) low-resolution spectra. The primary aim is to use a rapid spectral-fitting method, specifically the spectral emulator with the grouping optimization strategy, to create a comprehensive catalog for stars of all types within LAMOST, addressing the shortcomings in parameter estimations for both cold and hot stars present in the official LAMOST AFGKM-type catalog. This effort is part of our series of studies dedicated to establishing an empirical spectral library for LAMOST. Experimental results demonstrate that our method is effectively applicable to parameter prediction for LAMOST, with the single-machine processing time within $70$ hr. We observed that the internal error dispersions for T_eff, log g, [Fe/H], and [alpha/Fe] across different spectral types lie within the ranges of $15-594$ K, $0.03-0.27$ dex, $0.02-0.10$ dex, and $0.01-0.04$ dex, respectively, indicating a good consistency. A comparative analysis with external data highlighted deficiencies in the official LAMOST catalog and issues with MaStar parameters, as well as potential limitations of our method in processing spectra with strong emission lines and bad pixels. The derived atmospheric parameters as a part of this work are available at https://nadc.china-vo.org/res/r101402/ .

astro-ph.SR