SearcharxivSearch

arXiv subjects

Ali Luo

Publications and source records attributed to Ali Luo.

At least 19 recordsLinked to original sources

AstroSpecLM: A Spectrum-Language Model for Evidence-Grounded Astronomical Spectral Analysis

Astronomical spectra encode rich physical information, but drawing scientific conclusions from spectral features typically requires expert interpretation. This paper presents AstroSpecLM, a spectrum-language model that connects one-dimensional DESI spectra with Qwen3-4B to answer questions and provide explanations grounded in spectral evidence. Instead of generating question-answer pairs directly from templates or raw catalog fields, we first distill each spectrum into a compact set of catalog- and spectrum-derived facts, then use these facts as references to generate instruction-following conversations. The resulting model is competitive with specialist supervised baselines on classification and redshift estimation, while additionally producing natural-language explanations that reference specific spectral features. Our results indicate that grounding a language model in one-dimensional scientific spectra is feasible, and that fact-mediated instruction data yields a model capable of both prediction and explanation.

astro-ph.IM

Semi-supervised Source Detection in Astronomical Images: New Benchmark and Strong Baseline

Source detection in modern observational astronomy is a cornerstone for localizing and identifying stellar sources accurately. It is crucial for studies such as stellar population synthesis and cosmological parameter estimation. However, the characteristics of astronomical images, including high density, the effect of point spread functions and low signal-to-noise ratios, significantly challenge the latest advanced object detectors. Besides, fully-supervised detection methods are hardly practical, due to the significant difficulty in annotating dense, small, and faint sources in astronomical images. To tackle the scarcity of astronomical datasets, we introduce a new comprehensive benchmark (LAMOST-DET), comprising 18,400 astronomical images and 728,898 source instances. Upon the dataset, we further devise a novel semi-supervised learning framework coined Nova Teacher, capable of detecting dense sources effectively given sparse annotations. It integrates source light enhancement module, confidence-guided pseudo-supervision, and cross-view complementary mining in a dual-teacher paradigm. Extensive experiments on LAMOST-DET show that, Nova Teacher consistently improves previous competitors by 4.04% and 5.22% mAP under two semi-supervised settings. Additionally, our method competes against other detectors on a natural image dataset, validating its generalization ability to various scenarios. The source code is available at https://github.com/AcWiz/NovaTeacher.

cs.CV

The Stellar Abundances and Galactic Evolution Survey (SAGES). V. The First Data Release of the DDO51 Band

We present the first public data release of DDO51 band from the Stellar Abundances and Galactic Evolution Survey (SAGES), based on Nanshan One-meter Wide-field Telescope (NOWT) observations obtained between 2023 September and 2024 January. This release initiates the DDO51-band component of the survey, covering $\sim$ 2,500 deg$^2$ of the northern sky and including more than 10 million sources. The DDO51 filter is centered near the \ion{Mg}{1}~$b$ triplet and the adjacent MgH feature, offering sensitivity to stellar surface gravity. The data reduction pipeline incorporates an improved astrometric solution anchored to Gaia DR3 and a photometric calibration strategy tied to synthetic photometry from Gaia XP spectra. These procedures yield a point-source depth of $\sim$18.9 mag at S/N$\sim$10 and an internal photometric precision $\approx$6-7 mmag at the bright end. A preliminary color--color analysis using Gaia broadband photometry confirms the expected sensitivity of the DDO51 band to stellar surface gravity, demonstrating a clear photometric separation between dwarf and giant sequences for late-type stars. This dataset, when combined with existing SAGES photometry in other bands, provides a crucial tool for disentangling the substructures of the Milky Way. All data products from this release upon publication will be available.

astro-ph.SR

StarCLR: Contrastive Learning Representation for Astronomical Light Curves

With the rapid development of time-domain surveys, the availability of massive light curve data offers new opportunities for studying stellar evolution and variable star classification, while simultaneously posing challenges for feature extraction and modeling. We present StarCLR, a contrastive pretraining framework for large-scale light curves. By constructing positive pairs from partially overlapping sub-sequences, StarCLR encourages the model to learn temporal representations. We pretrain StarCLR on the TESS dataset and fine-tune it for variable star classification on three surveys with distinct observational characteristics, namely TESS (18 types), ZTF (11 types), and Gaia (24 types). StarCLR achieves macro-F1 scores of 84.35%, 87.82%, and 92.73%, and micro-F1 scores of 94.46%, 92.83%, and 99.49%, respectively. Compared with LSTM and Transformer trained from scratch, StarCLR performs better on TESS and ZTF, with the largest gains on sparsely sampled ZTF light curves, demonstrating promising generalization. For Gaia, which involves a broader class space, the evaluation is not directly comparable, and performance is likely influenced by astrophysical features, resulting in a more limited contribution from the pretrained backbone. Systematic ablations on embedding design, pooling strategy, and pretraining settings further indicate that the pretrained representations provide performance gains by capturing informative temporal characteristics of light curves. Looking ahead, with standardized datasets and more diverse labeling schemes, the generalization ability of StarCLR can be further enhanced.

astro-ph.SR

The Spectroscopic and Photometric Study of a Star Cluster Sample in Andromeda Halo

Halo star clusters serve as vital tracers for the formation and evolution of the Andromeda galaxy. In this work, we present physical parameters for 29 M31 halo star clusters, derived from a combination of spectroscopic and photometric data. Low-resolution spectra were acquired using the BFOSC spectrograph on the NAOC Xinglong 2.16-m telescope. For the photometric analysis, we utilized uSC and vSAGE bands from the SAGE survey, complemented by archival data from GALEX(NUV, FUV), PAN-STARRS(grizy) and the 2MASS(JHK). Ages and metallicities were determined via ULySS (Vazdekis et al. and pegase-hr) SSP model and the Bruzual & Charlot (2003) (BC03) stellar population synthesis models. The derived parameters show good agreement with literature values. Notably, for three of these clusters, this study represents the first combined photometric and spectroscopic analysis.

astro-ph.GA

Spec-o3: A Tool-Augmented Vision-Language Agent for Rare Celestial Object Candidate Vetting via Automated Spectral Inspection

Due to the limited generalization and interpretability of deep learning classifiers, The final vetting of rare celestial object candidates still relies on expert visual inspection--a manually intensive process. In this process, astronomers leverage specialized tools to analyze spectra and construct reliable catalogs. However, this practice has become the primary bottleneck, as it is fundamentally incapable of scaling with the data deluge from modern spectroscopic surveys. To bridge this gap, we propose Spec-o3, a tool-augmented vision-language agent that performs astronomer-aligned spectral inspection via interleaved multimodal chain-of-thought reasoning. Spec-o3 is trained with a two-stage post-training recipe: cold-start supervised fine-tuning on expert inspection trajectories followed by outcome-based reinforcement learning on rare-type verification tasks. Evaluated on five rare-object identification tasks from LAMOST, Spec-o3 establishes a new State-of-the-Art, boosting the macro-F1 score from 28.3 to 76.5 with a 7B parameter base model and outperforming both proprietary VLMs and specialized deep models. Crucially, the agent demonstrates strong generalization to unseen inspection tasks across survey shifts (from LAMOST to SDSS/DESI). Expert evaluations confirm that its reasoning traces are coherent and physically consistent, supporting transparent and trustworthy decision-making. Code, data, and models are available at https://github.com/Maxwell-Jia/spec-o3.

cs.CL

Variability of H$\alpha$ chromospheric activity of solar-like stars revealed by the time-domain data of LAMOST Medium-Resolution Spectroscopic Survey

The variability of H$\alpha$ chromospheric activity of solar-like stars is investigated by using the time-domain data of LAMOST Medium-Resolution Spectroscopic Survey (MRS). We use $R_\mathrm{H\alpha}$ index (ratio of H$\alpha$ luminosity to bolometric luminosity) to measure the H$\alpha$ activity intensity of a spectrum, and utilize the median of the $R_\mathrm{H\alpha}$ values of multiple observations ($R_\mathrm{H\alpha}^\mathrm{median}$) as the representative activity intensity of a stellar source. The H$\alpha$ variability of a stellar source is indicated by the extent of $R_\mathrm{H\alpha}$ fluctuation ($R_\mathrm{H\alpha}^\mathrm{EXT}$) of multiple observations. Our sample shows that $R_\mathrm{H\alpha}^\mathrm{EXT}$ of solar-like stars is about one order of magnitude smaller than $R_\mathrm{H\alpha}^\mathrm{median}$. The distribution of $\log R_\mathrm{H\alpha}^\mathrm{EXT}$ versus $\log R_\mathrm{H\alpha}^\mathrm{median}$ reveals the distinct behaviors between the stellar source categories with lower ($\log R_\mathrm{H\alpha}^\mathrm{median} < -4.85$) and higher ($\log R_\mathrm{H\alpha}^\mathrm{median} > -4.85$) activity intensity. For the former stellar source category, the top envelope of the distribution first increases and then decreases with $\log R_\mathrm{H\alpha}^\mathrm{median}$; while for the latter category, the top envelope of the distribution is largely along a positive correlation line. In addition, for the stellar sources with lower activity intensity, the large-$\log R_\mathrm{H\alpha}^\mathrm{EXT}$ objects near the top envelope of the $\log R_\mathrm{H\alpha}^\mathrm{EXT}$ versus $\log R_\mathrm{H\alpha}^\mathrm{median}$ distribution tend to have long-term and regular variations of H$\alpha$ activity; while for the stellar sources with higher activity intensity, the H$\alpha$ variations are more likely to be random fluctuations.

astro-ph.SR

In-Orbit GRB Identification Using LLM-based model for the CXPD CubeSat

To validate key technologies for wide field-of-view (FOV) X-ray polarization measurements, the Cosmic X-ray Polarization Detector (CXPD) CubeSat series has been developed as a prototype platform for the Low-Energy X-ray Polarization Detector (LPD) onboard the POLAR-2 mission. The wide-FOV design significantly increases the complexity of the background environment, posing notable challenges for real-time gamma-ray burst (GRB) identification. In this work, we propose an in-orbit GRB identification method based on machine learning, using simulated spectral data as input. A training dataset was constructed using a Geant4-based simulator, incorporating in-orbit background and GRB events modeled within the 2-10 keV energy range. To meet the computational constraints of onboard processing, we employ a multimodal large language model (MLLM), which is fine-tuned using low-rank adaptation (LoRA) based on miniCPM-V2.6 and quantized to 4-bit precision. The model achieves perfect classification accuracy on validation data and demonstrates strong regression performance in estimating GRB spectral indices, with an RMSE of 0.118. Furthermore, we validate the feasibility of onboard deployment through a simulated satellite data processing pipeline, highlighting the potential of our approach to enable future real-time GRB detection and spectral analysis in orbit.

astro-ph.IM

The Stellar Abundances and Galactic Evolution Survey (SAGES). IV. Surface Gravity Estimation and Giant-Dwarf Separation with the DDO51 Filter

Reliable estimation of stellar surface gravity (log $g$) for a large sample is crucial for evaluating stellar evolution models and understanding galactic structure; However, it is not easy to accomplish due to the difficulty in gathering a large spectroscopic data set. Photometric sky survey using a specific filter, on the other hand, can play a substantial role in the assessment of log $g$. The Stellar Abundances and Galactic Evolution Survey (SAGES) utilizes eight filters to provide accurate stellar parameters for $\sim10^{7}$ stars, with its DDO51 intermediate-band filter specifically designed for robust log $g$ determination. In this work, the observed SAGES $u_{\rm SC}$ and $v_{\rm SAGES}$ photometry, the synthetic photometry in $g$, $r$, $i$, and DDO51 bands derived from \textit{Gaia} XP spectra are employed to investigate the importance of the DDO51 filter in the determination of log $g$. We applied machine-learning-based extinction correction and employed XGBoost models, trained on stellar parameters from LAMOST, to predict log $g$ using photometric data. By comparing model predicted log $g$ with LAMOST values, we find that including DDO51 filter improve the accuracies of log $g$ estimates by 21.0\% (from 0.224\,dex to 0.177\,dex) overall, and by 26.5\% (from 0.302\,dex to 0.222\,dex ) for GK-type stars, as compared to those obtained without DDO51. The DDO51 filter is also validated to be particularly effective for metal-poor stars ([Fe/H]$<$-1.0), where it significantly mitigates systematic biases. Our findings highlight the diagnostic power of the SAGES DDO51 filter, providing enhanced stellar characterization vital for future in-depth studies of the Milky Way.

astro-ph.SR

Looking for Signs of Unresolved Binarity in the Continuum of LAMOST Stellar Spectra

We describe an attempt to derive the binarity rate of samples of 166 A-, F-, G-, and K-type stars from LAMOST DR5 and 1000 randomly selected presumably single stars from Gaia DR3 catalogs. To this end, we compared continua of the observed spectra with the continua of synthetic spectra in the 3700 to 9097 Angstrom range. The latter spectra were reduced to the LAMOST set of wavelengths, while the former ones were smoothed. Next, we searched for every observed star the nearest synthetic spectrum using a four-parameter representation - effective temperature, gravity, [Fe/H], and a range of interstellar absorption values. However, rms deviations of observed spectra from synthetic ones appeared to be not sufficient to claim that any of the stars is a binary. We conclude that comparison of the intensity of pairs of spectral lines remains the best way to detect binarity.

astro-ph.SR

Sky Background Building of Multi-objective Fiber spectra Based on Mutual Information Network

Sky background subtraction is a critical step in Multi-objective Fiber spectra process. However, current subtraction relies mainly on sky fiber spectra to build Super Sky. These average spectra are lacking in the modeling of the environment surrounding the objects. To address this issue, a sky background estimation model: Sky background building based on Mutual Information (SMI) is proposed. SMI based on mutual information and incremental training approach. It utilizes spectra from all fibers in the plate to estimate the sky background. SMI contains two main networks, the first network applies a wavelength calibration module to extract sky features from spectra, and can effectively solve the feature shift problem according to the corresponding emission position. The second network employs an incremental training approach to maximize mutual information between representations of different spectra to capturing the common component. Then, it minimizes the mutual information between adjoining spectra representations to obtain individual components. This network yields an individual sky background at each location of the object. To verify the effectiveness of the method in this paper, we conducted experiments on the spectra of LAMOST. Results show that SMI can obtain a better object sky background during the observation, especially in the blue end.

cs.CV

Phase II of the LAMOST-Kepler/K2 Survey. II. Time Domain of Medium-resolution Spectroscopic Observations from 2018 to 2023

The LAMOST-Kepler/K2 Medium-Resolution Spectroscopic Survey (LK-MRS) conducted time-domain medium-resolution spectroscopic observations of 20 LAMOST plates in the Kepler and K2 fields from 2018 to 2023, a phase designated as LK-MRS-I. A catalog of stellar parameters for a total of 36,588 stars, derived from the spectra collected during these five years, including the effective temperature, the surface gravity, the metallicity, the {\alpha}-element abundance, the radial velocity, and v sin i of the target stars, is released, together with the weighted averages and uncertainties. At S/N = 10, the measurement uncertainties are 120 K, 0.18 dex, 0.13 dex, 0.08 dex, 1.9 km/s, and 4.0 km/s for the above parameters, respectively. Comparisons with the parameters provided by the APOGEE and GALAH surveys validate the effective temperature and surface gravity measurements, showing minor discrepancies in metallicity and {\alpha}-element abundance values. We identified some peculiar star candidates, including 764 metal-poor stars, 174 very metal-poor stars, and 30 high-velocity stars. Moreover, we found 2,333 stars whose radial velocity seems to be variable. Using Kepler/K2 or TESS photometric data, we confirmed 371 periodic variable stars among the radial velocity variable candidates and classified their variability types. LK-MRS-I provides spectroscopic data being useful for studies of the Kepler and K2 fields. The LK-MRS project will continue collecting time-domain medium-resolution spectra for target stars during the third phase of LAMOST surveys, providing data to support further scientific research.

astro-ph.SR

StellarF: A Physics-Informed LoRA Framework for Stellar Flare Forecasting with Historical & Statistical Data

Stellar flare forecasting represents a critical frontier in astrophysics, offering profound insights into stellar activity mechanisms and exoplanetary habitability assessments. Yet the inherent unpredictability of flare activity, rooted in stellar diversity and evolutionary stages, underpins the field's core challenges: (1) sparse, incomplete, noisy lightcurve data from traditional observations; (2) ineffective multi-scale flare evolution capture via single representations; (3) poor physical interpretability in data-driven models lacking physics-informed priors. To address these challenges, we propose StellarF, a physics-informed framework synergizing general Al with astrophysical domain knowledge via three core components: a unified preprocessing pipeline for lightcurve refinement (missing-value imputation, temporal patch partitioning, adaptive sample filtering); a Low-Rank Adaptation (LoRA)-finetuned large language model (LLM) backbone enhanced by first-order difference augmentation, flare statistical information, and flare historical record modules for multimodal fusion instead of only simple representations; and a novel physics-informed loss embedding a minimum rising rate prior, appended to the cross-entropy loss, to align with flare physics. Extensive experiments on Kepler and TESS datasets show StellarF achieves state-of-the-art performance across key metrics, setting new benchmarks for flare forecasting. This work bridges general AI with astrophysics, offering a practical, physically interpretable paradigm for transient event forecasting in time-domain astronomy.

cs.LG

FALCO: a Foundation model of Astronomical Light Curves for time dOmain astronomy

Time-domain surveys have advanced astronomical research by revealing diverse variable phenomena, from stellar flares to transient events. The scale and complexity of survey data, along with the demand for rapid classification, present significant challenges for analysis. While machine learning offers solutions, most existing models are tailored to single tasks, struggle to generalize, and depend heavily on large, accurately labeled datasets. We introduce FALCO, a foundation model for astronomical light curve analysis in time-domain astronomy. This work presents the initial version of FALCO trained via self-supervised learning on unlabeled Kepler light curves using a Transformer-based architecture. The model has been evaluated on three distinct tasks and demonstrates strong generalization: achieving 95 percent accuracy in stellar variability classification across eight classes, an overall RMSE of 0.1305 dex in surface gravity estimation (notably improved to below 0.08 dex when log g is less than 1, and approximately 0.02 dex near log g equals 3), and 87 percent precision in flare identification. These results highlight the model's versatility and ability to learn generalizable representations from light curves, enabling straightforward adaptation to diverse tasks. We further analyze the impact of model scaling and sequence length, finding performance improves with larger models and longer input sequences. We also apply FALCO to derive surface gravity (log g) measurements for 179,732 Kepler stars from their light curves.

astro-ph.IM

The Mini-SiTian Array: A Pathfinder for the SiTian Project

The Mini-SiTian Array serves as a pathfinder for the SiTian project, which aims to survey the entire sky in $gri$ bands every 30 minutes, reaching a limiting magnitude of 21. This special issue features 11 papers covering the design, operation, data reduction, and early scientific results from two years of Mini-SiTian observations. The insights gained from these pathfinder experiments represent a significant milestone toward the full realization of the SiTian project.

astro-ph.IM

FLARE: A Framework for Stellar Flare Forecasting using Stellar Physical Properties and Historical Records

Stellar flare events are critical observational samples for astronomical research; however, recorded flare events remain limited. Stellar flare forecasting can provide additional flare event samples to support research efforts. Despite this potential, no specialized models for stellar flare forecasting have been proposed to date. In this paper, we present extensive experimental evidence demonstrating that both stellar physical properties and historical flare records are valuable inputs for flare forecasting tasks. We then introduce FLARE (Forecasting Light-curve-based Astronomical Records via features Ensemble), the first-of-its-kind large model specifically designed for stellar flare forecasting. FLARE integrates stellar physical properties and historical flare records through a novel Soft Prompt Module and Residual Record Fusion Module. Our experiments on the publicly available Kepler light curve dataset demonstrate that FLARE achieves superior performance compared to other methods across all evaluation metrics. Finally, we validate the forecast capability of our model through a comprehensive case study.

astro-ph.SR

The Stellar Abundances and Galactic Evolution Survey (SAGES). II. Machine Learning-Based Stellar parameters for 21 million stars from the First Data Release

Stellar parameters for large samples of stars play a crucial role in constraining the nature of stars and stellar populations in the Galaxy. An increasing number of medium-band photometric surveys are presently used in estimating stellar parameters. In this study, we present a machine-learning approach to derive estimates of stellar parameters, including [Fe/H], logg, and Teff, based on a combination of medium-band and broad-band photometric observations. Our analysis employs data primarily sourced from the SAGE Survey , which aims to observe much of the Northern Hemisphere. We combine the $uv$-band data from SAGES DR1 with photometric and astrometric data from Gaia EDR3, and apply the random forest method to estimate stellar parameters for approximately 21 million stars. We are able to obtain precisions of 0.09 dex for [Fe/H], 0.12 dex for logg, and 70 K for Teff. Furthermore, by incorporating 2MASS and WISE infrared photometric and GALEX ultraviolet data, we are able to achieve even higher precision estimates for over 2.2 million stars. These results are applicable to both giant and dwarf stars. Building upon this mapping, we construct a foundational dataset for research on metal-poor stars, the structure of the Milky Way, and beyond. With the forthcoming release of additional bands from SAGE Survey such DDO51 and H-alpha, this versatile machine learning approach is poised to play an important role in upcoming surveys featuring expanded filter sets

astro-ph.SR

LAMOST medium-resolution spectroscopic survey of Galactic Open Clusters (LAMOST-MRS-O): An overview of survey plan and preliminary results

As part of the LAMOST medium-resolution spectroscopic survey, the LAMOST-MRS-O is a non-time domain survey that aims to perform medium-resolution spectral observations for member stars in the open cluster area. This survey plans to obtain the spectroscopic parameters such as radial velocity and metal abundances of member stars and provide data support for further study on the chemical and dynamical characteristics and evolution of open clusters in combination with Gaia data. We have completed the observations on ten open cluster fields and obtained 235184 medium-resolution spectra of 133792 stars. Based on the data analyzed of LAMOST DR11V1.1, for some clusters of particular concern, it is found that the sampling ratio of members stars with Gmag < 15 mag can reach 70%, which indicates that the LAMOST-MRS-O has reached our initial design goal.

astro-ph.GA