SearcharxivSearch

arXiv subjects

Wujun Shao

Publications and source records attributed to Wujun Shao.

4 recordsLinked to original sources

M-EPDet: Real-Time Real-Bogus Classification and Transient Candidate Judgement for the EP-WXT Pipeline via Multi-Modal Data

The Wide-field X-ray Telescope (WXT) onboard the Einstein Probe (EP) produces a large post-detection candidate stream in which genuine astrophysical sources coexist with instrumental artifacts and Cosmic Ray events. We present M-EPDet, a three-step post-detection framework for real-time candidate vetting in EP-WXT lobster-eye Micro-pore Optics (MPO) data. The framework combines a ResNet-based Arm filter, a dual-branch temporal-spectral Cosmic Ray filter, and a background-aware Bayesian Blocks module for single-exposure variability screening. Using on-orbit EP-WXT observations, we report decoupled metrics for the cascading system. M-EPDet achieves a Real-Bogus Recall of 98.31\% ($98.53\% \times 99.78\%$) for genuine astrophysical sources, together with rejection rates of 92.99\% for instrumental artifacts and 98.18\% for Cosmic Ray events. In the final step, the Bayesian Blocks module flags 0.75\% of the post-filtration observations, corresponding to a 99.25\% reduction in candidate volume. The system is deployed in the EP-WXT pipeline as a lightweight real-time service, reducing the manual-inspection burden in candidate vetting.

astro-ph.IM

A Value-added Physical Properties Catalog for Low-redshift Galaxies from DESI Legacy Imaging Surveys DR10

Galaxy physical properties-such as star formation rate (SFR), stellar mass, and gas-phase metallicity-are essential for population studies and evolutionary analyses. Deriving these quantities for billions of galaxies in modern imaging surveys presents significant challenges due to limited spectroscopy and the computational costs associated with traditional spectral energy distribution fitting. As a result, many galaxies in large photometric surveys still lack homogeneous property estimates. This study introduces a multimodal deep learning model that integrates optical imaging with photometric catalog features to estimate SFR, stellar mass, and oxygen abundance in low-redshift galaxies. The model incorporates a ResNet-based convolutional neural network to extract spatial information from multiband images and a multilayer perceptron that processes catalog-level photometric features, leveraging complementary constraints from morphology, surface brightness, and broadband colors. Trained on reference measurements from the MPA-JHU DR8 catalog, the model is optimized for efficient large-scale estimation. When applied to the DESI Legacy Imaging Surveys (LS) DR10, the model generates a value-added catalog containing physical property estimates for approximately 547 million galaxies with redshifts z <= 0.5. Validation through comparisons with independent catalogs and exploration of key scaling relations demonstrates that while the derived properties are not intended for precision measurements of individual objects, they effectively capture the dominant astrophysical trends necessary for ensemble studies. This catalog represents the first homogeneous set of photometry-based SFR, stellar mass, and metallicity estimates for DESI LS DR10, providing a vital resource for statistical studies of galaxies in the local Universe and facilitating comparisons with current and future spectroscopic surveys.

astro-ph.GA

Deep learning-based astronomical multimodal data fusion: A comprehensive review

With the rapid advancements in observational technologies and the widespread implementation of large-scale sky surveys, diverse electromagnetic wave data (e.g., optical and infrared) and non-electromagnetic wave data (e.g., gravitational waves) have become increasingly accessible. Astronomy has thus entered an unprecedented era of data abundance and complexity. Astronomers have long relied on unimodal data analysis to perceive the universe, but these efforts often provide only limited insights when confronted with the current massive and heterogeneous astronomical data. In this context, multimodal data fusion (MDF), as an emerging method, provides new opportunities to enhance the value of astronomical data and deepening the understanding of the universe by integrating information from different modalities. Recent progress in artificial intelligence (AI), particularly in deep learning (DL), has greatly accelerated the development of multimodal research in astronomy. Therefore, a timely review of this field is essential. This paper begins by discussing the motivation and necessity of astronomical MDF, followed by an overview of astronomical data sources and major data modalities. It then introduces representative DL models commonly used in astronomical multimodal studies, the general fusion process as well as various fusion strategies, emphasizing their characteristics, applicability, advantages, and limitations. Subsequently, the paper surveys existing astronomical multimodal studies and datasets. Finally, the discussion section synthesizes key findings, identifies potential challenges, and suggests promising directions for future research. By offering a structured overview and critical analysis, this review aims to inspire and guide researchers engaged in DL-based MDF in astronomy.

astro-ph.IM

Astronomical Knowledge Entity Extraction in Astrophysics Journal Articles via Large Language Models

Astronomical knowledge entities, such as celestial object identifiers, are crucial for literature retrieval and knowledge graph construction, and other research and applications in the field of astronomy. Traditional methods of extracting knowledge entities from texts face challenges like high manual effort, poor generalization, and costly maintenance. Consequently, there is a pressing need for improved methods to efficiently extract them. This study explores the potential of pre-trained Large Language Models (LLMs) to perform astronomical knowledge entity extraction (KEE) task from astrophysical journal articles using prompts. We propose a prompting strategy called Prompt-KEE, which includes five prompt elements, and design eight combination prompts based on them. Celestial object identifier and telescope name, two most typical astronomical knowledge entities, are selected to be experimental object. And we introduce four currently representative LLMs, namely Llama-2-70B, GPT-3.5, GPT-4, and Claude 2. To accommodate their token limitations, we construct two datasets: the full texts and paragraph collections of 30 articles. Leveraging the eight prompts, we test on full texts with GPT-4 and Claude 2, on paragraph collections with all LLMs. The experimental results demonstrated that pre-trained LLMs have the significant potential to perform KEE tasks in astrophysics journal articles, but there are differences in their performance. Furthermore, we analyze some important factors that influence the performance of LLMs in entity extraction and provide insights for future KEE tasks in astrophysical articles using LLMs.

astro-ph.IM