SearcharxivSearch

arXiv subjects

Mingshen Zhou

Publications and source records attributed to Mingshen Zhou.

3 recordsLinked to original sources

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review

We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent papers) and metadata mismatches (real papers with one or more incorrect fields). We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.

astro-ph.IM

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.

cs.CL

Tracing the light: Identification for the optical counterpart candidates of binary black-holes during O3

The accretion disks of active galactic nuclei (AGN) are widely considered the ideal environments for binary black hole (BBH) mergers and the only plausible sites for their electromagnetic (EM) counterparts. Graham et al.(2023) identified seven AGN flares that are potentially associated with gravitational-wave (GW) events detected by the LIGO-Virgo-KAGRA (LVK) Collaboration during the third observing run. In this article, utilizing an additional three years of Zwicky Transient Facility (ZTF) public data after their discovery, we conduct an updated analysis and find that only three flares can be identified. By implementing a joint analysis of optical and GW data through a Bayesian framework, we find two flares exhibit a strong correlation with GW events, with no secondary flares observed in their host AGN up to 2024 October 31. Combining these two most robust associations, we derive a Hubble constant measurement of $H_{0}= 72.1^{+23.9}_{-23.1} \ \mathrm{km \ s^{-1} Mpc^{-1}}$ and incorporating the multi-messenger event GW170817 improves the precision to $H_{0}=73.5^{+9.8}_{-6.9} \ \mathrm{km \ s^{-1} Mpc^{-1}}$. Both results are consistent with existing measurements reported in the literature.

astro-ph.HE