SearcharxivSearch

arXiv subjects

Tu Guo

Publications and source records attributed to Tu Guo.

4 recordsLinked to original sources

PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research

The paradigm of agentic science requires AI systems to conduct robust reasoning and engage in long-horizon, autonomous exploration. However, current scientific benchmarks remain confined to domain knowledge comprehension and complex reasoning, failing to evaluate the exploratory nature and procedural complexity of real-world research. In this work, we present research-oriented evaluations in theoretical and computational physics, a natural testbed with comprehensive domain knowledge, complex reasoning, and verifiable end-to-end workflows without reliance on experiments. Here we introduce PRL-Bench (Physics Research by LLMs), a benchmark designed to systematically map the capability boundaries of LLMs in executing end-to-end physics research. Constructed from 100 curated papers from the latest issues of Physical Review Letters since August 2025 and validated by domain experts, PRL-Bench covers five major theory- and computation-intensive subfields of modern physics: astrophysics, condensed matter physics, high-energy physics, quantum information, and statistical physics. Each task in the benchmark is designed to replicate the core properties of authentic scientific research, including exploration-oriented formulation, long-horizon workflows, and objective verifiability, thereby reconstructing the essential reasoning processes and research workflows of real physics research. Evaluation across frontier models shows that performance remains limited, with the best overall score below 50, revealing a pronounced gap between current LLM capabilities and the demands of real scientific research. PRL-Bench serves a reliable testbed for accessing next generation AI scientists advancing AI systems toward autonomous scientific discovery.

cs.LG

Impact of Resonant Compton Scattering on Magnetar X-Ray Polarization with QED Vacuum Resonance

Recent obeservations have revealed significant soft X-ray polarizations from several quiescent magnetars, including the intriguing $90^\deg$ polarization angle (PA) swing as a function of photon energy for some sources. We present a general semi-analytical framework for calculating energy-dependent soft X-ray polarization signatures from magnetars, consistently incorporating both QED vacuum resonance in the atmosphere and resonant Compton scattering (RCS) in the magnetosphere. Starting from the polarized radiative transfer equation for RCS and treating vacuum-resonance-induced mode conversion as an input, we employ a first-order approximation in RCS optical depth to evaluate the effect of different magnetospheric plasma density (which depends on magnetic twist), drift velocity and temperature, and viewing geometry on the observed radiation. Our analysis reveals that magnetic twist and plasma drift velocity are the critical parameters controlling the impact of RCS on both the absolute polarization degree and its variation across the soft X-ray spectrum. We find that sufficiently strong RCS can wash out the PA swing caused by vacuum resonance. Furthermore, in addition to the QED vacuum resonance effect, significant relativistic signatures arising from plasma drift velocity ($\beta_0 \gtrsim 0.5$) may introduce an extra $90^\circ$ PA swing in the spectrum. Our calculation framework, based on single-scattering approximation, bypasses the need for complex, multi-dimensional Monte Carlo simulations, providing an analytical pathway for modeling full-surface emission and rotational-phase-resolved radiation from magnetic neutron stars, in support of current and future X-ray polarization missions.

astro-ph.HE

PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research

Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because research requires deep domain expertise, long-horizon reasoning, and reliable numerical computation. We introduce PRL-Bench, a research-reproduction benchmark adapted from 100 Physical Review Letters papers across major areas of modern physics. PRL-Bench distills realistic research workflows into traceable tasks with explicit intermediate artifacts and diverse evaluation rubrics; each task is estimated by domain experts to require more than six hours for a specialized PhD student to reproduce independently. Evaluations show that existing agents remain unreliable on extended research workflows. We therefore present PhysMaster, a scientific agent combining adaptive MCTS-based multi-trajectory exploration with hierarchical memory to improve long-horizon robustness and knowledge accumulation. PhysMaster achieves the highest overall PRL-Bench score of 51.08, outperforming Codex, OpenHands, OpenClaw, and ReAct, and yields relative improvements of 14.13 percent to 93.38 percent across backbone models. Error analysis shows that PhysMaster substantially reduces failures from incomplete long-horizon execution, while remaining bottlenecks lie in physics knowledge and analytical reasoning. Together, PRL-Bench and PhysMaster provide a rigorous benchmark and effective system for advancing autonomous AI research in frontier physics.

cs.AI

Power corrections in the determination of heavy meson LCDAs: A renormalon-based estimation

At leading power accuracy the QCD light-cone distribution amplitudes (LCDAs) for a heavy meson can be matched onto the LCDAs in the framework of heavy-quark effective theory (HQET) through a factorization formula. We examine the power corrections to this factorization in the renormalon model, which can associate the power corrections originating from high-twist contributions to the divergent series in a matching kernel. Our analysis indicates that the dominant power corrections originate from the virtual part of the vertex bubble chain diagrams, which generate poles at $w=n+\frac{1}{2},\forall n\in \mathbb{N}$ and $w=1$ in the Borel plane. Employing phenomenological models for both HQET and QCD LCDA, we present a numerical estimate. The results indicate that the power corrections in the peak region are approximately $22\%$ for the D meson and $7\%$ for the $\overline{\mathrm{B}}$ meson. These findings showcase the magnitude and the potential importance of power corrections in achieving high-precision phenomenological predictions for heavy mesons.

hep-ph