SearcharxivSearch

arXiv subjects

Shengdu Chai

Publications and source records attributed to Shengdu Chai.

6 recordsLinked to original sources

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We present ResearchClawBench, a benchmark for evaluating autonomous scientific research across 40 tasks from 10 scientific domains. Each task is grounded in a real published paper, provides related literature and raw data, and hides the target paper during evaluation. Expert-curated multimodal rubrics decompose the target scientific artifacts into weighted criteria, enabling evaluation of target-paper-level re-discovery while leaving room for new discovery. We evaluate seven autonomous research (auto-research) agents under a unified protocol and seventeen native LLMs through the lightweight ResearchHarness. Current systems remain far from reliable re-discovery: the strongest autonomous agent, Claude Code, averages 21.5, and the strongest ResearchHarness LLM, Claude-Opus-4.7, averages 20.7, with an LLM frontier mean of only 26.5. Error analysis shows that failures concentrate in experimental protocol mismatch, evidence mismatch, and missing scientific core. ResearchClawBench provides a reproducible evaluation frontier for measuring progress toward autonomous scientific research.

cs.LG

Revisiting the Broken Symmetry Phase of Solid Hydrogen: A Neural Network Variational Monte Carlo Study

The crystal structure of high-pressure solid hydrogen remains a fundamental open problem. Although the research frontier has mostly shifted toward ultra-high pressure phases above 400 GPa, we show that even the broken symmetry phase observed around 130~GPa requires revisiting due to its intricate coupling of electronic and nuclear degrees of freedom. Here, we develop a first principle quantum Monte Carlo framework based on a deep neural network wave function that treats both electrons and nuclei quantum mechanically within the constant pressure ensemble. Our calculations reveal an unreported ground-state structure candidate for the broken symmetry phase with $Cmcm$ space group symmetry, and we test its stability up to 96 atoms. The predicted structure quantitatively matches the experimental equation of state and X-ray diffraction patterns. Furthermore, our group-theoretical analysis shows that the $Cmcm$ structure is compatible with existing Raman and infrared spectroscopic data. Crucially, static density functional theory calculation reveals the $Cmcm$ structure as a dynamically unstable saddle point on the Born-Oppenheimer potential energy surface, demonstrating that a full quantum many-body treatment of the problem is necessary. These results shed new light on the phase diagram of high-pressure hydrogen and call for further experimental verifications.

cond-mat.str-el

Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows

Despite advances in scientific AI, a coherent framework for Scientific General Intelligence (SGI)-the ability to autonomously conceive, investigate, and reason across scientific domains-remains lacking. We present an operational SGI definition grounded in the Practical Inquiry Model (PIM: Deliberation, Conception, Action, Perception) and operationalize it via four scientist-aligned tasks: deep research, idea generation, dry/wet experiments, and experimental reasoning. SGI-Bench comprises over 1,000 expert-curated, cross-disciplinary samples inspired by Science's 125 Big Questions, enabling systematic evaluation of state-of-the-art LLMs. Results reveal gaps: low exact match (10--20%) in deep research despite step-level alignment; ideas lacking feasibility and detail; high code executability but low execution result accuracy in dry experiments; low sequence fidelity in wet protocols; and persistent multimodal comparative-reasoning challenges. We further introduce Test-Time Reinforcement Learning (TTRL), which optimizes retrieval-augmented novelty rewards at inference, enhancing hypothesis novelty without reference answer. Together, our PIM-grounded definition, workflow-centric benchmark, and empirical insights establish a foundation for AI systems that genuinely participate in scientific discovery.

cs.AI

From Optimal Observables to Machine Learning: an Effective-Field-Theory Analysis of $e^+e^- \to W^+W^-$ at Future Lepton Colliders

We apply machine-learning techniques to the effective-field-theory analysis of the $e^+e^- \to W^+W^-$ processes at future lepton colliders, and demonstrate their advantages in comparison with conventional methods, such as optimal observables. Compared to traditional algorithms, we show that simulation-based inference methods are more robust to detector effects and backgrounds, and could in principle produce unbiased results with sufficient Monte Carlo simulation samples that accurately describe experiments. This is crucial for the analyses at future lepton colliders given the outstanding precision of the $e^+e^- \to W^+W^-$ measurement ($\sim 10^{-4}$ in terms of anomalous triple gauge couplings or even better) that can be reached. Our framework can be generalized to other effective-field-theory analyses, such as the one of $e^+e^- \to t\bar{t}$ or similar processes at muon colliders.

hep-ph

Measurement of the Chern Number for Non-Hermitian Chern Insulators

The identification of the topological invariant of a topological system is crucial in experiments. However, due to the inherent non-Hermitian features, such determination is notably challenging in non-Hermitian systems. Here, we propose that the magnetic effect can be utilized to measure the Chern number of the non-Hermitian Chern insulator. We find that the splitting of non-Hermitian bands under the magnetic field is Chern number dependent. Consequently, one can easily identify the Chern number by analyzing these splitting sub-bands. From the experimental perspective, the measurement of non-Hermitian bands is demonstrated in LC electric circuits. Furthermore, we find that the non-Hermiticity can drive open (closed) orbits of sub-bands in the Hermitian limit closed (open), which can also be identified by our proposal. These phenomena highlight the distinctive capabilities of non-Hermitian systems. Our results facilitate the detection of Chern numbers for non-Hermitian systems and may motivate further studies of their topological properties.

cond-mat.mes-hall

Accommodating the CDF W-boson Mass Measurement in the Beautiful Mirror Model

The W-boson mass measurement recently reported by the CDF II experiment exhibits a significant deviation from both the Standard Model prediction and previous measurements. There is also a long-standing deviation between the Standard Model prediction of the forward-backward asymmetry of the bottom quark ($A^{0,b}_{\rm FB}$) and its measurement at the LEP experiment. The Beautiful Mirror model, proposed to resolve the $A^{0,b}_{\rm FB}$ discrepancy, introduces vector-like quarks that modify the W-boson mass at one-loop level. In this study, we find an interesting region in the model parameter space that could potentially explain both discrepancies, which puts the new quarks in the multi-TeV region. This region is mostly consistent with current LHC bounds from direct searches and Higgs coupling measurements, but will be thoroughly probed at the High Luminosity LHC. As such, the Beautiful Mirror model as an explanation of the $m_W$ and $A^{0,b}_{\rm FB}$ discrepancies could be confirmed or falsified in the near future.

hep-ph