SearcharxivSearch

arXiv subjects

Zhengde Zhang

Publications and source records attributed to Zhengde Zhang.

5 recordsLinked to original sources

Data Preservation in High Energy Physics: Global Report 2026

This document summarizes the contributions to the 5th DPHEP workshop March 5-6, 2026, CERN, and reflects the advancements since 2024, as well as future milestones and tendencies. Impressive progress in HEP data preservation is observed. Legacy data revival was showcased through successful reanalysis of archived data using contemporary methods, demonstrating the long-term scientific value of preservation. Sustainability challenges were noted, emphasizing the need for long-term funding and institutional support to maintain data preservation infrastructure, particularly for legacy experiments transitioning to archival modes. Innovative transverse projects display constant progress towards common technologies for a robust and transferrable DP. In particular, there is a clear shift toward automation, with increasing use of AI and machine learning for data curation, metadata extraction, and workflow optimization. Open science momentum is growing, with wider adoption of FAIR principles and open data policies, and experiments committing to public releases.

hep-ex

Rongzai agent: A Large Language Model-Based Autonomous Assistant for Rietveld Refinement of Neutron Diffraction Data

Neutron diffraction (ND) is an indispensable technique for determining atomic positions (especially light elements) and thus serves as a critical probe for revealing microscopic structures in materials science. However, traditional Rietveld refinement of ND data relies heavily on manual operation of specialized software, which is time-consuming, labor-intensive, and highly dependent on user expertise, severely hindering automated analysis. The automation of Rietveld refinement has long been a long-standing and challenging problem in crystallography. To address this challenge, this paper presents the Dr.Sai-Rongzai agent, an autonomous refinement assistant based on a large language model (LLM), a specialist knowledge base, and the GSAS-II refinement engine, achieving for the first time an intelligent refinement that integrates knowledge-driven decision-making. The agent accomplishes a fully automated workflow from natural language task parsing to autonomous decision-making, execution of refinement strategies, and report generation. Evaluation on five representative samples shows that the Rongzai agent achieves lower Rwp values than human specialists on three samples (2.88% vs. 4.42%, 5.06% vs. 5.40%, 7.60% vs. 9.00%), while on the other two samples its results are very close to those of the specialists. The agent is currently deployed at the China Spallation Neutron Source (CSNS) and is open for external user registration, providing an intelligent and user-friendly analytical tool for materials research. This work fully leverages the cutting-edge advantages of LLM, offers a new path to solve the long-standing problem of automated refinement, takes a key step toward intelligent and fully automated crystallographic analysis, and holds great potential to accelerate AI for Science discoveries in neutron-based materials characterization.

cond-mat.mtrl-sci

Dr.Sai: An agentic AI for real-world physics analysis at BESIII

High Energy Physics (HEP) experiments like BESIII produce petabyte-scale data. Extracting physics results requires complex workflows (simulation, reconstruction, statistical analysis, etc.) that traditionally take experts months or years. Current manual methods are labor-intensive, prone to bias, and limit large-scale systematic scans. As data grows, this paradigm slows discovery. Large Language Models (LLMs) offer a solution. Their natural language understanding and code generation capabilities allow them to interpret scientific tasks and integrate with HEP tools (e.g., ROOT, BOSS) to act as an "AI partner" for autonomous analysis. We present Dr.Sai, an LLM-powered multi-agent system that translates natural language into rigorous physics workflows. As validation, Dr.Sai performed large-scale re-measurements of ten J/psi decay branching fractions - without manual coding. It successfully navigated the real BESIII computing environment and produced results matching established benchmarks. The article details Dr.Sai's architecture, the validation results, and performance evaluation. This work provides a blueprint for autonomous discovery, with relevance to other data-intensive fields like astronomy and genomics.

hep-ex

AI Agents, Language, Deep Learning and the Next Revolution in Science

Modern science is reaching a critical inflection point. Instruments across disciplines, from particle physics and astronomy to genomics and climate modeling, now produce data of such scale, diversity, and interdependence that traditional analytical methods can no longer keep pace. This growing imbalance between data generation and data understanding signals the need for a new scientific paradigm. We propose that intelligent, human-supervised AI agents operating over deep-learning algorithms, represent the next evolution of the scientific method. Built upon large language models and multimodal learning, these agents can interpret scientific intent, design and execute analytical workflows, and ensure traceability through domain-specific languages that preserve human oversight and accountability. Particle physics, a historic incubator of computational innovation, offers the ideal testbed for this transition. At the Institute of High Energy Physics of the Chinese Academy of Sciences, the Dr. Sai system embodies this vision, a multi-agent reasoning framework deployed within collider research at the CEPC. This emerging approach does not replace human scientists but extends their cognitive reach, enabling discovery to scale with complexity and redefining how knowledge itself is produced in the age of intelligent machines. The significance of this paradigm transcends particle physics, offering a blueprint for all data-driven sciences facing the same complexity ceiling.

hep-ex

Xiwu: A Basis Flexible and Learnable LLM for High Energy Physics

Large Language Models (LLMs) are undergoing a period of rapid updates and changes, with state-of-the-art (SOTA) model frequently being replaced. When applying LLMs to a specific scientific field, it's challenging to acquire unique domain knowledge while keeping the model itself advanced. To address this challenge, a sophisticated large language model system named as Xiwu has been developed, allowing you switch between the most advanced foundation models and quickly teach the model domain knowledge. In this work, we will report on the best practices for applying LLMs in the field of high-energy physics (HEP), including: a seed fission technology is proposed and some data collection and cleaning tools are developed to quickly obtain domain AI-Ready dataset; a just-in-time learning system is implemented based on the vector store technology; an on-the-fly fine-tuning system has been developed to facilitate rapid training under a specified foundation model. The results show that Xiwu can smoothly switch between foundation models such as LLaMA, Vicuna, ChatGLM and Grok-1. The trained Xiwu model is significantly outperformed the benchmark model on the HEP knowledge question-and-answering and code generation. This strategy significantly enhances the potential for growth of our model's performance, with the hope of surpassing GPT-4 as it evolves with the development of open-source models. This work provides a customized LLM for the field of HEP, while also offering references for applying LLM to other fields, the corresponding codes are available on Github.

hep-ph