SearcharxivSearch

arXiv · 2603.01051

CelloAI Benchmarks: Toward Repeatable Evaluation of AI Assistants

Abstract

Large Language Models (LLM) are increasingly used for software development, yet existing benchmarks for LLM-based coding assistance do not reflect the constraints of High Energy Physics (HEP) and High Performance Computing (HPC) software. Code correctness must respect science constraints and changes must integrate into large, performance-critical codebases with complex dependencies and build systems. The primary contribution of this paper is the development of practical, repeatable benchmarks that quantify LLM performance on HEP/HPC-relevant tasks. We introduce three evaluation tracks -- code documentation benchmarks measure the ability of an LLM to generate Doxygen-style comments, code generation benchmarks evaluate end-to-end usability on representative GPU kernels, and graphical data analysis benchmarks evaluate vision-enabled LLMs. These benchmarks provide a unified framework for measuring progress in scientific coding assistance across documentation quality, code generation robustness, and multimodal validation analysis. By emphasizing repeatability, automated scoring, and domain-relevant failure modes, the suite enables fair comparisons of models and settings while supporting future work on methods that improve reliability for HEP/HPC software development.

Explore related subjects

Keep this discovery

BibTeXRIS

Mohammad Atif, Kriti Chopra, Fang-Ying Tsai, Ozgur O. Kilic, Tianle Wang, Zhihua Dong, Douglas Benjamin, Charles Leggett, Meifeng Lin, Paolo Calafiura, Salman Habib. 2026-03-01. CelloAI Benchmarks: Toward Repeatable Evaluation of AI Assistants. https://arxiv.org/abs/2603.01051

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Production of Light Nuclei and Hypernuclei in Heavy-Ion Collisions

We review recent STAR and ALICE measurements of light-nucleus and hypernucleus yields, femtoscopic correlations, and collective flow presented at SQM 2026. Statistical-hadronization calculations provide a useful baseline for integrated yields but do not simultaneously describe all measured light-nucleus ratios across collision energies and system sizes. For bound states with mass number $A<4$, current coalescence calculations provide a broadly consistent description of yields, femtoscopic correlations, and collective flow, although the quantitative hypertriton comparison depends on the assumed few-body wave function. The suppressed production of resonant $^{4}$Li relative to compact $^{4}$He indicates an effect of nuclear structure and late-stage dynamics. However, the quantitative model comparison also depends on the treatment of feed-down from unstable states. In high-multiplicity $p$+$p$ collisions, pion-deuteron femtoscopy further indicates that most observed (anti)deuterons are formed through nucleon fusion after strong decays of short-lived resonances. Taken together, these measurements show that production chronology and internal nuclear structure leave measurable imprints on the physics observables.

hep-ex

Search for the process $e^+e^-\to f_1(1285)$ at the SND detector

In the experiment with the SND detector at the VEPP-2000 $e^+e^-$ collider, a search is performed for the direct production of the $C$-even $f_1(1285)$ resonance in $e^+e^-$ collisions. The analysis is based on data with an integrated luminosity of about 200 pb$^{-1}$, accumulated in the center-of-mass energy range of 1.14--1.46 GeV, of which about 72 pb$^{-1}$ were recorded near the maximum of the $f_1(1285)$ resonance. The $f_1(1285)$ production cross section at the resonance maximum $\sigma(e^+e^-\to f_1)=(31\pm 13\pm 2)$ pb and the branching fraction $B(f_1(1285)\to e^+e^-)=(3.5\pm 1.4\pm 0.3)\times 10^{-9}$ have been measured. The significance of the observation of the $e^+e^-\to f_1(1285)$ process is $2.5\sigma$. Since the significance is low, we also present the upper limits at the 90% confidence level: $\sigma(e^+e^-\to f_1)<48\mbox{ pb}$ and $B(f_1(1285)\to e^+e^-)<5.4\times 10^{-9}$.

hep-ex

Projected Sensitivity to Slow Muonphilic Dark Matter with Accelerator Muon Beams

The nature of dark matter (DM) remains one of the most enduring open questions in modern physics, and muonphilic DM has emerged as a promising scenario that complements traditional DM candidates. Following the recently established cosmic-ray muon scattering approach, we investigate the sensitivity for probing slow muonphilic DM with accelerator muon beams. A Geant4-based simulation framework is developed, incorporating the detector geometry from the PKMu muon tomography system and a dedicated elastic $\mu$-DM scattering process. The projected sensitivity is found to be largely insensitive to both the beam energy and the transverse beam size when the beam is fully contained within the detector acceptance. For a benchmark beam intensity of $10^5/\rm{s}$, the simulated pure-muon beam surpasses the existing cosmic-ray limit of $1.61\times10^{-17}$ cm$^2$ at $m_{\rm DM}=1$ GeV within approximately 11 seconds. A realistic muon beam phase-space distribution based on simulations for the High Intensity heavy-ion Accelerator Facility (HIAF) is also implemented, yielding projected limits that improve upon the cosmic-ray results by nearly two orders of magnitude in a one-day exposure. These results demonstrate that a beam-muon scattering experiment offers a robust and promising route toward significantly improved sensitivity to slow muonphilic DM.

hep-ex