Searcharxiv⌕ Search

arXiv subjects

Zhi-Yuan Chen

Publications and source records attributed to Zhi-Yuan Chen.

5 recordsLinked to original sources

AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents

While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions. Unlike mathematical reasoning where errors are often rectifiable via backtracking, tool-use failures frequently induce irreversible side effects, making accurate step-level verification critical. However, existing process-level benchmarks are predominantly confined to closed-world mathematical domains, failing to capture the dynamic and open-ended nature of tool execution. To bridge this gap, we introduce AgentProcessBench, the first benchmark dedicated to evaluating step-level effectiveness in realistic, tool-augmented trajectories. The benchmark comprises 1,000 diverse trajectories and 8,509 human-labeled step annotations with 89.1% inter-annotator agreement. It features a ternary labeling scheme to capture exploration and an error propagation rule to reduce labeling ambiguity. Extensive experiments reveal key insights: (1) weaker policy models exhibit inflated ratios of correct steps due to early termination; (2) distinguishing neutral and erroneous actions remains a significant challenge for current models; and (3) process-derived signals provide complementary value to outcome supervision, significantly enhancing test-time scaling. We hope AgentProcessBench can foster future research in reward models and pave the way toward general agents. The code and data are available at https://github.com/RUCBM/AgentProcessBench.

cs.AI↗

$P_{c\bar cs}(4459)^{0}$, $P_{c\bar c s}(4338)^0$ and mass spectrum of strange hidden-charm pentaquarks

Strange hidden-charm pentaquark states have been systematically investigated within a diquark-triquark model. Through a Gaussian expansion method, masses of some diquarks, triquarks and strange hidden-charmed pentaquark states from S-wave to P-wave excitations have been calculated with the non-relativistic Semay and Silvestre-Brac potentials in terms of the same parameters employed for tetraquark states. Masses of pentaquark states in S-wave excitations are found between $4200$ MeV and $4590$ MeV, while masses of all P-wave excitations are found above $4600$ MeV. Mass splittings between the S-wave and P-wave pentaquark states are about $350-570$ MeV. In comparison to the experimental data, $P_{c\bar cs}(4459)^{0}$ observed by LHCb in decay channel $Ξ_{b}^{-}\rightarrow J/ψΛK^-$ is assumed as the $|1; 0, 1/2; 3/2, 0\rangle_{3/2}$ $[sq][\bar{c}cq]$ pentaquark state with $J^P={3\over 2}^-$, while $P_{c\bar c s}(4338)^0$ observed in the decay channel $B^{-}\rightarrow J/ψΛ\bar{p}$ is very possibly the $|0; 1, 1/2; 1/2, 0\rangle_{1/2}$ $[cq][\bar{c}sq]$ pentaquark state with $J^P={1\over 2}^-$. We predict a lowest strange hidden-charm pentaquark state with $J^P={1\over 2}^-$ around $4200$ MeV.

hep-ph↗

Spectrum of $[cq][\bar{s}\bar{q}]$ tetraquarks: Nature of $D^*_{s0}(2317)$, $D_{s1}(2460)$ and $T^*_{c\bar s0}(2900)$

Motivated by the recent observations of exotic open-charm tetraquark candidates \(T^a_{c\bar{s}0}(2900)^{++}\) and \(T^a_{c\bar{s}0}(2900)^{0}\), we systematically calculate the mass spectra of \([cq][\bar{s}\bar{q}]\) tetraquarks within a nonrelativistic constituent quark potential model. In the model, the tetraquark states are treated as diquark-antidiquark bound systems with an interior interaction similar to the quark-antiquark interaction in conventional mesons. The well established states \(D_{s0}^*(2317)\) with \(J^P=0^+\) and \(D_{s1}(2460)\) with \(J^P=1^+\) could be identified as the two ground states of the \([cq][\bar{s}\bar{q}]\) system. \(T^a_{c\bar{s}0}(2900)^{0}\) and \(T^a_{c\bar{s}0}(2900)^{++}\) could be naturally interpreted as radially excited \(0^+\) tetraquark states with different interior components. Their large mass difference may result from their different interior structure instead of an isospin symmetry breaking. Whether \(T^a_{c\bar{s}0}(2900)^{0}\) and \(T^a_{c\bar{s}0}(2900)^{++}\) belong to an isospin triplet deserves further experimental investigation. In addition, there may be another \(0^+\) \([cq][\bar{s}\bar{q}]\) tetraquark state with mass around $2450$ MeV, which is composed of a $cq$ diquark and a $\bar s\bar q$ antidiquark both with spin-0. In the energy region $2640-2700$ MeV, there may be a $J^P=2^+$ \([cq][\bar{s}\bar{q}]\) tetraquark state composed of the $cq$ diquark and the $\bar s\bar q$ antidiquark both with spin-1.

hep-ph↗

Beyond the Surface: Measuring Self-Preference in LLM Judgments

Recent studies show that large language models (LLMs) exhibit self-preference bias when serving as judges, meaning they tend to favor their own responses over those generated by other models. Existing methods typically measure this bias by calculating the difference between the scores a judge model assigns to its own responses and those it assigns to responses from other models. However, this approach conflates self-preference bias with response quality, as higher-quality responses from the judge model may also lead to positive score differences, even in the absence of bias. To address this issue, we introduce gold judgments as proxies for the actual quality of responses and propose the DBG score, which measures self-preference bias as the difference between the scores assigned by the judge model to its own responses and the corresponding gold judgments. Since gold judgments reflect true response quality, the DBG score mitigates the confounding effect of response quality on bias measurement. Using the DBG score, we conduct comprehensive experiments to assess self-preference bias across LLMs of varying versions, sizes, and reasoning abilities. Additionally, we investigate two factors that influence and help alleviate self-preference bias: response text style and the post-training data of judge models. Finally, we explore potential underlying mechanisms of self-preference bias from an attention-based perspective. Our code and data are available at https://github.com/zhiyuanc2001/self-preference.

cs.CL↗

Large Language Model-based Human-Agent Collaboration for Complex Task Solving

In recent developments within the research community, the integration of Large Language Models (LLMs) in creating fully autonomous agents has garnered significant interest. Despite this, LLM-based agents frequently demonstrate notable shortcomings in adjusting to dynamic environments and fully grasping human needs. In this work, we introduce the problem of LLM-based human-agent collaboration for complex task-solving, exploring their synergistic potential. In addition, we propose a Reinforcement Learning-based Human-Agent Collaboration method, ReHAC. This approach includes a policy model designed to determine the most opportune stages for human intervention within the task-solving process. We construct a human-agent collaboration dataset to train this policy model in an offline reinforcement learning environment. Our validation tests confirm the model's effectiveness. The results demonstrate that the synergistic efforts of humans and LLM-based agents significantly improve performance in complex tasks, primarily through well-planned, limited human intervention. Datasets and code are available at: https://github.com/XueyangFeng/ReHAC.

cs.CL↗