SearcharxivSearch

arXiv subjects

Zelong Zhang

Publications and source records attributed to Zelong Zhang.

7 recordsLinked to original sources

Benchmarking LLM Tool-Use in the Wild

Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inherently wild, being intricate, messy, and flexible. We identify three key challenges from user behaviour: compositional tasks that demand efficient orchestration of tool-call topologies, implicit intent spread across dialogue turns that require contextual inference, and instruction transition, which mixes task queries, clarifications, and casual conversation, forcing LLMs to adjust their policies on the fly. Existing benchmarks overlook these behaviors, making the apparent progress of LLMs on tool-use spurious. To address this, we introduce WildToolBench, an LLM tool-use benchmark grounded in real-world user behavior patterns. Comprehensive evaluations of 57 LLMs reveal that no model achieves an accuracy of more than 15%, indicating a substantial gap in the robustness of LLMs' agentic ability. Controlled experiments and in-depth analyses further indicate that the real challenge for LLM tool-use lies not in artificially complex tasks, but in the wild nature of user behavior, emphasizing the need to reconsider the interactions among LLMs, users, and tools.

cs.HC

Astrophysical neutrino flux measurement and search for tau neutrino induced cascades with 11 years of IceCube data

IceCube measured the diffuse astrophysical neutrino flux for all flavors up to PeV energies. The high energy (TeV-PeV) IceCube cascade sample is particularly effective at selecting electron and tau neutrinos. We present the results of Single Power Law (SPL) and Broken Power Law (BPL) flux measurements based on 11 years of cascade data. From this cascade sample, we study the identification of high energy (~100 TeV) tau neutrinos by detecting a double cascade signature produced by a charged-current neutrino interaction and the subsequent decay of the tau lepton. A Boosted Decision Tree (BDT) is employed to search for the double cascade signature, achieving a significantly improved signal-to-background ratio of 9:1 compared to previous analyses. The selected sample has a tau-neutrino purity of approximately 90% and a weighted mean reconstruction error on the tau decay length of about 4m. We present sensitivities for a maximum likelihood fit of the flavor composition and constraints on the astrophysical neutrino flavor ratios using this sample in combination with IceCube's northern muon neutrino-induced track sample.

astro-ph.HE

$C^3$-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking

Agents based on large language models leverage tools to modify environments, revolutionizing how AI interacts with the physical world. Unlike traditional NLP tasks that rely solely on historical dialogue for responses, these agents must consider more complex factors, such as inter-tool relationships, environmental feedback and previous decisions, when making choices. Current research typically evaluates agents via multi-turn dialogues. However, it overlooks the influence of these critical factors on agent behavior. To bridge this gap, we present an open-source and high-quality benchmark $C^3$-Bench. This benchmark integrates attack concepts and applies univariate analysis to pinpoint key elements affecting agent robustness. In concrete, we design three challenges: navigate complex tool relationships, handle critical hidden information and manage dynamic decision paths. Complementing these challenges, we introduce fine-grained metrics, innovative data collection algorithms and reproducible evaluation methods. Extensive experiments are conducted on 49 mainstream agents, encompassing general fast-thinking, slow-thinking and domain-specific models. We observe that agents have significant shortcomings in handling tool dependencies, long context information dependencies and frequent policy-type switching. In essence, $C^3$-Bench aims to expose model vulnerabilities through these challenges and drive research into the interpretability of agent performance. The benchmark is publicly available at https://github.com/TencentHunyuan/C3-Benchmark.

cs.AI

Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions

Large language models (LLMs) demonstrate strong potential as agents for tool invocation due to their advanced comprehension and planning capabilities. Users increasingly rely on LLM-based agents to solve complex missions through iterative interactions. However, existing benchmarks predominantly access agents in single-mission scenarios, failing to capture real-world complexity. To bridge this gap, we propose the Multi-Mission Tool Bench. In the benchmark, each test case comprises multiple interrelated missions. This design requires agents to dynamically adapt to evolving demands. Moreover, the proposed benchmark explores all possible mission-switching patterns within a fixed mission number. Specifically, we propose a multi-agent data generation framework to construct the benchmark. We also propose a novel method to evaluate the accuracy and efficiency of agent decisions with dynamic decision trees. Experiments on diverse open-source and closed-source LLMs reveal critical factors influencing agent robustness and provide actionable insights to the tool invocation society.

cs.AI

Temperature Effect on Interactions of Oil Droplet with Water-wetted Shale Kerogen at Reservoir Temperatures

Understanding the thermodynamics of the interfacial interactions between oil and kerogen is imperative for recovering hydrocarbon in tight reservoirs, especially in unconventional shale that retains abundant hydrocarbon in kerogen nanopores. The temperature effect on the interactions of light oil with a type II kerogen in water was investigated using molecular dynamics simulation. Non-polar and polar light oil droplets were modeled with clusters of 30 octane molecules and 30 octanethiol molecules, respectively. Kerogen was modeled with a molecular fragment from a type II kerogen. The free energy calculations were performed at constant volume and temperature with umbrella sampling at temperatures in the range of 300--500 K (27--227 °C, 80--440 °F), comparable to the reservoir conditions of common shale plays. The result shows that the free energy of desorption of an oil droplet scales linearly with temperature. For oil droplets, the desorption free energy cannot be quantitatively scaled up from that of a single oil molecule. Additionally, the free energy of desorption exhibited a strong temperature dependence, suggesting a significant entropic contribution to the free energy. The contact angle of oil droplets was estimated by the morphologies of the oil cluster in contact with the kerogen surface, identified at the lowest free energy point in the free energy profile. The cosine of the contact angle is linearly correlated with the free energy of the desorption. This study provides a thermodynamic basis and molecular details on how temperature affects the oil interactions with kerogen, providing a valuable insight to strategy for improving unconventional oil recovery.

cond-mat.mtrl-sci

Measurement of the astrophysical diffuse neutrino flux in a combined fit of IceCube's high energy neutrino data

The IceCube Neutrino Observatory has discovered a diffuse neutrino flux of astrophysical origin and measures its properties in various detection channels. With more than 10 years of data, we use multiple data samples from different detection channels for a combined fit of the diffuse astrophysical neutrino spectrum. This leverages the complementary information of different neutrino event signatures. For the first time, we use a coherent modelling of the signal and background, as well as the detector response and corresponding systematic uncertainties. The detector response is continuously varied during the simulation in order to generate a general purpose Monte Carlo set, which is central to our approach. We present a combined fit yielding a measurement of the diffuse astrophysical neutrino flux properties with unprecedented precision.

astro-ph.HE

A Combined Fit of the Diffuse Neutrino Spectrum using IceCube Muon Tracks and Cascades

The IceCube Neutrino Observatory first observed a diffuse flux of high energy astrophysical neutrinos in 2013. Since then, this observation has been confirmed in multiple detection channels such as high energy starting events, cascades, and through-going muon tracks. Combining these event selections into a high statistics global fit of 10 years of IceCube's neutrino data could strongly improve the understanding of the diffuse astrophysical neutrino flux: challenging or confirming the simple unbroken power-law flux model as well as the astrophysical neutrino flux composition. One key component of such a combined analysis is the consistent modelling of systematic uncertainties of different event selections. This can be achieved using the novel SnowStorm Monte Carlo method which allows constraints to be placed on multiple systematic parameters from a single simulation set. We will report on the status of a new combined analysis of through-going muon tracks and cascades. It is based on a consistent all flavor neutrino signal and background simulation using, for the first time, the SnowStorm method to analyze IceCube's high-energy neutrino data. Estimated sensitivities for the energy spectrum of the diffuse astrophysical neutrino flux will be shown.

astro-ph.HE