SearcharxivSearch

arXiv subjects

Xiaobo Xue

Publications and source records attributed to Xiaobo Xue.

7 recordsLinked to original sources

Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls

Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool and orchestrating the order of calls. However, correctly filling the parameters of a tool call is equally critical for successful execution and has received far less attention. In domains such as cloud networking, even frontier models correctly complete fewer than half of tool calls. Inspired by recent analyses showing that LLM hidden states encode rich information about model predictions, we discover that while the model generates a parameter value, its hidden state contains a strong correctness signal: a simple linear probe can accurately predict whether the value will be correct. Based on this observation, we propose a unified probe-guided framework with two complementary approaches: probe-filtered bootstrapped training (PBT), which uses the probe to filter reliable self-generated calls for fine-tuning, and probe-guided reranking (PGR), which uses the probe to select better candidates during inference. To support systematic evaluation, we release ParamBench, a benchmark built from real cloud-network APIs that categorizes every instance into five difficulty levels according to parameter nesting depth, cross-parameter dependencies, and the reasoning required to derive values from earlier calls. Extensive experiments across 5 open models on ParamBench and 6 external benchmarks demonstrate that our method substantially improves parameter generation, raising the average exact match from 19.7% to 59.6%.

cs.AI

Scalable LLM Agent Tool Access in the Cloud

LLM agents increasingly rely on tool calling to act on external systems, and the Model Context Protocol (MCP) has quickly become its de facto interface. Operating MCP at cloud scale, however, becomes difficult. On the tool provider side, legacy services are not directly callable through MCP; the rapid protocol development also creates ongoing compatibility cost. On the agent side, the number of accessible tool is limited by the LLM context window and inference overhead; mounting a large tool set increases token usage and inference latency and can reduce task success rate. Moreover, for stateful MCP backends with multiple replicas, preserving session affinity increases client-side complexity. We present a cloud-scale gateway system for MCP service. It breaks the direct-connect model on the data plane and offloads legacy service integration, consolidating incompatible MCP variants, access control, tool recommendation, and session-aware routing to the gateway. Hybrid retrieval sustains 98% Top-15 recall; it scales agent tool access to 3,000+ with high tool selection accuracy, and reduces tool selection time by $8.9\times$ and token usage by $23.8\times$, with low per-call overhead, stable under scale-out. Finally, we share the lessons learned from deploying the gateway system in production.

cs.DC

AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts

Large language model agents repeatedly encounter related tasks, yet systems that learn from trajectories commit every lesson to one predefined artifact form. A local constraint, a reusable procedure, and a delegated objective require different amounts of runtime ownership, so one form either under-specifies the correction or wraps it in execution machinery it does not need. We present AutoRefine, which treats trajectory learning as typed artifact compilation. It contrasts failed against successful trajectories to derive a type-neutral, evidence-linked intervention specification, then compiles that specification into the first Rule, Skill, or bounded Subagent that closes it under a runtime-relative ownership order: the selected schema must own every specified observation, state variable, dependent decision, and completion condition. Validation is stated in the same terms: a type-specific contract gate tests whether the generated object realizes its declared boundary, and a replay gate admits it only when it improves the correction cases linked to its source failures without regression on preservation cases. With GPT-5.6-terra as the shared backbone, AutoRefine records the highest success on ALFWorld, ScienceWorld, TravelPlanner, and SpreadsheetBench, and ties the best result on SkillCraft; on TravelPlanner it reaches 80.56% success against 50.0% for the strongest baseline. Removing boundary closure or replay validation costs 15.00 and 16.11 percentage points, the two largest losses among our construction and admission policies. In a longitudinal TravelPlanner stream, the repository holds 89--91% held-out success after 60 learning tasks with no net loss on previously solved tasks, and frozen repositories improve all 25 evaluated source--target pairs, more within a domain (14.20 points on average) than across domains (6.99).

cs.AI

Blueprint First, Model Second: A Framework for Deterministic LLM Workflow

While powerful, the inherent non-determinism of large language model (LLM) agents limits their application in structured operational environments where procedural fidelity and predictable execution are strict requirements. This limitation stems from current architectures that conflate probabilistic, high-level planning with low-level action execution within a single generative process. To address this, we introduce the \textsc{Source Code Agent} framework, a new paradigm built on the ``Blueprint First, Model Second'' philosophy that decouples workflow logic from the generative model. An expert-defined operational procedure is first codified into a source code-based Execution Blueprint, which is then executed by a deterministic engine. The LLM is strategically invoked as a specialized tool to handle bounded, complex sub-tasks within the workflow, but never to decide the workflow's path. We evaluate on the TravelPlanner benchmark for constraint-aware travel planning. The \textsc{Source Code Agent} achieves a 35.56\% final pass rate, a 97.6\% improvement over the state-of-the-art ATLAS baseline (18.00\%) on the same Claude-Sonnet-4 backbone. Critically, it reduces constraint violations by 96.0\% (11 vs 275) while improving execution efficiency by 27.1\% (10.2$\pm$0.7 steps vs 14.0). Two production incident-diagnosis deployments and additional results on ScienceWorld and ALFWorld confirm that the architecture transfers beyond travel planning to procedurally well-defined, constraint-intensive workflows. Our work enables the verifiable and reliable deployment of autonomous agents in applications governed by strict procedural logic.

cs.SE

Prospects for 10^{-18} Instability Laser Referenced on Thermal Atomic Ensembles

A thermal atomic ensemble-based laser source with superior frequency stability is proposed that relies on the accumulated contributions from an abundance of nonzero-transverse-velocity atomic ensembles. Compared with the traditional case in which only atoms with near-zero transverse velocities are utilized, the amplitude of the optical Ramsey fringes for a thermal calcium beam can be dramatically enhanced by three orders of magnitude or more, thus, the signal-to-noise ratio can be improved 33-fold. Based on the recent results of atomic interferometry-based laser stabilization, a quantum projection noise-limited frequency instability less than 2E-17/tau^0.5 is feasible. Such an ultrastable laser has promising applications in diverse areas, including metrology and astronomy.

physics.atom-ph

Ultralong Faraday laser as an optical frequency standard

In this letter, we introduce the concept and experimentally demonstrate an ultralong Faraday laser as an optical frequency standard in principle. The ultralong Faraday laser is based on the Faraday anomalous dispersion optical filter (FADOF) with ultra-narrow bandwidth and the ultralong fiber extended cavity of $800$ m. The ultra-narrow FADOF is based on atomic transition line of isotope $^{87}$Rb, which has an ultra-narrow bandwidth of $26.0$ MHz and a transmission of $23.6\%$ at $780$ nm. Fibers of length $150$ m and $800$ m are used as ultralong fiber extended cavities, which provide optical feedback and give extremely small FSR of $0.667$ MHz and $0.125$ MHz, respectively. The mechanism of the proposed ultralong Faraday laser is to combine FADOF's ultra-narrow bandwidth and ultralong cavity's small free spectral range to limit the lasing frequency within FADOF bandwidth covered by the semiconductor gain. The active lasing frequency of the ultralong Faraday laser is determined by the center frequency of FADOF transmission, which is corresponding to atomic transition $5^{2}S_{1/2},\ F = 2 \ \rightarrow\ 5^{2}P_{3/2},\ F^{\prime} = 2,\ 3$ of isotope $^{87}$Rb. The Allan deviation of the fractional frequency of the ultralong Faraday laser output signal in 0.1 s -- 1 s measuring time is around $1\times10^{-11}$.

physics.atom-ph

Hanle detection for optical clocks

Considering the strong inhomogeneous spatial polarization and intensity distribution of spontaneous decay fluorescence due to the Hanle effect, we propose and demonstrate a universe Hanle detection configuration of electron-shelving method for optical clocks. Experimental results from Ca atomic beam optical frequency standard with 423 nm electron-shelving method show that a designed Hanle detection geometry with optimized magnetic field direction, detection laser beam propagation and polarization direction, and detector position can improve the fluorescence collection rate by more than one order of magnitude comparing with that of inefficient geometry. With the fixed 423 nm fluorescence, the improved 657 nm optical frequency standard signal intensity is presented. And the potential application of the Hanle detection geometry designed for facilitating the fluorescence collection for optical lattice clock with a limited solid angle of the fluorescence collection has been discussed. This Hanle detection configuration is also effective for ion detection in ion optical clock and quantum information experiments. Besides, a cylinder fluorescence collection structure is designed to increase the solid angle of the fluorescence collection in Ca atomic beam optical frequency standard.

physics.atom-ph