SearcharxivSearch

arXiv subjects

Zelong Huang

Publications and source records attributed to Zelong Huang.

5 recordsLinked to original sources

LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling

Looped Transformers scale latent computation by repeatedly applying shared blocks, but sequential looping increases latency and KV-cache memory with the loop count. Parallel loop Transformers (PLT) alleviate this cost through cross-loop position offsets (CLP) and shared-KV gated sliding-window attention, making loop count a practical design choice. We therefore study PLT loop-count selection through a gain--cost view: an extra loop may refine representations, but CLP also introduces a positional mismatch at each loop boundary. We instantiate this study by training LoopCoder-v2, a family of 7B PLT coders with different loop counts, from scratch on 18T tokens, followed by matched instruction tuning and evaluation. Empirically, the two-loop variant delivers broad gains over the non-looped baseline across code generation, code reasoning, agentic software engineering, and tool-use benchmarks, improving SWE-bench Verified from 43.0 to 64.4 points and Multi-SWE from 14.0 to 31.0 points. In contrast, variants with three or more loops regress, revealing a strongly non-monotonic loop-count effect. Our diagnostics show that loop 2 provides the main productive refinement, while later loops yield diminishing, oscillatory updates and reduced representational diversity. Because the CLP-induced mismatch remains roughly fixed as refinement gains shrink, the offset cost increasingly dominates. This gain--cost trade-off explains PLT's saturation at two loops and provides diagnostics for loop-count selection.

cs.LG

IQuest-Coder-V1 Technical Report

In this report, we introduce the IQuest-Coder-V1 series-(7B/14B/40B/40B-Loop), a new family of code large language models (LLMs). Moving beyond static code representations, we propose the code-flow multi-stage training paradigm, which captures the dynamic evolution of software logic through different phases of the pipeline. Our models are developed through the evolutionary pipeline, starting with the initial pre-training consisting of code facts, repository, and completion data. Following that, we implement a specialized mid-training stage that integrates reasoning and agentic trajectories in 32k-context and repository-scale in 128k-context to forge deep logical foundations. The models are then finalized with post-training of specialized coding capabilities, which is bifurcated into two specialized paths: the thinking path (utilizing reasoning-driven RL) and the instruct path (optimized for general assistance). IQuest-Coder-V1 achieves state-of-the-art performance among competitive models across critical dimensions of code intelligence: agentic software engineering, competitive programming, and complex tool use. To address deployment constraints, the IQuest-Coder-V1-Loop variant introduces a recurrent mechanism designed to optimize the trade-off between model capacity and deployment footprint, offering an architecturally enhanced path for efficacy-efficiency trade-off. We believe the release of the IQuest-Coder-V1 series, including the complete white-box chain of checkpoints from pre-training bases to the final thinking and instruction models, will advance research in autonomous code intelligence and real-world agentic systems.

cs.AI

Close the Loop: Synthesizing Infinite Tool-Use Data via Multi-Agent Role-Playing

Enabling Large Language Models (LLMs) to reliably invoke external tools remains a critical bottleneck for autonomous agents. Existing approaches suffer from three fundamental challenges: expensive human annotation for high-quality trajectories, poor generalization to unseen tools, and quality ceilings inherent in single-model synthesis that perpetuate biases and coverage gaps. We introduce InfTool, a fully autonomous framework that breaks these barriers through self-evolving multi-agent synthesis. Given only raw API specifications, InfTool orchestrates three collaborative agents (User Simulator, Tool-Calling Assistant, and MCP Server) to generate diverse, verified trajectories spanning single-turn calls to complex multi-step workflows. The framework establishes a closed loop: synthesized data trains the model via Group Relative Policy Optimization (GRPO) with gated rewards, the improved model generates higher-quality data targeting capability gaps, and this cycle iterates without human intervention. Experiments on the Berkeley Function-Calling Leaderboard (BFCL) demonstrate that InfTool transforms a base 32B model from 19.8% to 70.9% accuracy (+258%), surpassing models 10x larger and rivaling Claude-Opus, and entirely from synthetic data without human annotation.

cs.CL

Comprehensive Investigation of Fundamental Mode Profiles in Monolithic Nonplanar Ring Oscillators

Nonplanar ring oscillators (NPROs) are building blocks for high-performance single-frequency lasers and ring-laser gyroscopes that have profoundly improved the state-of-the-art laser technologies, fundamental research and precision measurements. However, a comprehensive investigation of fundamental mode profiles in monolithic NPROs has been a missing part even though they will affect the performance of the lasers or ring-laser gyroscopes. Here, we present a comprehensive finite-element modeling of the output beam profiles of monolithic NPROs by combining ABCD transmission matrix and generalized Huygens-Fresnel integral. We theoretically investigate the effects of geometric parameters of monolithic NPROs on their output mode profiles. In particular, we focus on the thermal effect inside the monolithic NPRO and calculate the equivalent focal-length of the thermal-lens by using a ray-tracing-method. Furthermore, we experimentally characterize the output laser beam profile, reconstruct the beam profile at the A-facet of the monolithic NPRO, and compare the experimental results with the simulation, thereby validating the accuracy and reliability of our model. The investigation may facilitate monolithic NPRO design and subsequently improve the performances of NPRO lasers and gyroscopes in the future.

physics.optics

Sub-kHz single-frequency pulsed semiconductor laser based on NPRO injection locking

We report a single-frequency, narrow-linewidth semiconductor pulsed laser based on pump current modulation and optical injection locking technique. A monolithic non-planar ring oscillator laser is employed as the seed source to guarantee the single-frequency narrow-linewidth performance. Simultaneously, pulse operation is achieved by directly modulating the pump current of the semiconductor laser. The single-frequency pulsed laser (SFPL) has achieved a pulse repetition rate of 50 kHz-1 MHz, a pulse duration ranging from 120 ns to a quasi-continuous state, and a peak power of 160 mW. Moreover, the SFPL has reached a pulsed laser linewidth as narrow as 905 Hz, optical spectrum signal-to-noise ratio of better than 65 dB at a center wavelength of 1064.45 nm. Such extremely narrow-linewidth, repetition-rate and pulse-width tunable SFPL has great potential for applications in coherent LIDAR, metrology, remote sensing, and nonlinear frequency conversion.

physics.optics