SearcharxivSearch

arXiv subjects

Chao Song

Publications and source records attributed to Chao Song.

At least 19 recordsLinked to original sources

DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation

On-policy distillation (OPD) trains student models on their own rollouts to reduce exposure bias. However, in multi-turn agent scenarios, early student errors can lead a trajectory away from the teacher's familiar domain. Existing curriculum learning methods regulate how much teacher support is used according to training progress, but cannot determine when it is needed. In light of this, we propose DASH-OPD, Discrepancy-Aware Switching with Hysteresis for OPD, the first agentic OPD method that can switch executors adaptively and bidirectionally. On each turn, DASH-OPD calculates a mean log-probability ratio between the two executors over action tokens as their discrepancy. Student-to-teacher ratios on student turns form drift signals, while teacher-to-student ratios on teacher turns form recovery signals. These signals are normalized and accumulated over multiple turns into drift and recovery evidence. DASH-OPD switches executors when the evidence exceeds its corresponding switching threshold. This multi-turn accumulation makes the switching hysteretic, preventing high-frequency switches caused by transient fluctuations. Across WebShop, ALFWorld, and ScienceWorld at two student-model scales, DASH-OPD outperforms five baselines in all 14 task-performance comparisons while yielding the shortest trajectories in nine of ten turn-count comparisons, offering the strongest overall performance-efficiency trade-off. This paper is a work in progress. Code, training logs, and model checkpoints will be released later.

cs.LG

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning

Experience-driven self-evolution is critical for large language model (LLM) agents to improve through open-world interaction. However, existing experience learning methods mostly rely on single-agent loops, where the same agent executes tasks, summarizes outcomes, and determines memory content. This setup makes agents vulnerable to the Self-Confirmation Trap: wrong-but-self-consistent trajectories are misidentified as successful experience, leading to cumulative errors during retrieval and reuse. To address this issue, we propose EDV, an Execute-Distill-Verify framework for reliable experience learning. In the Execute stage, multiple heterogeneous agents explore the same task space in parallel to generate diverse candidate trajectories. In the Distill stage, a dedicated third-party agent comparatively analyzes these trajectories to produce candidate experiences, reducing executor-centric summarization bias. In the Verify stage, the execution group validates candidates via a consensus mechanism, and only approved experiences are written into shared or private memory. By decoupling the three stages, EDV transforms experience learning from isolated self-reflection into collaborative construction, filtering erroneous and noisy content before memory insertion. We evaluate EDV on three challenging long-horizon benchmarks: tau2-bench, Mind2Web and MMTB. Results show EDV consistently outperforms strong baselines, validating that reliable experience construction is essential for robust agent self-evolution. Our code is available at https://github.com/shidingz/EDV.

cs.CL

Tensor Train Decomposition-based 3D Implicit Full Waveform Inversion with Multi-scale Structural Similarity

Three-dimensional full waveform inversion (3DFWI) is a powerful technique for reconstructing high-resolution subsurface velocity models. However, its application is often limited by high memory requirements, computational costs, and sensitivity to cycle skipping. To overcome these challenges, we propose a novel tensor train (TT) decomposition-based 3D implicit full waveform inversion framework (TT-3DIFWI) combined with a multi-scale structural similarity (M-SSIM) objective function. In this framework, the 3D velocity model is represented by TT decomposition as a product of a series of low-rank core tensors. Then, three axis-specific implicit neural network representations (INR) based on one-dimensional vector coordinates as input are constructed to predict these core tensors, rather than directly predicting the velocity model. This INR reparameterization method based on TT decomposition can significantly reduce the memory consumption of INR training while maintaining the accuracy and resolution of the 3D velocity model reconstruction. Meanwhile, the low-rank structure of TT decomposition also ensures the structural consistency of the reconstruction velocity, thereby improving the accuracy and continuity of the inversion result. Furthermore, the M-SSIM objective function can compare the multi-scale structural differences between predicted and observed data, and utilize the ultra-low frequency features to reduce cycle skipping. Numerical experiments on synthetic and challenging land datasets demonstrate that TT-3DIFWI with M-SSIM achieves accurate and continuous velocity reconstruction, even with poor initial models or missing low-frequency data.

physics.geo-ph

A superconducting surface-code processor with lattice-surgery logical operations

Fault-tolerant logical operations are fundamental for scalable quantum computation. Here, we report the experimental realization of lattice-surgery operations between a pair of distance-three surface-code logical qubits on a planar superconducting processor. During repeated syndrome extraction cycles, the logical qubits exhibit per-cycle error rates of $0.0365(2)$ and $0.0282(1)$, respectively, after leakage events are rejected. By leveraging joint initialization and lattice splitting, we deterministically prepare a logical Bell state, confirming genuine bipartite entanglement via the error-corrected logical state fidelity. We further execute a two-qubit Deutsch-Jozsa algorithm at the logical level to demonstrate algorithmic utility in a fault-tolerant framework. Finally, to achieve universal control, we implement magic-state injection and gate teleportation to realize continuous non-Clifford rotations about the logical $X$ axis. For the logical $R_{X}(\pi/4)$ gate, we achieve a logical gate fidelity of $0.943_{-9}^{+10}$ conditioned on the absence of detected errors. These results establish lattice surgery as a practical and versatile paradigm for logical computation in near-term surface-code architectures, representing a critical milestone toward scalable fault-tolerant quantum advantage in superconducting circuits.

quant-ph

Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe

Reinforcement Learning (RL) is essential for evolving Large Language Models (LLMs) into autonomous agents capable of long-horizon planning, yet a practical recipe for scaling RL in complex, multi-turn environments remains elusive. This paper presents a systematic empirical study using TravelPlanner, a challenging testbed requiring tool orchestration to satisfy multifaceted constraints. We decompose the agentic RL design space along 5 axes: reward shaping, model scaling, data composition, algorithm selection, and environmental stability. Our controlled experiments yield 7 key takeaways, e.g., (1) reward and algorithm choices are scale-dependent as smaller models benefit from staged rewards and enhanced exploration, whereas larger models converge efficiently with simpler dense rewards, (2) ~ 1K training samples with a balanced difficulty mixture mark a sweet spot for both in-domain and out-of-domain performance, and (3) environmental stability is critical to prevent policy degradation. Based on our distilled recipe, our RL-trained models achieve state-of-the-art performance on TravelPlanner, significantly outperforming leading LLMs.

cs.LG

VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning

Large models are increasingly becoming autonomous agents that interact with real-world environments and use external tools to augment their static capabilities. However, most recent progress has focused on text-only large language models, which are limited to a single modality and therefore have narrower application scenarios. On the other hand, multimodal large models, while offering stronger perceptual capabilities, remain limited to static knowledge and lack the ability to access and leverage up-to-date web information. In this paper, we propose VSearcher, turning static multimodal model into multimodal search agent capable of long-horizon, multi-turn tool use in real-world web environments, including text search, image search, and web browsing, via reinforcement learning. Specifically, we introduce Iterative Injection Data Synthesis pipeline to generate large-scale, complex multimodal QA questions, which are further filtered with comprehensive metrics to ensure high quality and sufficient difficulty. We then adopt an SFT-then-RL training pipeline to turn base multimodal models to agent capable of multi-turn tool calling in real-world web environments. Besides, we propose a multimodal search benchmark MM-SearchExam dedicated to evaluating search capabilities of multimodal search agents, which proves highly challenging for recent proprietary models. Extensive evaluations across multiple multimodal search benchmarks reveal effectiveness of our method. VSearcher achieves superior performance compared to recent multimodal search agents and even surpasses several proprietary models on multimodal web search tasks.

cs.CV

Teleportation transition of surface codes on a superconducting quantum processor

The topological surface code is a leading candidate for harnessing long-range entanglement to protect logical quantum information against errors, and teleportation of logical states is desirable for robust quantum information processing. Nevertheless, scaling up the surface code in quantum teleportation poses a formidable challenge to experiment. Here on a superconducting quantum processor with 125 qubits, we demonstrate the robust teleportation of topological rotated surface code prepared by a linear-depth unitary circuit, with code distances up to 7. We obtain the teleportation phase diagram by tuning the local entangling gates uniformly across a finite threshold. Furthermore, we show that the entangling threshold can be boosted by coherent qubit rotations that inject magic resources beyond the Clifford regime, restoring the duality symmetry of the topological phase, which serves as a guiding principle to minimize the entanglement resource. Our results shed light on simulating and leveraging topological quantum matter on quantum devices, and pave the way to the ultimate goal of distributed fault tolerant quantum computation.

quant-ph

EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone Generation

Designing enzyme backbones with substrate-specific functionality is a critical challenge in computational protein engineering. Current generative models excel in protein design but face limitations in binding data, substrate-specific control, and flexibility for de novo enzyme backbone generation. To address this, we introduce EnzyBind, a dataset with 11,100 experimentally validated enzyme-substrate pairs specifically curated from PDBbind. Building on this, we propose EnzyControl, a method that enables functional and substrate-specific control in enzyme backbone generation. Our approach generates enzyme backbones conditioned on MSA-annotated catalytic sites and their corresponding substrates, which are automatically extracted from curated enzyme-substrate data. At the core of EnzyControl is EnzyAdapter, a lightweight, modular component integrated into a pretrained motif-scaffolding model, allowing it to become substrate-aware. A two-stage training paradigm further refines the model's ability to generate accurate and functional enzyme structures. Experiments show that our EnzyControl achieves the best performance across structural and functional metrics on EnzyBind and EnzyBench benchmarks, with particularly notable improvements of 13\% in designability and 13\% in catalytic efficiency compared to the baseline models. The code is released at https://github.com/Vecteur-libre/EnzyControl.

q-bio.BM

Fock space prethermalization and time-crystalline order on a quantum processor

Periodically driven quantum many-body systems exhibit a wide variety of exotic nonequilibrium phenomena and provide a promising pathway for quantum applications. A fundamental challenge for stabilizing and harnessing these highly entangled states of matter is system heating by energy absorption from the drive. Here, we propose and demonstrate a disorder-free mechanism, dubbed Fock space prethermalization (FSP), to suppress heating. This mechanism divides the Fock-space network into linearly many sparse sub-networks, thereby prolonging the thermalization timescale even for initial states at high energy densities. Using 72 superconducting qubits, we observe an FSP-based time-crystalline order that persists over 120 cycles for generic initial Fock states. The underlying kinetic constraint of approximately conserved domain wall (DW) numbers is identified by measuring site-resolved correlators. Further, we perform finite-size scaling analysis for DW and Fock-space dynamics by varying system sizes, which reveals size-independent regimes for FSP-thermalization crossover and links the dynamical behaviors to the eigenstructure of the Floquet unitary. Our work establishes FSP as a robust mechanism for breaking ergodicity, and paves the way for exploring novel nonequilibrium quantum matter and its applications.

quant-ph

ACPO: Adaptive Curriculum Policy Optimization for Aligning Vision-Language Models in Complex Reasoning

Aligning large-scale vision-language models (VLMs) for complex reasoning via reinforcement learning is often hampered by the limitations of existing policy optimization algorithms, such as static training schedules and the rigid, uniform clipping mechanism in Proximal Policy Optimization (PPO). In this work, we introduce Adaptive Curriculum Policy Optimization (ACPO), a novel framework that addresses these challenges through a dual-component adaptive learning strategy. First, ACPO employs a dynamic curriculum that orchestrates a principled transition from a stable, near on-policy exploration phase to an efficient, off-policy exploitation phase by progressively increasing sample reuse. Second, we propose an Advantage-Aware Adaptive Clipping (AAAC) mechanism that replaces the fixed clipping hyperparameter with dynamic, sample-wise bounds modulated by the normalized advantage of each token. This allows for more granular and robust policy updates, enabling larger gradients for high-potential samples while safeguarding against destructive ones. We conduct extensive experiments on a suite of challenging multimodal reasoning benchmarks, including MathVista, LogicVista, and MMMU-Pro. Results demonstrate that ACPO consistently outperforms strong baselines such as DAPO and PAPO, achieving state-of-the-art performance, accelerated convergence, and superior training stability.

cs.AI

Combinatorial optimization enhanced by shallow quantum circuits with 104 superconducting qubits

A pivotal task for quantum computing is to speed up solving problems that are both classically intractable and practically valuable. Among these, combinatorial optimization problems have attracted tremendous attention due to their broad applicability and natural fitness to Ising Hamiltonians. Here we propose a quantum sampling strategy, based on which we design an algorithm for accelerating solving the ground states of Ising model, a class of NP-hard problems in combinatorial optimization. The algorithm employs a hybrid quantum-classical workflow, with a shallow-circuit quantum sampling subroutine dedicated to navigating the energy landscape. Using up to 104 superconducting qubits, we demonstrate that this algorithm outputs favorable solutions against even a highly-optimized classical simulated annealing (SA) algorithm. Furthermore, we illustrate the path toward quantum speedup based on the time-to-solution metric against SA running on a single-core CPU with just 100 qubits. Our results indicate a promising alternative to classical heuristics for combinatorial optimization, a paradigm where quantum advantage might become possible on near-term superconducting quantum processors with thousands of qubits and without the assistance of error correction.

quant-ph

Linear Preference Optimization: Decoupled Gradient Control via Absolute Regularization

DPO (Direct Preference Optimization) has become a widely used offline preference optimization algorithm due to its simplicity and training stability. However, DPO is prone to overfitting and collapse. To address these challenges, we propose Linear Preference Optimization (LPO), a novel alignment framework featuring three key innovations. First, we introduce gradient decoupling by replacing the log-sigmoid function with an absolute difference loss, thereby isolating the optimization dynamics. Second, we improve stability through an offset constraint combined with a positive regularization term to preserve the chosen response quality. Third, we implement controllable rejection suppression using gradient separation with straightforward estimation and a tunable coefficient that linearly regulates the descent of the rejection probability. Through extensive experiments, we demonstrate that LPO consistently improves performance on various tasks, including general text tasks, math tasks, and text-to-speech (TTS) tasks. These results establish LPO as a robust and tunable paradigm for preference alignment, and we release the source code, models, and training data publicly.

cs.LG

Fast ground penetrating radar dual-parameter full waveform inversion method accelerated by hybrid compilation of CUDA kernel function and PyTorch

This study proposes a high-performance dual-parameter full waveform inversion framework (FWI) for ground-penetrating radar (GPR), accelerated through the hybrid compilation of CUDA kernel functions and PyTorch. The method leverages the computational efficiency of GPU programming while preserving the flexibility and usability of Python-based deep learning frameworks. By integrating customized CUDA kernels into PyTorch's automatic differentiation mechanism, the framework enables accurate and efficient inversion of both dielectric permittivity and electrical conductivity. Experimental evaluations on synthetic data and real wavefield data demonstrate that the proposed method achieves dual-parameter FWI for GPR data while maintaining high accuracy. Moreover, the framework is flexible and extensible, supporting optional regularization strategies such as total variation and multi-scale inversion. These features make the proposed approach a practical and scalable framework for rapid GPR-based subsurface imaging in applications including civil engineering, environmental monitoring, and geophysical exploration.

physics.geo-ph

Experimental realization of the bucket-brigade quantum random access memory

Quantum random access memory (QRAM) enables efficient classical data access for quantum computers -- a prerequisite for many quantum algorithms to achieve quantum speedup. Despite various proposals, the experimental realization of QRAM remains largely unexplored. Here, we experimentally investigate the circuit-based bucket-brigade QRAM with a superconducting quantum processor. To facilitate the experimental implementation, we introduce a hardware-efficient gate decomposition scheme for quantum routers, which effectively reduces the depth of the QRAM circuit by more than 30% compared to the conventional controlled-SWAP-based implementation. We further propose an error mitigation method to boost the QRAM query fidelity. With these techniques, we are able to experimentally implement the QRAM architectures with two and three layers, achieving query fidelities up to 0.800 $\pm$ 0.026 and 0.604$\pm$0.005, respectively. Additionally, we study the error propagation mechanism and the scalability of our QRAM implementation, providing experimental evidence for the noise resilience nature of the bucket-brigade QRAM architecture. Our results highlight the potential of superconducting quantum processors for realizing a scalable QRAM architecture.

quant-ph

Noise Analysis and Hierarchical Adaptive Body State Estimator For Biped Robot Walking With ESVC Foot

The ESVC(Ellipse-based Segmental Varying Curvature) foot, a robot foot design inspired by the rollover shape of the human foot, significantly enhances the energy efficiency of the robot walking gait. However, due to the tilt of the supporting leg, the error of the contact model are amplified, making robot state estimation more challenging. Therefore, this paper focuses on the noise analysis and state estimation for robot walking with the ESVC foot. First, through physical robot experiments, we investigate the effect of the ESVC foot on robot measurement noise and process noise. and a noise-time regression model using sliding window strategy is developed. Then, a hierarchical adaptive state estimator for biped robots with the ESVC foot is proposed. The state estimator consists of two stages: pre-estimation and post-estimation. In the pre-estimation stage, a data fusion-based estimation is employed to process the sensory data. During post-estimation, the acceleration of center of mass is first estimated, and then the noise covariance matrices are adjusted based on the regression model. Following that, an EKF(Extended Kalman Filter) based approach is applied to estimate the centroid state during robot walking. Physical experiments demonstrate that the proposed adaptive state estimator for biped robot walking with the ESVC foot not only provides higher precision than both EKF and Adaptive EKF, but also converges faster under varying noise conditions.

cs.RO

Model Analysis And Design Of Ellipse Based Segmented Varying Curved Foot For Biped Robot Walking

This paper presents the modeling, design, and experimental validation of an Ellipse-based Segmented Varying Curvature (ESVC) foot for bipedal robots. Inspired by the segmented curvature rollover shape of human feet, the ESVC foot aims to enhance gait energy efficiency while maintaining analytical tractability for foot location based controller. First, we derive a complete analytical contact model for the ESVC foot by formulating spatial transformations of elliptical segments only using elementary functions. Then a nonlinear programming approach is engaged to determine optimal elliptical parameters of hind foot and fore foot based on a known mid-foot. An error compensation method is introduced to address approximation inaccuracies in rollover length calculation. The proposed ESVC foot is then integrated with a Hybrid Linear Inverted Pendulum model-based walking controller and validated through both simulation and physical experiments on the TT II biped robot. Experimental results across marking time, sagittal, and lateral walking tasks show that the ESVC foot consistently reduces energy consumption compared to line, and flat feet, with up to 18.52\% improvement in lateral walking. These findings demonstrate that the ESVC foot provides a practical and energy-efficient alternative for real-world bipedal locomotion. The proposed design methodology also lays a foundation for data-driven foot shape optimization in future research.

cs.RO

Simulating fluid vortex interactions on a superconducting quantum processor

Vortex interactions are commonly observed in atmospheric turbulence, plasma dynamics, and collective behaviors in biological systems. However, accurately simulating these complex interactions is highly challenging due to the need to capture fine-scale details over extended timescales, which places computational burdens on traditional methods. In this study, we introduce a quantum vortex method, reformulating the Navier--Stokes (NS) equations within a quantum mechanical framework to enable the simulation of multi-vortex interactions on a quantum computer. We construct the effective Hamiltonian for the vortex system and implement a spatiotemporal evolution circuit to simulate its dynamics over prolonged periods. By leveraging eight qubits on a superconducting quantum processor with gate fidelities of 99.97\% for single-qubit gates and 99.76\% for two-qubit gates, we successfully reproduce natural vortex interactions. This method bridges classical fluid dynamics and quantum computing, offering a novel computational platform for studying vortex dynamics. Our results demonstrate the potential of quantum computing to tackle longstanding challenges in fluid dynamics and broaden applications across both natural and engineering systems.

quant-ph

Experimental Detection of Dissipative Quantum Chaos

More than four decades of research on chaos in isolated quantum systems have led to the identification of universal signatures -- such as level repulsion and eigenstate thermalization -- that serve as cornerstones in our understanding of complex quantum dynamics. The emerging field of dissipative quantum chaos explores how these properties manifest in open quantum systems, where interactions with the environment play an essential role. We report the first experimental detection of dissipative quantum chaos and integrability by measuring the complex spacing ratios (CSRs) of open many-body quantum systems implemented on a high-fidelity superconducting quantum processor. Employing gradient-based tomography, we retrieve a ``donut-shaped'' CSR distribution for chaotic dissipative circuits, a hallmark of level repulsion in open quantum systems. For an integrable circuit, spectral correlations vanish, evidenced by a sharp peak at the origin in the CSR distribution. As we increase the depth of the integrable dissipative circuit, the CSR distribution undergoes an integrability-to-chaos crossover, demonstrating that intrinsic noise in the quantum processor is a dissipative chaotic process. Our results reveal the universal spectral features of dissipative many-body systems and establish present-day quantum computation platforms, which are predominantly used to run unitary simulations, as testbeds to explore dissipative many-body phenomena.

quant-ph