SearcharxivSearch

arXiv subjects

Zhiguo Huang

Publications and source records attributed to Zhiguo Huang.

13 recordsLinked to original sources

Toward General Quantum Control with Physics-Informed Large Language Models

Quantum control is essential for quantum information science and technology, yet designing high-fidelity control protocols remains challenging due to complex optimization landscapes, hardware noise, and long pulse sequences. Existing numerical solvers often require problem-specific engineering and produce opaque control amplitudes, while naive large language models (LLMs) lack the physical consistency and long-horizon precision for reliable quantum control synthesis. Here we introduce VF-QCTRL, a physics-informed large language model framework for general quantum control that combines symbolic reasoning with optimization to propose analytic control ans\"atze and coherently refine their parameters through feedback. To systematically evaluate LLM-driven quantum control, we develop QCTRL-BENCH, a benchmark spanning sixteen tasks across single- and multi-qubit systems, closed and open quantum dynamics, noiseless and noisy settings, and both analytic and numerical protocols. Across the benchmark, VF-QCTRL demonstrates strong universality, accuracy, efficiency, and interpretability: it applies to generic quantum control systems without task-specific training, achieves performance competitive with or exceeding state-of-the-art conventional solvers in both noiseless and noisy regimes with query efficiency, exhibits favorable inference-time scaling and pulse resolution scaling, and derives physically interpretable analytical protocols directly from prompts. Our results establish physics-informed LLM-based quantum control as a promising paradigm for accurate, efficient, interpretable, and training-free quantum control protocol design across a broad range of quantum systems.

quant-ph

StepAudio 2.5 Technical Report

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks. However, existing unified foundations often struggle to match the depth of specialized systems across automatic speech recognition (ASR), text-to-speech synthesis (TTS), and realtime spoken interaction. Bridging this gap remains an open challenge. This report presents StepAudio 2.5, a unified audio-language foundation model that matches or exceeds specialized systems across all three capabilities. Rather than treating these tasks as architecturally distinct, we operate on the premise that once text and audio share a multimodal representational space, task specialization becomes a matter of operational regimes: data construction, optimization targets, and decoding constraints. Guided by this insight, we advance the post-training paradigm from standard supervised learning to task-tailored Reinforcement Learning from Human Feedback (RLHF), using it as the primary mechanism to define complex optimization targets. We leverage this RLHF-centric alignment, alongside specialized decoding, to shape a shared backbone into three distinct operational modes. Concretely, the ASR branch advances transcription efficiency via verifiable multi-token decoding; the TTS branch achieves controllable, expressive synthesis through preference-based RLHF and context-rich supervision; and the Realtime branch realizes low-latency, persona-consistent dialogue via generative reward modeling within an RLHF framework. On standard benchmarks, StepAudio 2.5 achieves state-of-the-art results across ASR, TTS, and Realtime, demonstrating that a singular audio-language foundation can successfully internalize the distinct deployment objectives of speech understanding, generation, and live interaction.

eess.AS

Robust and High-Fidelity Controlled Two-Qubit Gates via Asymmetric Parallel Resonant Excitation

Implementing high-fidelity controlled two-qubit gates in dipole-dipole interacting systems, such as rare-earth-ion crystals, in hindered by spectral inhomogeneity and weak coupling. Existing method often rely on detuned pulses, making them susceptible to frequency errors and AC Stark shifts. We propose a robust resonant scheme for arbitrary controlled two-qubit gates that utilizes asymmetric excitation and pulse engineering to achieve decoupled, parallel qubit control. Simulations on rare-earth-ion ensemble qubits demonstrate gate fidelities exceeding 99% within a 170 kHz detuning range with off-resonant excitation below 0.2%. This approach offers a robust, scalable route for quantum computing in spectrally crowded systems.

quant-ph

Towards Verifiable and Self-Correcting AI Physicists for Quantum Many-Body Simulations

While large language models (LLMs) promise to revolutionize automated scientific discovery, their application in rigorous real-world physical research is stalled by two critical barriers: a lack of realistic evaluation benchmarks and systemic LLM hallucinations. Here, we address both problems. We introduce QMP-Bench, a pioneering end-to-end research-level benchmark in quantum many-body simulation consisting of $100$ tasks extracted from $21$ high-impact prestigious journals, presenting a challenge even for current frontier LLMs. To establish a paradigm for reliable and transparent AI physicists, we present PhysVEC, a multi-agent framework that enforces self-verifiable and error correction in AI research. PhysVEC seamlessly integrates programming and scientific verifiers to guarantee coding correctness and principle-based physical validity, yielding interpretable evidence and error correction at each step. PhysVEC significantly outperforms existing LLM baselines on various scenarios in QMP-Bench and presents a favorable inference-time scaling, successfully transforming unreliable AI generations into accurate physical reproductions, paving a robust and trustworthy path towards future automated scientific discovery.

physics.comp-ph

Scalable high-fidelity and near-deterministic preparation of large photon-number states

The scalable preparation of large photon-number (Fock) states is a long-standing frontier in quantum science, with direct implications for quantum metrology and bosonic quantum information processing. Despite substantial progress at small photon numbers, extending state generation to large photon numbers while maintaining high fidelity and operating deterministically remains a significant challenge. Here we demonstrate a scalable and experimentally accessible control protocol for generating large photon-number states using only native spin--oscillator operations. The protocol alternates Jaynes--Cummings interactions with phase-space displacements to imprint photon-number--dependent phases and convert them into selective interference in photon-number space. It already achieves high preparation fidelity unconditionally, while an optional final qubit projection removes residual qubit--field correlations and further enhances the fidelity. Conditioned on this final projection, photon-number state preparation with fidelities exceeding $0.95$ is achieved for photon numbers in the few-hundred regime, with a success probability exceeding $0.90$, placing the protocol in a near-deterministic operating regime. The resulting control sequences remain shallow and are robust against detuning, control noise, and experimentally relevant dissipation. Our results establish a practical route to scalable, high-fidelity photon-number state preparation at large photon numbers and provide a versatile interference-engineering toolbox for nonclassical bosonic state synthesis.

quant-ph

Step-GUI Technical Report

Recent advances in multimodal large language models unlock unprecedented opportunities for GUI automation. However, a fundamental challenge remains: how to efficiently acquire high-quality training data while maintaining annotation reliability? We introduce a self-evolving training pipeline powered by the Calibrated Step Reward System, which converts model-generated trajectories into reliable training signals through trajectory-level calibration, achieving >90% annotation accuracy with 10-100x lower cost. Leveraging this pipeline, we introduce Step-GUI, a family of models (4B/8B) that achieves state-of-the-art GUI performance (8B: 80.2% AndroidWorld, 48.5% OSWorld, 62.6% ScreenShot-Pro) while maintaining robust general capabilities. As GUI agent capabilities improve, practical deployment demands standardized interfaces across heterogeneous devices while protecting user privacy. To this end, we propose GUI-MCP, the first Model Context Protocol for GUI automation with hierarchical architecture that combines low-level atomic operations and high-level task delegation to local specialist models, enabling high-privacy execution where sensitive data stays on-device. Finally, to assess whether agents can handle authentic everyday usage, we introduce AndroidDaily, a benchmark grounded in real-world mobile usage patterns with 3146 static actions and 235 end-to-end tasks across high-frequency daily scenarios (8B: static 89.91%, end-to-end 52.50%). Our work advances the development of practical GUI agents and demonstrates strong potential for real-world deployment in everyday digital interactions.

cs.CV

Relativistic Calculations of Energy Levels, Field Shift Factors, and Polarizabilities of Mercury and Copernicium

Mercury (Hg) and superheavy element copernicium (Cn) are investigated using equation-of-motion relativistic coupled-cluster (EOM-RCC) and configuration interaction plus many-body perturbation theory (CI+MBPT) methods. Key atomic properties including ionization potentials (IP), excitation energies (EEs), isotope field shift factors (F), and static electric dipole polarizabilities ({\alpha}) are calculated for ground and low-lying excited states. To evaluate the theoretical accuracy, calculations for both Hg and Cn are performed, with experimental data of Hg serving as benchmarks. Furthermore, basis set dependence has been systematically evaluated in the EOM-RCC calculations, with corresponding uncertainty estimates having been provided. The calculated atomic properties could provide valuable insights into the electronic structure and chemical behavior of superheavy elements.

physics.atom-ph

An artificially intelligent magnetic resonance spectroscopy quantification method: Comparison between QNet and LCModel on the cloud computing platform CloudBrain-MRS

Objctives: This work aimed to statistically compare the metabolite quantification of human brain magnetic resonance spectroscopy (MRS) between the deep learning method QNet and the classical method LCModel through an easy-to-use intelligent cloud computing platform CloudBrain-MRS. Materials and Methods: In this retrospective study, two 3 T MRI scanners Philips Ingenia and Achieva collected 61 and 46 in vivo 1H magnetic resonance (MR) spectra of healthy participants, respectively, from the brain region of pregenual anterior cingulate cortex from September to October 2021. The analyses of Bland-Altman, Pearson correlation and reasonability were performed to assess the degree of agreement, linear correlation and reasonability between the two quantification methods. Results: Fifteen healthy volunteers (12 females and 3 males, age range: 21-35 years, mean age/standard deviation = 27.4/3.9 years) were recruited. The analyses of Bland-Altman, Pearson correlation and reasonability showed high to good consistency and very strong to moderate correlation between the two methods for quantification of total N-acetylaspartate (tNAA), total choline (tCho), and inositol (Ins) (relative half interval of limits of agreement = 3.04%, 9.3%, and 18.5%, respectively; Pearson correlation coefficient r = 0.775, 0.927, and 0.469, respectively). In addition, quantification results of QNet are more likely to be closer to the previous reported average values than those of LCModel. Conclusion: There were high or good degrees of consistency between the quantification results of QNet and LCModel for tNAA, tCho, and Ins, and QNet generally has more reasonable quantification than LCModel.

physics.med-ph

Determination of Landé $g_J$ factor and Zeeman coefficients in ground-state $^{171}$Yb$^+$ and their applications to quantum frequency standards

We report the determination of the Landé $g_J$ factor and Zeeman coefficients for the ground-state of $^{171}$Yb$^+$, relevant to microwave quantum frequency standards (QFSs). The $g_J$ factor is obtained by using two independent methods: multiconfiguration Dirac-Hartree-Fock and multireference configuration interaction, yielding a consistent value of 2.002615(70). The first- and second-order Zeeman coefficients are determined as 14,010.78(49) Hz/$μ$T and 31.0869(22) mHz/$μ$T$^2$, respectively, based on the calculated $g_J$ factor. These coefficients enable reduced magnetic-field-induced uncertainties, improving the accuracy of the $^{171}$Yb$^+$ microwave QFSs. The results reported in this work also offer potential for improved constraints on variations in fundamental constants through frequency comparisons, and advancing trapped-ion quantum computers based on the ground-state hyperfine splitting of $^{171}$Yb$^+$.

physics.atom-ph

Enhancing the expressivity of quantum neural networks with residual connections

In the recent noisy intermediate-scale quantum era, the research on the combination of artificial intelligence and quantum computing has been greatly developed. Inspired by neural networks, developing quantum neural networks with specific structures is one of the most promising directions for improving network performance. In this work, we propose a quantum circuit-based algorithm to implement quantum residual neural networks (QResNets), where the residual connection channels are constructed by introducing auxiliary qubits to the data-encoding and trainable blocks of the quantum neural networks. Importantly, we prove that when this particular network architecture is applied to a $l$-layer data-encoding, the number of frequency generation forms can be extended from one, namely the difference of the sum of generator eigenvalues, to $\mathcal{O}(l^2)$. And the flexibility in adjusting the corresponding Fourier coefficients can also be improved due to the diversity of spectrum construction methods and the additional optimization degrees of freedom in the generalized residual operators. These results indicate that the residual encoding scheme can achieve better spectral richness and enhance the expressivity of various parameterized quantum circuits. Extensive numerical demonstrations in regression tasks of fitting various functions and applications in image classification with MNIST datasets are offered to present the expressivity enhancement. Our work lays the foundation for a complete quantum implementation of the classical residual neural networks and explores a new strategy for quantum feature map in quantum machine learning.

quant-ph

A full circuit-based quantum algorithm for excited-states in quantum chemistry

Utilizing quantum computer to investigate quantum chemistry is an important research field nowadays. In addition to the ground-state problems that have been widely studied, the determination of excited-states plays a crucial role in the prediction and modeling of chemical reactions and other physical processes. Here, we propose a non-variational full circuit-based quantum algorithm for obtaining the excited-state spectrum of a quantum chemistry Hamiltonian. Compared with previous classical-quantum hybrid variational algorithms, our method eliminates the classical optimization process, reduces the resource cost caused by the interaction between different systems, and achieves faster convergence rate and stronger robustness against noise without barren plateau. The parameter updating for determining the next energy-level is naturally dependent on the energy measurement outputs of the previous energy-level and can be realized by only modifying the state preparation process of ancillary system, introducing little additional resource overhead. Numerical simulations of the algorithm with hydrogen, LiH, H2O and NH3 molecules are presented. Furthermore, we offer an experimental demonstration of the algorithm on a superconducting quantum computing platform, and the results show a good agreement with theoretical expectations. The algorithm can be widely applied to various Hamiltonian spectrum determination problems on the fault-tolerant quantum computers.

quant-ph

Distance and Hop-wise Structures Encoding Enhanced Graph Attention Networks

Numerous works have proven that existing neighbor-averaging Graph Neural Networks cannot efficiently catch structure features, and many works show that injecting structure, distance, position or spatial features can significantly improve performance of GNNs, however, injecting overall structure and distance into GNNs is an intuitive but remaining untouched idea. In this work, we shed light on the direction. We first extracting hop-wise structure information and compute distance distributional information, gathering with node's intrinsic features, embedding them into same vector space and then adding them up. The derived embedding vectors are then fed into GATs(like GAT, AGDN) and then Correct and Smooth, experiments show that the DHSEGATs achieve competitive result. The code is available at https://github.com/hzg0601/DHSEGATs.

cs.LG

Behavior Pattern and Compiled Information Based Performance Prediction in MOOCs

With the development of MOOCs massive open online courses, increasingly more subjects can be studied online. Researchers currently show growing interest in the field of MOOCs, including dropout prediction, cheating detection and achievement prediction. Previous studies on achievement prediction mainly focused on students' video and forum behaviors, and few researchers have considered how well students perform their assignments. In this paper, we choose a C programming course as the experimental subject, which involved 1528 students. This paper mainly focuses on the students' accomplishment behaviors in programming assignments and compiled information from programming assignments. In this paper, feature sequences are extracted from the logs according to submission times, submission order and plagiarism. The experimental results show that the students who did not pass the exam had obvious sequence patterns but that the students who passed the test did not have an obvious sequence pattern. Then, we extract 23 features from the compiled information of students' programming assignments and select the most distinguishing features to predict the students' performances. The experimental results show that we can obtain an accuracy rate of 0.7049 for predicting students' performances.

cs.IR