SearcharxivSearch

arXiv subjects

Wenhui Huang

Publications and source records attributed to Wenhui Huang.

At least 19 recordsLinked to original sources

GeoWorldAD: Geometry World Action Model for Autonomous Driving

Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although recent Vision/Video-Action models learn policies directly from visual observations and scale well with advances in vision transformers and large-scale training data, they often lack explicit geometric grounding and future-aware spatial guidance, limiting their ability to balance collision avoidance and driving progress. In this work, we propose GeoWorldAD, a geometry world action model that grounds trajectory planning in ego-aligned 3D space and anticipates short-horizon scene evolution with latent future geometry tokens. Present geometry provides essential spatial constraints for safe planning, while future geometry reveals how surrounding agents and ego-centric free space may evolve, reducing overly conservative decisions without sacrificing safety. To efficiently exploit these geometric cues, GeoWorldAD progressively aggregates multi-scale present geometry and latent future geometry through iterative trajectory refinement. Experiments on NAVSIM v1 and v2 demonstrate state-of-the-art performance, highlighting the effectiveness of explicit 3D geometry grounding and future geometry world modeling for safe and efficient autonomous driving.

cs.RO

Weak-Strong Steady-State Microbunching Accelerator Light Source

We propose a phase space manipulation involving one energy modulation sandwiched by two dispersion sections which converts a bunched particle beam or bunch train to ultra-high-harmonic density modulation, while the energy modulation in principle can be arbitrarily weak. The same scheme can also be used for energy bunching, creating energy levels in a bunched beam. We further propose a mechanism invoking three laser modulators in a storage ring to longitudinally focus the electron beam both weakly and strongly, such that a microbunch train and its high-density-harmonics or energy bunching form and sustain turn-by-turn. We call this mechanism weak-strong steady-state microbunching (Weak-Strong SSMB). The longitudinal beta function can vary by seven orders of magnitude along such a ring, with the minimal value squeezed to 10 nm. An example application of Weak-Strong SSMB for kW coherent EUV radiation is presented. Extension to X-ray can be anticipated. An energy-leveled electron beam enables $\gamma$-ray frequency comb production. The ideas can be scaled to wavelengths like RF and THz, for bunch length and energy spread control, ultrashort X-ray and coherent THz generation. Our work establishes a new paradigm for longitudinal dynamics study, accelerator light source development, and opens great potential for accelerator physics and technology.

physics.acc-ph

Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset

Urban transportation systems face growing safety challenges that require scalable intelligence for emerging smart mobility infrastructures. While recent advances in foundation models and large-scale multimodal datasets have strengthened perception and reasoning in intelligent transportation systems (ITS), existing research remains largely centered on microscopic autonomous driving (AD), with limited attention to city-scale traffic analysis. In particular, open-ended safety-oriented visual question answering (VQA) and corresponding foundation models for reasoning over heterogeneous roadside camera observations remain underexplored. To address this gap, we introduce the Land Transportation Dataset (LTD), a large-scale open-source vision-language dataset for open-ended reasoning in urban traffic environments. LTD contains 11.6K high-quality VQA pairs collected from heterogeneous roadside cameras, spanning diverse road geometries, traffic participants, illumination conditions, and adverse weather. The dataset integrates three complementary tasks: fine-grained multi-object grounding, multi-image camera selection, and multi-image risk analysis, requiring joint reasoning over minimally correlated views to infer hazardous objects, contributing factors, and risky road directions. To ensure annotation fidelity, we combine multi-model vision-language generation with cross-validation and human-in-the-loop refinement. Building upon LTD, we further propose UniVLT, a transportation foundation model trained via curriculum-based knowledge transfer to unify microscopic AD reasoning and macroscopic traffic analysis within a single architecture. Extensive experiments on LTD and multiple AD benchmarks demonstrate that UniVLT achieves SOTA performance on open-ended reasoning tasks across diverse domains, while exposing limitations of existing foundation models in complex multi-view traffic scenarios.

cs.CV

AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving

Integrating vision-language models (VLMs) into end-to-end (E2E) autonomous driving (AD) systems has shown promise in improving scene understanding. However, existing integration strategies suffer from several limitations: they either struggle to resolve distribution misalignment between reasoning and action spaces, underexploit the general reasoning capabilities of pretrained VLMs, or incur substantial inference latency during action policy generation, which degrades driving performance. To address these challenges, we propose AutoMoT in this work, an end-to-end AD framework that unifies reasoning and action generation within a single vision-language-action (VLA) model. Our approach leverages a mixture-of-transformer (MoT) architecture with joint attention sharing, which preserves the general reasoning capabilities of pre-trained VLMs while enabling efficient fast-slow inference through asynchronous execution at different task frequencies. Extensive experiments on multiple benchmarks, under both open- and closed-loop settings, demonstrate that AutoMoT achieves competitive performance compared to state-of-the-art methods. We further investigate the functional boundary of pre-trained VLMs in AD, examining when AD-tailored fine-tuning is necessary. Our results show that pre-trained VLMs can achieve competitive multi-task scene understanding performance through semantic prompting alone, while fine-tuning remains essential for action-level tasks such as decision-making and trajectory planning. We refer to https://automot-website.github.io/ for the demonstration videos and qualitative results.

cs.CV

Quantum dynamics of microwave photons in synthetic frequency dimension

Synthetic frequency dimension offers a powerful approach to simulate lattice models and control photon dynamics. However, extending this concept into the quantum regime, particularly at the single-photon level, has remained challenging in photonic platforms. Here, we demonstrate quantum-state initialization and detection of single-photon evolutions within a synthetic frequency lattice by integrating a superconducting qubit with a 16-meter aluminum coaxial cable. A tunable superconducting quantum interference device (SQUID)-based modulator is employed to synthesize lattice couplings and artificial gauge fields. We observe single-photon quantum random walks and Bloch oscillations, as well as nonadiabatic, unidirectional frequency conversion under rapid temporal modulation of the lattice Hamiltonian, together with band-structure measurements. The lattice connectivity can be readily reconfigured to construct higher-dimensional lattices using multiple drive tones. Our results establish superconducting quantum circuits as a versatile platform for programmable Hamiltonians and extensible synthetic lattices with flexible single-photon control.

quant-ph

OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic

Recently, two-stage fine-tuning strategies, e.g., acquiring essential driving knowledge through supervised fine-tuning (SFT) and further enhancing decision-making and planning via reinforcement fine-tuning (RFT), have shown strong potential in advancing the knowledge-driven autonomous driving (AD) paradigm. However, the learning nature of SFT still limits the generalization of reasoning, thereby constraining the full potential of driving performance. Meanwhile, current RFT approaches are primarily applied to downstream tasks, since scene understanding is an open-ended problem where corresponding rewards are difficult to quantify. To address these limitations, we propose OpenREAD, an OPEN-ended REasoning reinforced vision-language model (VLM)-based autonomous driving (AD) framework that enables end-to-end RFT across the full spectrum from high-level reasoning to low-level trajectory planning. Specifically, we begin by constructing large-scale Chain-of-Thought (CoT) annotations on open-source driving-related knowledge datasets, and employ the powerful Qwen3 large language model (LLM) as the critic in RFT to quantify reasoning quality for open-ended questions during reward modeling. Extensive experiments confirm that joint end-to-end RFT yields substantial improvements in both upstream and downstream tasks, enabling OpenREAD to achieve state-of-the-art performance on reasoning and planning benchmarks.

cs.CV

Thermalizing channel states for rapid qubit heating

Although known for negatively impacting the operation of superconducting qubits, thermal baths are shown to exert qubit control in a positive way, provided they are properly engineered. We demonstrate an experimental method to engineer the transduction of microwave driving into heat flow through a leaky resonator. Given the precise conversion, a qubit receiving the heat flow obtains a quasi-thermal equilibrium with arbitrary target temperature in hundreds of nanoseconds. We show that the dynamics of the quantum transducing process is described by thermalizing channel states, generated from the double dressings of the resonator by the semi-classical driving and the qubit-resonator coupling. Their spectrum, coupling, and driving strength determine the channel rate of energy flow, along with the relaxation rates of photon leakage into the bath. The analytical prediction is shown to match well with the experimental measurements on an Xmon qubit circuit.

quant-ph

MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning

Integrating visual-language instructions into visuomotor policies is gaining momentum in robot learning for enhancing open-world generalization. Despite promising advances, existing approaches face two challenges: limited language steerability when no generated reasoning is used as a condition, or significant inference latency when reasoning is incorporated. In this work, we introduce MoTVLA, a mixture-of-transformers (MoT)-based vision-language-action (VLA) model that integrates fast-slow unified reasoning with behavior policy learning. MoTVLA preserves the general intelligence of pre-trained VLMs (serving as the generalist) for tasks such as perception, scene understanding, and semantic planning, while incorporating a domain expert, a second transformer that shares knowledge with the pretrained VLM, to generate domain-specific fast reasoning (e.g., robot motion decomposition), thereby improving policy execution efficiency. By conditioning the action expert on decomposed motion instructions, MoTVLA can learn diverse behaviors and substantially improve language steerability. Extensive evaluations across natural language processing benchmarks, robotic simulation environments, and real-world experiments confirm the superiority of MoTVLA in both fast-slow reasoning and manipulation task performance.

cs.RO

High-performance multiplexed readout of superconducting qubits with a tunable broadband Purcell filter

Fast, high-fidelity, and low back-action readout plays a crucial role in the advancement of quantum error correction (QEC). Here, we demonstrate high-performance multiplexed readout of superconducting qubits using a tunable broadband Purcell filter, effectively resolving the fundamental trade-off between measurement speed and photon-noise-induced dephasing. By dynamically tuning the filter parameters, we suppress photon-noise-induced dephasing by a factor of 7 in idle status, while enabling rapid, high-fidelity readout in measurement status. We achieve 99.6\% single-shot readout fidelity with 100~ns readout pulse, limited primarily by relaxation errors during readout. Using a multilevel readout protocol, we further attain 99.9\% fidelity in 50~ns. Simultaneous readout of three qubits using 100~ns pulses achieves an average fidelity of 99.5\% with low crosstalk. Additionally, the readout exhibits high quantum-nondemolition (QND) performance: 99.4\% fidelity over repeated measurements and a low leakage rate below 0.1\%. Building on the tunable broadband filter, we further propose a scalable readout scheme for surface code QEC with enhanced multiplexing capability, offering a promising solution for fast and scalable QEC.

quant-ph

Terahertz-Driven Nano-tip Field-Emission Electron Gun and Cascaded Acceleration

This paper reports two versions of terahertz (THz)-driven nanotip field-emission electron guns: single-layer reflective guns (SLRGs) and double-layer reflective guns (DLRGs). Both guns use nanotip emitters and accelerate electrons through the electric field of the THz wave. SLRGs employ a reflective structure to superimpose the initial and subsequent half-cycles of the THz electric field, enhancing the field amplitude and acceleration efficiency. Experiments have demonstrated that SLRGs achieve higher acceleration efficiency than single-layer nonreflective guns (SLNRGs) for identical THz input energies. This constitutes direct experimental verification of the efficacy of the reflective structure. Theoretically, SLRGs operating in single-feed mode can match the acceleration efficiency of dual-feed SLNRGs while reducing operational complexity. DLRGs demonstrate THz-driven cascaded electron acceleration through precise scanning of the delay between two incident THz beams. This represents a direct experimental demonstration of cascaded acceleration in THz-driven electron sources. The experimental results of DLRGs align closely with the results of electron dynamics predicted by simulations. This establishes the foundation for developing multilayer high-acceleration-efficiency THz-driven high-energy electron guns. The ability to manipulate the THz for each layer individually holds promising potential for improving the beam quality of THz electron guns.

physics.acc-ph

Sub-5-fs compression and synchronization of relativistic electron bunches enabled by a high-gradient $\alpha$-magnet and low-jitter photoinjector

Generating high-brightness relativistic electron bunches with few-femtosecond duration, while simultaneously achieving few-fs synchronization with ultrafast lasers, remains an outstanding challenge at the frontier of accelerator physics and ultrafast science. In this Letter, we present the beam physics and experimental demonstration of a new method that, for the first time, enables simultaneous control of bunch duration and synchronization with few-fs precision. Timing stabilization is achieved using a tailored high-gradient $\alpha$-magnet that optimizes the correlation between time of flight and momentum, together with a photocathode RF gun designed to suppress the effect of RF-to-laser timing jitter. Compression is realized by manipulating the time-momentum correlation in phase space, primarily through space-charge effects. Sub-5-fs rms bunch duration and synchronization are demonstrated. This method establishes a new regime in electron bunch control, unlocking new capabilities for ultrafast beam physics and applications.

physics.acc-ph

Observation and Interpretation of Field Emission Saturation Induced by an Ultra-fast Intense Terahertz Field

Field emission under ultra-fast intense terahertz fields provides a promising approach for generating electron bunches with ultrashort pulse duration and high charge densities. It is generally believed that the field emission current described by traditional field emission theory increases dramatically with the applied electric field. However, we conducted extensive field emission experiments using quasi-single-cycle strong-field terahertz radiation at various energy levels and different temperatures and observed an intriguing phenomenon where the emitted charge reached saturation. A novel model is proposed to interpret this phenomenon, which considers the contribution of surface valence electrons and the dynamic replenishment of free electrons from the bulk to the surface. The experimentally observed convex relationship between the emitted charge and terahertz energy is consistent with the model prediction, unlike the concave relationship derived from the traditional field emission formula. In addition, another observed counter-intuitive phenomenon, the inverse correlation between the cathode temperature and saturated emission charge, is also well interpreted by the model. This work offers comprehensive insights into field emission dynamics under ultra-fast intense fields, paving the way for generating electron bunches with unprecedented temporal resolution.

physics.acc-ph

State Similarity in Modular Superconducting Quantum Processors with Classical Communications

As quantum devices continue to scale, distributed quantum computing emerges as a promising strategy for executing large-scale tasks across modular quantum processors. A central challenge in this paradigm is verifying the correctness of computational outcomes when subcircuits are executed independently following circuit cutting. Here we propose a cross-platform fidelity estimation algorithm tailored for modular architectures. Our method achieves substantial reductions in sample complexity compared to previous approaches designed for single-processor systems. We experimentally implement the protocol on modular superconducting quantum processors with up to 6 qubits to verify the similarity of two 11-qubit GHZ states. Beyond verification, we show that our algorithm enables a federated quantum kernel method that preserves data privacy. As a proof of concept, we apply it to a 5-qubit quantum phase learning task using six 3-qubit modules, successfully extracting phase information with just eight training samples. These results establish a practical path for scalable verification and trustworthy quantum machine learning of modular quantum processors.

quant-ph

A Low-Noise and High-Stability DC Source for Superconducting Quantum Circuits

With the rapid scaling of superconducting quantum processors, electronic control systems relying on commercial off-the-shelf instruments face critical bottlenecks in signal density, power consumption, and crosstalk mitigation. Here we present a custom dual-channel direct current (DC) source module (QPower) dedicated for large-scale superconducting quantum processors. The module delivers a voltage range of $\pm$7 V with 200 mA maximum current per channel, while achieving the following key performance benchmarks: noise spectral density of 20 nV/$\sqrt{\mathrm{Hz}}$ at 10 kHz, output ripple $<$500 $\mu$V$_{\mathrm{pp}}$ within 20 MHz bandwidth, and long-term voltage drift $<$5 $\mu$V$_{\mathrm{pp}}$ over 12 hours. Integrated into the control electronics of a 66-qubit quantum processor, QPower enables qubit coherence times of $T_1 = 87.6~\mu\mathrm{s}$ and Ramsey $T_2 = 5.1~\mu\mathrm{s}$, with qubit resonance frequency drift constrained to $\pm$40 kHz during 12-hour operation. This modular design is compact in size and efficient in energy consumption, providing a scalable DC source solution for intermediate-scale quantum processors with stringent noise and stability requirements, with potential extensions to other quantum hardware platforms and precision measurement.

quant-ph

Logical multi-qubit entanglement with dual-rail superconducting qubits

Recent advances in quantum error correction (QEC) across hardware platforms have demonstrated operation near and beyond the fault-tolerance threshold, yet achieving exponential suppression of logical errors through code scaling remains a critical challenge. Erasure qubits, which enable hardware-level detection of dominant error types, offer a promising path toward resource-efficient QEC by exploiting error bias. Single erasure qubits with dual-rail encoding in superconducting cavities and transmons have demonstrated high coherence and low single-qubit gate errors with mid-circuit erasure detection, but the generation of multi-qubit entanglement--a fundamental requirement for quantum computation and error correction--has remained an outstanding milestone. Here, we demonstrate a superconducting processor integrating four dual-rail erasure qubits that achieves the logical multi-qubit entanglement with error-biased protection. Each dual-rail qubit, encoded in pairs of tunable transmons, preserves millisecond-scale coherence times and single-qubit gate errors at the level of $10^{-5}$. By engineering tunable couplings between logical qubits, we generate high-fidelity entangled states resilient to physical qubit noise, including logical Bell states (98.8% fidelity) and a three-logical-qubit Greenberger-Horne-Zeilinger (GHZ) state (93.5% fidelity). A universal gate set is realized through a calibrated logical controlled-NOT (CNOT) gate with 96.2% process fidelity, enabled by coupler-activated $XX$ interactions in the protected logical subspace. This work advances dual-rail architectures beyond single-qubit demonstrations, providing a blueprint for concatenated quantum error correction with erasure qubits.

quant-ph

Logical operations with a dynamical qubit in Floquet-Bacon-Shor code

Quantum error correction (QEC) protects quantum systems against inevitable noises and control inaccuracies, providing a pathway towards fault-tolerant (FT) quantum computation. Stabilizer codes, including surface code and color code, have long been the focus of research and have seen significant experimental progress in recent years. Recently proposed time-dynamical QEC, including Floquet codes and generalized time-dynamical code implementations, opens up new opportunities for FT quantum computation. By employing a periodic schedule of low-weight parity checks, Floquet codes can generate additional dynamical logical qubits, offering enhanced error correction capabilities and potentially higher code performance. Here, we experimentally implement the Floquet-Bacon-Shor code on a superconducting quantum processor. We encode a dynamical logical qubit within a $3\times 3$ lattice of data qubits, alongside a conventional static logical qubit. We demonstrate FT encoding and measurement of the two-qubit logical states, and stabilize these states using repeated error detection. We showcase universal single-qubit logical gates on the dynamical qubit. Furthermore, by implementing a logical CNOT gate, we entangle the dynamical and static logical qubits, generating an error-detected logical Bell state with a fidelity of 75.9\%. Our results highlight the potential of Floquet codes for resource-efficient FT quantum computation.

quant-ph

Experimental observation of exact quantum critical states

Anderson localization physics features three fundamental types of eigenstates: extended, localized, and critical, with the third one exhibiting the exotic properties in-between the former two. Confirming the presence of critical states is challenging, as it typically necessitates either advancing the analysis to the thermodynamic limit or identifying a universal mechanism which can rigorously determine these states. Here we report the unambiguous experimental realization of critical states, governed by a rigorous mechanism for exact quantum critical states, and further observe a generalized mechanism that quasiperiodic zeros in hopping couplings protect the critical states. We implement a programmable quasiperiodic mosaic model with tunable couplings and on-site potentials through a multiple superconducting qubit quantum system. By measuring the time-evolving observables, we identify the coexisting delocalized dynamics and incommensurately distributed zeros in the couplings, which are the defining features of the critical states. We map the localized-to-critical phase transition and demonstrate that critical states persist until quasiperiodic zeros are removed by strong long-range couplings, highlighting a novel generalized mechanism discovered in this experiment and shown with rigorous theory. Finally, we resolve the energy-dependent transition between localized and critical states, revealing the presence of anomalous mobility edges.

quant-ph