SearcharxivSearch

arXiv subjects

Yuetao Chen

Publications and source records attributed to Yuetao Chen.

16 recordsLinked to original sources

Standard-quantum-limit-surpassing vector polarimetry using Rydberg atoms in an SU(1,1) interferometer

Vector polarimetry is an important application frontier for Rydberg-atom-based sensing. While prior research has largely concentrated on developing novel measurement schemes, high-sensitivity vector polarimetry remains an open question. Here we propose a theoretical framework for high-sensitivity detection of radio-frequency (RF) electric field polarization direction, which is particularly suitable for weak-field detection. Under a static magnetic field, the asymmetry in coupling between the Zeeman sublevels of the Rydberg atom and the RF field's polarization components enables the polarization angles to be determined from the atomic absorption index, which is retrieved via homodyne detection by incorporating the Rydberg atom system into an SU(1,1) interferometer. We derive the sensitivity of the polarization angles along with the corresponding standard quantum limit (SQL) and quantum Cram\'{e}r--Rao bound (QCRB). Our results demonstrate a sensitivity surpassing the SQL across wide angular ranges using either dual coherent states or a coherent state combined with a squeezed vacuum state as input. Significantly, the optimal sensitivity reaches below \SI{e-6}{\degree}, with sensitivities better than \SI{e-3}{\degree} maintained over most of the angular domain. This work establishes a foundation for high-precision vector polarimetry, thereby advancing the development of Rydberg-atom-based quantum sensing and contributing to a deeper understanding of light--matter interactions.

quant-ph

PRISM: Parametrically Refactoring Inference for Speculative Sampling Draft Models

Large Language Models (LLMs), constrained by their auto-regressive nature, suffer from slow decoding. Speculative decoding methods have emerged as a promising solution to accelerate LLM decoding, attracting attention from both systems and AI research communities. Recently, the pursuit of better draft quality has driven a trend toward parametrically larger draft models, which inevitably introduces substantial computational overhead. While existing work attempts to balance the trade-off between prediction accuracy and compute latency, we address this fundamental dilemma through architectural innovation. We propose PRISM, which disaggregates the computation of each predictive step across different parameter sets, refactoring the computational pathways of draft models to successfully decouple model capacity from inference cost. Through extensive experiments, we demonstrate that PRISM outperforms all existing draft architectures, achieving exceptional acceptance lengths while maintaining minimal draft latency for superior end-to-end speedup. We also re-examine scaling laws with PRISM, revealing that PRISM scales more effectively with expanding data volumes than other draft architectures. Through rigorous and fair comparison, we show that PRISM boosts the decoding throughput of an already highly optimized inference engine by more than 2.6x.

cs.AI

Make Every Draft Count: Hidden State based Speculative Decoding

Speculative decoding has emerged as a pivotal technique to accelerate LLM inference by employing a lightweight draft model to generate candidate tokens that are subsequently verified by the target model in parallel. However, while this paradigm successfully increases the arithmetic intensity of memory-bound inference, it causes significant compute inefficiency: the majority of draft tokens fail verification and are discarded, resulting in waste of computation. Motivated by the goal of recollecting this wasted computation, we propose a novel system that transforms discarded drafts into reusable tokens. Our key insight is to perform auto-regressive prediction at the hidden states level and postpone the integrating token information after the hidden states generation, so the draft hidden states are not contaminated by incorrect tokens, enabling hidden state reuse. To implement such a system, first we introduce a draft model architecture based on auto-regressive hidden states, which preserves richer semantics than token-based drafters to facilitate draft repurposing. Second, we design an efficient token information injection mechanism that leverages our specialized draft model to construct high-quality draft token trees and enables resampling tokens from verification failures. Third, we eliminate the overhead hidden in our design to further maximize hardware utilization. We conducted extensive evaluations against various baselines, demonstrating up to a 3.3x speedup against standard speculative decoding.

cs.CL

Quantum Coherence in Reflected and Refracted Beams: A Van Cittert-Zernike Approach

Recent advances in quantum optics have highlighted the critical role of spatial propagation in controlling the quantum coherence of light beams. However, the evolution of quantum coherence for light beams undergoing fundamental optical processes at dielectric interfaces remains unexplored. Furthermore, manipulating multiphoton correlations typically requires complex interactions that challenge few-photon level implementation. Here, we introduce a quantum van Cittert-Zernike theorem for light beams, describing how their coherence-polarization properties are influenced by reflection and refraction, as well as how these properties evolve upon subsequent propagation. Our work demonstrates that the quantum statistics of photonic systems can be controllably modified through the inherent polarization coupling arising from reflection and refraction at an interface, without relying on conventional light-matter interactions. Our approach reveals regimes where thermal light can exhibit sub-Poissonian statistics with fluctuations below the shot-noise level through post-selected measurements, and this statistical property can be tuned by the incident angle. Remarkably, this quantum statistical modification is governed by a scaling law linking beam collimation to far-field thermalization. Our work establishes a robust, decoherence-avoiding mechanism for quantum state control, advancing the fundamental understanding of coherence in quantum optics and opening new avenues for applications in quantum information and metrology.

quant-ph

Quantum metrology of hopping strength in a one-dimensional electronic chain

The electron hopping between the two sites in a lattice is of fundamental importance in condensed matter physics. Precise control of the hopping strength allows for the prospect of manipulating the properties of electronic materials, such as topological properties, superconductivity, etc. In this framework, measuring the hopping strength of an electronic lattice with high precision is perhaps the most relevant step in controlling the properties of electronic materials. Here, we design a critical quantum metrological protocol to measure the hopping strength in a cavity electronic chain coupling system featuring a pseudo-superradiant phase transition. We show that the cavity ground state, which is initially a squeezed vacuum state, can be utilized as a quantum probe to achieve a high quantum precision of the hopping strength, which can be optimally saturated in either the loss or lossless case. Remarkably, in the presence of chain loss, we find that increasing the electron current in the chain is beneficial for enhancing precision, and the arbitrarily large precision could be obtained by increasing the chain size, in principle. Our results provide an effective method to measure the hopping strength in the electronic chain with high precision, so it has potential applications in critical quantum metrology, condensed matter physics, etc.

quant-ph

Beam Spliter and Localization Induced by Controlled Perturbations after Time Boundary

The recent investigation into the phenomena of refraction and reflection at temporal boundaries, conducted through the lens of spacetime duality, has attracted considerable scholarly interest. This duality unveils insights into the propagation behaviors of beams at the temporal boundaries of perturbed systems. We have delineated a temporal boundary effect amenable to external control through a specifically tailored driving force, augmented by a time-varying constituent within the driving signal. We then unveil the phenomenon of beam splitting-both in time refraction and reflection induced by perturbing the lattice's hopping parameter over time. By introducing varying intensities of aperiodic disorder to the coupling coefficients, we have exercised authority over the reflection angles. Our results lay the groundwork for delving into the temporal evolution of crystalline attributes via temporal boundary effects, while also enabling deliberate manipulation of spatial distribution frequency patterns by regulating the form and magnitude of noise. The results offer a manageable avenue for scrutinizing condensed-matter phenomena through providing an experimentally feasible solution.

physics.optics

DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training

Diffusion Transformers (DiTs) have shown remarkable performance in generating high-quality videos. However, the quadratic complexity of 3D full attention remains a bottleneck in scaling DiT training, especially with high-definition, lengthy videos, where it can consume up to 95% of processing time and demand specialized context parallelism. This paper introduces DSV to accelerate video DiT training by leveraging the dynamic attention sparsity we empirically observe. DSV uses a two-stage algorithm to capture the dynamic sparsity patterns via low-rank based approximation of the original query and key. It employs custom kernels to efficiently identify critical key-value pairs and compute the sparse attention. To accommodate the new sparsity dimension, DSV adopts a hybrid sparsity-aware context parallelism that re-balances the skewed workload across attention heads and blocks due to sparsity heterogeneity. DSV achieves up to 3.02x higher training throughput, scaling to 128 GPUs and 520k token lengths, without quality loss.

cs.DC

DeepServe: Serverless Large Language Model Serving at Scale

In this paper, we propose DEEPSERVE, a scalable and serverless AI platform designed to efficiently serve large language models (LLMs) at scale in cloud environments. DEEPSERVE addresses key challenges such as resource allocation, serving efficiency, and cold start latencies through four main design components. First, DEEPSERVE uses a simple serverless abstraction called the request-job-task model, which helps manage diverse AI workloads across posttraining and model-serving tasks. Second, DEEPSERVE integrates an in-house serving engine named FLOWSERVE using a microkernel-inspired design, NPU-centric execution, and SPMD-based parallelism to optimize LLM serving. Third, DEEPSERVE includes novel scheduling policies tailored for a configuration with both PD-disaggregated and PD-colocated instances. Fourth, DEEPSERVE includes optimizations such as pre-warmed pods, DRAM pre-loading, and NPU-fork, which allow DEEPSERVE to scale up to 64 instances in seconds. DEEPSERVE has been in production for over a year, operating on a large Ascend NPU cluster and providing industrystandard APIs for fine-tuning, agent serving, and model serving to our customers.

cs.DC

Echo: Simulating Distributed Training At Scale

Simulation offers unique values for both enumeration and extrapolation purposes, and is becoming increasingly important for managing the massive machine learning (ML) clusters and large-scale distributed training jobs. In this paper, we build Echo to tackle three key challenges in large-scale training simulation: (1) tracing the runtime training workloads at each device in an ex-situ fashion so we can use a single device to obtain the actual execution graphs of 1K-GPU training, (2) accurately estimating the collective communication without high overheads of discrete-event based network simulation, and (3) accounting for the interference-induced computation slowdown from overlapping communication and computation kernels on the same device. Echo delivers on average 8% error in training step -- roughly 3x lower than state-of-the-art simulators -- for GPT-175B on a 96-GPU H800 cluster with 3D parallelism on Megatron-LM under 2 minutes.

cs.LG

Squeezed Displaced Schrödinger-cat state as a signature of the PT-symmetry phase transition

Parity-time (PT ) symmetric systems are gain-loss systems whose dynamics are governed by non-Hermitian Hamiltonians with degeneracies at exceptional-points (EPs) and has been studied in various photonic, electrical, mechanical systems, and so on. However, it is still an open question how to capture PT symmetry phase transition in electronic system where the transport properties of electron will be dramatically effected. Fortunately, the hybridization between photon and electron offers a novel way not only to control but also probe material properties. Here, we investigate a cavity coupled to a non-Hermitian Su-Schrieffer-Heeger (SSH) chain within mean-field ansatzs. We find that Squeezed Displaced Schrodinger cat (SDSc) will emerge with high fidelity in cavity ground state when PT -symmetry is broken and the fidelity will experience a sharp drop from almost 1 to 0 as PT symmetry recovers. Additionally, in semiclassical limit, we find that there exists local extrema at two sides of $x=0$ in semiclassical photon Hamiltonian $H_{\rm eff}(x, p)$, a clear signature of the emergence of SDSc state in cavity ground state. Thus, the appearance of SDSc state can be used to capture PT-symmetry phase transition which can not be modified by cavity mode. Besides, we exploit the cavity ground state to estimate the phase in the optical interferometer, and show that the quantum Fisher information and nonclassicality will sharply decline at EPs. This reveals that PT-symmetry breaking in electronic materials can also be captured by the quantum Fisher information and nonclassicality in phase estimation.

quant-ph

Evaluating the quantum optimal biased bound in a unitary evolution process

Seeking the available precision limit of unknown parameters is a significant task in quantum parameter estimation. One often resorts to the widely utilized quantum Cramer-Rao bound (QCRB) based on unbiased estimators to finish this task. Nevertheless, most actual estimators are usually biased in the limited number of trials. For this reason, we introduce two effective error bounds for biased estimators based on a unitary evolution process in the framework of the quantum optimal biased bound. Furthermore, we show their estimation performance by two specific examples of the unitary evolution process, including the phase encoding and the SU(2) interferometer process. Our findings will provide an useful guidance for finding the precision limit of unknown parameters.

quant-ph

Global quantum thermometry based on the optimal biased bound

Thermometry is a fundamental parameter estimation problem which is crucial in the development process of natural sciences. One way to solve this problem is to the extensive used local thermometry theory, which makes use of the classical and quantum Cramér-Rao bound as benchmarks of thermometry precision. However, such a thermometry theory can only be used for decreasing temperature fluctuations around a known temperature value and hardly tackle the precision thermometry problem over a wide temperature range. For this reason, we derive two basic bounds on thermometry precision in the global setting and further show their thermometry performance by two specific applications, i.e., noninteracting spin-1/2 gas and a general N-level thermal equilibrium quantum probe.

quant-ph

Gamify Stencil Dwarf on Cloud for Democratizing Scientific Computing

Stencil computation is one of the most important kernels in various scientific computing. Nowadays, most Stencil-driven scientific computing still relies heavily on supercomputers, suffering from expensive access, poor scalability, and duplicated optimizations. This paper proposes Tetris, the first system for high-performance Stencil on heterogeneous CPU+GPU, towards democratizing Stencil-driven scientific computing on Cloud. In Tetris, polymorphic tiling tetrominoes are first proposed to bridge different hardware architectures and various application contexts with a perfect spatial and temporal tessellation automatically. Tetris is contributed by three main components: (1) Underlying hardware characteristics are first captured to achieve a sophisticated Pattern Mapping by register-level tetrominoes; (2) An efficient Locality Enhancer is first presented for data reuse on spatial and temporal dimensions simultaneously by cache/SMEM-level tetrominoes; (3) A novel Concurrent Scheduler is first designed to exploit the full potential of on-cloud memory and computing power by memory-level tetrominoes. Tetris is orthogonal to (and complements) the optimizations or deployments for a wide variety of emerging and legacy scientific computing applications. Results of thermal diffusion simulation demonstrate that the performance is improved by 29.6x, reducing time cost from day to hour, while preserving the original accuracy.

cs.DC

Simultaneous multiple angular displacement estimation precision enhanced by the intramode correlation

The angular displacement estimation is one of significant branches of quantum parameter estimation. However, most of the studies have focused on the single-angular displacement estimation, while the multiple angular displacement estimation in ideal and noisy scenarios is still elusive. In this paper, we investigate the simultaneous multiple angular displacement estimation based on an orbital angular momentum (OAM), together with inputting (d + 1)-mode NOON-like states as the probe state. By revealing the role of the intramode correlation of the probe state, this allows us to give a reasonable explanation for the corresponding quantum Cramer-Rao bound (QCRB) behaviors with and without photon losses. Our analyses suggest that the QCRB for the multiple angular displacement estimation is always positively related to the intramode correlation, especially for the multimode entangled squeezed vacuum state showing the best performance compared to another probe state. More importantly, strengthening the robustness of multiple angular-displacement estimation systems can be achieved by increasing the OAM quantum number.

quant-ph

Evaluating the quantum Ziv-Zakai bound in noisy environments

In the highly non-Gaussian regime, the quantum Ziv-Zakai bound (QZZB) provides a lower bound on the available precision, demonstrating the better performance compared with the quantum Cramér-Rao bound. However, evaluating the impact of a noisy environment on the QZZB without applying certain approximations proposed by Tsang [Phys. Rev. Lett. 108, 230401 (2012)] remains a difficult challenge. In this paper, we not only derive the general form of the QZZB with the photon loss and the phase diffusion by invoking the technique of integration within an ordered product of operators, but also show its estimation performance for several different Gaussian resources, such as a coherent state (CS), a single-mode squeezed vacuum state (SMSVS) and a two-mode squeezed vacuum state (TMSVS). Our results indicate that compared with the SMSVS and the TMSVS, the QZZB for the CS always shows the better estimation performance under the photon-loss environment. More interestingly, for the phase-diffusion environment, the estimation performance of the QZZB for the TMSVS can be better than that for the CS throughout a wide range of phase-diffusion strength. Our findings will provide a useful guidance for investigating the noisy quantum parameter estimation.

quant-ph

High-precision estimation of the parameters in the reservoir via the two-level system

A scheme is proposed to estimate the system and environmental parameter, the detuning, temperature and the squeezing strength with a high precision by the two-level atom system. It hasn't been reported that the squeezing strength estimation through quantum Fisher information. We find entangled state and optimal superposition state are beneficial for parameter estimation with one-qubit probe by calculating quantum Fisher information and fidelity. And the fidelity between initial and final states of the atom can be improved via the two-qubit probe. Moreover, the phenomenon of quantum Fisher information return occurs when the detuning or the temperature is estimated. Our work provides a basis for precision measurement technology and quantum information processing.

quant-ph