SearcharxivSearch

arXiv subjects

Zhedong Zhang

Publications and source records attributed to Zhedong Zhang.

At least 19 recordsLinked to original sources

Quantum Ghost Spectroscopy Reveals Hidden Electronic Coherence in Molecular Aggregates

Ultrafast spectroscopy of molecular systems is fundamentally constrained by the Fourier uncertainty principle: high temporal resolution smears out electronic state signatures, while high spectral resolution obscures dynamic information. Here we overcome this limitation using time-resolved quantum ghost spectroscopy (tr-QGS) with entangled photon pairs, which enables independent control of temporal and spectral scales. We apply this approach to perylene bismide (PBI-1) trimers for energy transfer,by combining a quantum description of light-molecule interaction with time-dependent density matrix renormalization group (TD-DMRG) simulations. This explicitly includes five vibrational modes and nonadiabatic coupling between electronic states. Our simulations reveal that tr-QGS uniquely captures electronic coherence oscillating at 0.7 eV for >50 fs, a signature of nonadiabatic coupling that was obscured in conventional time-resolved fluorescence due to Fourier-limited broadening. Moreover, we observe a direct transfer from electronic to vibrational coherence at 200 fs, providing real-time visualization of vibronic relaxation pathways. The entangled photon correlation enables a sensitivity below the shot-noise limit and suppresses photobleaching artifacts that plague classical measurements. These results establish tr-QGS as a transformative tool for interrogating nonadiabatic dynamics in molecular aggregates, light-harvesting complexes, and photocatalysts, offering a route to reveal quantum coherence in chemistry with unprecedented time-energy precision.

quant-ph

Fröhlich Condensation of Bosons: Graph texture of curl flux network for nonequilibrium properties

Nonequilibrium condensates of bosons subject to energy pump and dissipation are investigated, manifesting the Fröhlich coherence proposed in 1968. A quantum theory is developed to capture such a nonequilibrium nature, yielding a certain graphic structure arising from the detailed-balance breaking. The results show a network of probability curl fluxes that reveals a graph topology. The winding number associated with the flux network is thus identified as a new order parameter for the phase transition towards the Fröhlich condensation (FC), not attainable by the symmetry breaking. Our work demonstrates a global property of the FCs, in significant conjunction with the coherence of cavity polaritons that may exhibit robust cooperative phases driven far from equilibrium.

cond-mat.stat-mech

Quantum-Enhanced Sensing of Excited-State Dynamics with Correlated Photons

The squeezed photons, as a quantum-correlated light with reduced noise, have emerged as a great resource for sensing the structures of matter. Here we study the transient absorption (TA) scheme using the squeezed photons whose spectral correlation of amplitudes can be tailored. A microscopic theory is developed, revealing a highly time-energy-resolved nature of the signal that is not attainable by conventional TA scheme. Such a capability is elaborated by applying to monolayer transition metal dichalcogenide materials (TMDs), achieving a real-time monitoring of valley excitons and their dynamics. Moreover, we show the intermediate squeezing regime-not the strong squeezing-which the time-resolved spectroscopy is in favor of. Our work offers a new paradigm for studying nonequilibrium dynamics of matter, in light of the photocatalysis and optoelectronics.

quant-ph

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing

Movie dubbing aims to synthesize speech that preserves the vocal identity of a reference audio while synchronizing with the lip movements in a target video. Existing methods fail to achieve precise lip-sync and lack naturalness due to explicit alignment at the duration level. While implicit alignment solutions have emerged, they remain susceptible to interference from the reference audio, triggering timbre and pronunciation degradation in in-the-wild scenarios. In this paper, we propose a novel flow matching-based movie dubbing framework driven by the Cognitive Synchronous Diffusion Transformer (CoSync-DiT), inspired by the cognitive process of professional actors. This architecture progressively guides the noise-to-speech generative trajectory by executing acoustic style adapting, fine-grained visual calibrating, and time-aware context aligning. Furthermore, we design the Joint Semantic and Alignment Regularization (JSAR) mechanism to simultaneously constrain frame-level temporal consistency on the contextual outputs and semantic consistency on the flow hidden states, ensuring robust alignment. Extensive experiments on both standard benchmarks and challenging in-the-wild dubbing benchmarks demonstrate that our method achieves the state-of-the-art performance across multiple metrics.

cs.SD

Terahertz cavity hybridization of collective proteins vibrations

Hybrid light-matter states have transformed photonics, yet their realization with driven collective vibrations in biological systems remains an open challenge. Here we show that optically pumped R-phycoerythrin proteins at room temperature support coherent sub-terahertz vibrational modes consistent with Frohlich condensation, and that these modes hybridize with confined terahertz cavity photons in a microfluidic cavity platform. The resulting spectra exhibit a resolved doublet, power- and concentration-dependent redistribution of spectral weight, and linewidth narrowing indicative of cavity-modified dissipation. Quantitative analysis reveals collective square-root of N-scaling of the coupling strength, with cooperativity and splitting-to-linewidth ratios exceeding unity, consistent with the onset of strong collective coupling driven by the vibrational molecular mode. A microscopic nonequilibrium analysis further indicates that the relaxation timescale toward the Frohlich polariton state is on the order of 1-10 microseconds. These findings identify terahertz cavities as a platform for stabilizing and controlling collective molecular vibration dynamics and open opportunities for cavity-engineered vibrational spectroscopy, label-free biosensing and photonic control of energy transport in complex biomolecular systems.

cond-mat.other

Theory for Entangled-Photons Stimulated Raman Scattering versus Nonlinear Absorption for Polyatomic Molecules

Quantum entanglement offers an incredible resource for enhancing the sensing and spectroscopic probes. Here we develop a microscopic theory for the stimulated Raman scattering (SRS) using entangled photons. We demonstrate that the time-energy correlation of the photon pairs can optimize the signal for polyatomic molecules. Our results show that the spectral-line intensity of the entangled-photon SRS (ESRS) is of the same order of magnitude as the one for the entangled two-photon absorption (ETPA); the parameter window is thus identified to do so. Moreover, the vibrational coherence is found to play an important role for enhancing the ESRS against the ETPA intensity. Our work paves a firm road for extending the schemes of molecular spectroscopy with quantum light, based on the observation of the ETPA in experiments.

quant-ph

InstructDubber: Instruction-based Alignment for Zero-shot Movie Dubbing

Movie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character's visual performance. However, existing alignment approaches based on visual features face two key limitations: (1)they rely on complex, handcrafted visual preprocessing pipelines, including facial landmark detection and feature extraction; and (2) they generalize poorly to unseen visual domains, often resulting in degraded alignment and dubbing quality. To address these issues, we propose InstructDubber, a novel instruction-based alignment dubbing method for both robust in-domain and zero-shot movie dubbing. Specifically, we first feed the video, script, and corresponding prompts into a multimodal large language model to generate natural language dubbing instructions regarding the speaking rate and emotion state depicted in the video, which is robust to visual domain variations. Second, we design an instructed duration distilling module to mine discriminative duration cues from speaking rate instructions to predict lip-aligned phoneme-level pronunciation duration. Third, for emotion-prosody alignment, we devise an instructed emotion calibrating module, which finetunes an LLM-based instruction analyzer using ground truth dubbing emotion as supervision and predicts prosody based on the calibrated emotion analysis. Finally, the predicted duration and prosody, together with the script, are fed into the audio decoder to generate video-aligned dubbing. Extensive experiments on three major benchmarks demonstrate that InstructDubber outperforms state-of-the-art approaches across both in-domain and zero-shot scenarios.

cs.SD

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing

Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief reference audio. Existing methods focus primarily on reducing the word error rate while ignoring the importance of lip-sync and acoustic quality. To address these issues, we propose a large language model (LLM) based flow matching architecture for dubbing, named FlowDubber, which achieves high-quality audio-visual sync and pronunciation by incorporating a large speech language model and dual contrastive aligning while achieving better acoustic quality via the proposed voice-enhanced flow matching than previous works. First, we introduce Qwen2.5 as the backbone of LLM to learn the in-context sequence from movie scripts and reference audio. Then, the proposed semantic-aware learning focuses on capturing LLM semantic knowledge at the phoneme level. Next, dual contrastive aligning (DCA) boosts mutual alignment with lip movement, reducing ambiguities where similar phonemes might be confused. Finally, the proposed Flow-based Voice Enhancing (FVE) improves acoustic quality in two aspects, which introduces an LLM-based acoustics flow matching guidance to strengthen clarity and uses affine style prior to enhance identity when recovering noise into mel-spectrograms via gradient vector field prediction. Extensive experiments demonstrate that our method outperforms several state-of-the-art methods on two primary benchmarks.

cs.MM

Photon-resolved Floquet theory approach to spectroscopic quantum sensing

Spectroscopic methods play a vital role in quantum sensing, which uses the quantized nature of atoms or molecules to reach astonishing precision for sensing of, e.g., electric or magnetic fields. In the theoretical treatment, one typically invokes semiclassical methods to describe the light-matter interaction between quantum emitters, e.g., atoms or molecules, and a strong coherent laser field. However, these semiclassical approaches struggle to predict the stochastic measurement fluctuations beyond the mean value, necessary to predict the sensitivity of spectroscopic quantum sensing protocols. Here, we develop a theoretical framework based on the recently developed Photon-resolved Floquet theory (PRFT) which is capable to predict the measurement statistics describing higher order statistics of coherent quantum states of light. The PRFT constructs flow equations for the cumulants of the photonic measurement statistics utilizing only the semiclassical dynamics of the matter system. We apply the PRFT to spectroscopic quantum sensing using dissipative two-level and four-level systems (describing electric field sensing with Rydberg atoms), and demonstrate how to calculate the Fisher information of the measurement statistics with respect to various system parameters. In doing so, we demonstrate that the PRFT is a flexible tool allowing to improve the sensitivity of spectroscopic quantum sensing devices by several orders of magnitudes.

quant-ph

Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing

Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker's voice demonstrated in a short reference audio clip. This task demands the model bridge character performances and complicated prosody structures to build a high-quality video-synchronized dubbing track. The limited scale of movie dubbing datasets, along with the background noise inherent in audio data, hinder the acoustic modeling performance of trained models. To address these issues, we propose an acoustic-prosody disentangled two-stage method to achieve high-quality dubbing generation with precise prosody alignment. First, we propose a prosody-enhanced acoustic pre-training to develop robust acoustic modeling capabilities. Then, we freeze the pre-trained acoustic system and design a disentangled framework to model prosodic text features and dubbing style while maintaining acoustic quality. Additionally, we incorporate an in-domain emotion analysis module to reduce the impact of visual domain shifts across different movies, thereby enhancing emotion-prosody alignment. Extensive experiments show that our method performs favorably against the state-of-the-art models on two primary benchmarks. The demos are available at https://zzdoog.github.io/ProDubber/.

cs.SD

Generating High-quality Symbolic Music Using Fine-grained Discriminators

Existing symbolic music generation methods usually utilize discriminator to improve the quality of generated music via global perception of music. However, considering the complexity of information in music, such as rhythm and melody, a single discriminator cannot fully reflect the differences in these two primary dimensions of music. In this work, we propose to decouple the melody and rhythm from music, and design corresponding fine-grained discriminators to tackle the aforementioned issues. Specifically, equipped with a pitch augmentation strategy, the melody discriminator discerns the melody variations presented by the generated samples. By contrast, the rhythm discriminator, enhanced with bar-level relative positional encoding, focuses on the velocity of generated notes. Such a design allows the generator to be more explicitly aware of which aspects should be adjusted in the generated music, making it easier to mimic human-composed music. Experimental results on the POP909 benchmark demonstrate the favorable performance of the proposed method compared to several state-of-the-art methods in terms of both objective and subjective metrics.

cs.SD

StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing

Given a script, the challenge in Movie Dubbing (Visual Voice Cloning, V2C) is to generate speech that aligns well with the video in both time and emotion, based on the tone of a reference audio track. Existing state-of-the-art V2C models break the phonemes in the script according to the divisions between video frames, which solves the temporal alignment problem but leads to incomplete phoneme pronunciation and poor identity stability. To address this problem, we propose StyleDubber, which switches dubbing learning from the frame level to phoneme level. It contains three main components: (1) A multimodal style adaptor operating at the phoneme level to learn pronunciation style from the reference audio, and generate intermediate representations informed by the facial emotion presented in the video; (2) An utterance-level style learning module, which guides both the mel-spectrogram decoding and the refining processes from the intermediate embeddings to improve the overall style expression; And (3) a phoneme-guided lip aligner to maintain lip sync. Extensive experiments on two of the primary benchmarks, V2C and Grid, demonstrate the favorable performance of the proposed method as compared to the current stateof-the-art. The code will be made available at https://github.com/GalaxyCong/StyleDubber.

cs.CL

Collective Quantum Entanglement in Molecular Cavity Optomechanics

We propose an optomechanical scheme for reaching quantum entanglement in vibration polaritons. The system involves $N$ molecules, whose vibrations can be fairly entangled with plasmonic cavities. We find that the vibration-photon entanglement can exist at room temperature and is robust against thermal noise. We further demonstrate the quantum entanglement between the vibrational modes through the plasmonic cavities, which shows a delocalized nature and an incredible enhancement with the number of molecules. The underlying mechanism for the entanglement is attributed to the strong vibration-cavity coupling which possesses collectivity. Our results provide a molecular optomechanical scheme which offers a promising platform for the study of noise-free quantum resources and macroscopic quantum phenomena.

quant-ph

Two-dimensional UV femtosecond stimulated Raman spectroscopy for molecular polaritons: dark states and beyond

We have developed a femtosecond ultra-voilet (UV) stimulated Raman spectroscopy (UV-FSRS) for $N$ molecules in optical cavities. The scheme enables a real-time monitoring of collective dynamics of molecular polaritons and their coupling to vibrations, along with a crosstalk between polariton and dark states. Through multidimensional projections of the UV-FSRS signal, we identify clear signature of the dark states, e.g., pathways and timescales that used to be invisible in resonant technique. A microscopic theory is developed for the UV-FSRS, so as to reveal the polaritonic population and coherence dynamics that interplay with each other. The resulting signal makes the dark states visible, thereby providing a new technique for probing dark state dynamics and their correlation with polariton modes.

quant-ph

Entangled Photons Enabled Ultrafast Stimulated Raman Spectroscopy for Molecular Dynamics

Quantum entanglement has emerged as a great resource for interactions between molecules and radiation. We propose a new paradigm of stimulated Raman scattering with entangled photons. A quantum ultrafast Raman spectroscopy is developed for condensed-phase molecules, to monitor the exciton populations and coherences. Analytic results are obtained, showing a time-frequency scale not attainable by classical light. The Raman signal presents an unprecedented selectivity of molecular correlation functions, as a result of the Hong-Ou-Mandel interference. This is a typical quantum nature, advancing the spectroscopy for clarity. Our work suggests a new scheme of optical signals and spectroscopy, with potential to unveil advanced information about complex materials.

quant-ph

Multidimensional Coherent Spectroscopy of Molecular Polaritons: Langevin Approach

We present a microscopic theory for nonlinear optical spectroscopy of N molecules in an optical cavity. A quantum Langevin analytical expression is derived for the time- and frequency-resolved signals accounting for arbitrary numbers of vibrational excitations. We identify clear signatures of the polariton-polaron interaction from multidimensional projections of the signal, e.g., pathways and timescales. Cooperative dynamics of cavity polaritons against intramolecular vibrations is revealed, along with a cross talk between long-range coherence and vibronic coupling that may lead to localization effects. Our results further characterize the polaritonic coherence and the population transfer that is slower.

quant-ph

Multiple-Photon Resonance Enabled Quantum Interference in Emission Spectroscopy of N_2^+

Quantum interference occurs frequently in the interaction of laser radiation with materials, leading to a series of fascinating effects such as lasing without inversion, electromagnetically induced transparency, Fano resonance, etc. Such quantum interference effects are mostly enabled by single-photon resonance with transitions in the matter, regardless of how many optical frequencies are involved. Here, we demonstrate quantum interference driven by multiple photons in the emission spectroscopy of nitrogen ions that are resonantly pumped by ultrafast infrared laser pulses. In the spectral domain, Fano resonance is observed in the emission spectrum, where a laser-assisted dynamic Stark effect creates the continuum. In the time domain, the fast-evolving emission is measured, revealing the nature of free-induction decay (FID) arising from quantum radiation and molecular cooperativity. These findings clarify the mechanism of coherent emission of nitrogen ions pumped with MIR pump laser and are likely to be universal. The present work opens a route to explore the important role of quantum interference during the interaction of intense laser pulses with materials near multiple photon resonance.

physics.optics

Quantum Fluctuations and Coherence of a Molecular Polariton Condensate

A full quantum theory beyond the mean-field regime is developed for an exciton polariton condensate, to gain a complete understanding of quantum fluctuations. We find analytical solution for the polariton density matrix, showing the polariton nonlinearity causing fast relaxation correlated with the pump so as to yield the condensation at threshold. Increasing the pump intensity, a nonequilibrium phase transition towards the condensation of lower polaritons emerges, with a statistics transiting from a thermal, through a super-Poissonian and to a nonclassical distribution beyond the understanding at the level of off-diagonal long-range order. The results signify the role of dark states for polariton fluctuations, and lead to a nonclassical counting statistics of emitted photons, which elaborates the role of the key parameters, e.g., pump, detuning and temperature.

cond-mat.mes-hall