SearcharxivSearch

arXiv subjects

Jiahua Li

Publications and source records attributed to Jiahua Li.

11 recordsLinked to original sources

Relaxation-Rectified Anti-Jaynes-Cummings Cascades for Autonomous Fock-State Stabilization

We propose an autonomous scheme for stabilizing prescribed Fock states in a Kerr cavity by rectifying anti-Jaynes--Cummings interactions with auxiliary-qubit relaxation. Full Lindblad simulations demonstrate steady states dominated by the Fock states $|1\rangle$ through $|4\rangle$, accompanied by pronounced Wigner negativity. An analytical birth--death model captures the steady-state populations and identifies the operating regime for stabilization. These results establish auxiliary-qubit relaxation as a resource for autonomous reservoir engineering and provide a route to nonclassical bosonic-state preparation in a single Kerr cavity and, through a boundary reservoir, in a Kerr-cavity chain.

quant-ph

Wigner-Negative Magnon Steady States from Incoherent Qubit Pumping

We show that incoherently pumped qubits can realize a cascaded dissipative mechanism for stabilizing Wigner-negative magnon steady states. The mechanism combines qubit pumping with dispersive magnon-number selectivity to direct the steady-state population toward selected magnon Fock states. In the single-qubit case, the single-magnon population can approach unity, accompanied by strong antibunching and pronounced Wigner negativity. Extending the same principle to multiple qubits yields Wigner-negative steady states dominated by higher magnon Fock components. We further derive an analytical birth--death model that captures the mechanism and agrees with numerical results. These results establish incoherent qubit pumping as a controllable dissipative resource for generating nonclassical magnon states in hybrid quantum systems.

quant-ph

A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages

Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in English but struggle with non-English languages, producing inaccurate mouth shapes and rigid facial expressions. These limitations are mainly caused by English-dominated training datasets and the lack of cross-language generalization ability.To address these challenges, we propose Multilingual Experts (MuEx), a novel framework featuring a Phoneme-Guided Mixture-of-Experts (PG-MoE) architecture that employs phonemes and visemes as universal intermediaries to bridge the gap between audio and visual modalities, enabling lifelike multilingual TFS. We extract speech and visual features as phonemes and visemes, respectively, which represent the basic units of speech sounds and mouth movements, to alleviate linguistic differences and dataset bias.Furthermore, we introduce the Phoneme-Viseme Alignment Mechanism (PV-Align), which establishes robust cross-modal correspondences between phonemes and visemes to improve audiovisual synchronization. In addition, we construct a Multilingual Talking Face Dataset (MTFD) comprising 12 diverse languages with 95.04 hours of high-quality videos for training and evaluating multilingual TFS performance.Extensive experiments demonstrate that MuEx achieves superior performance across all languages in MTFD and exhibits effective zero-shot generalization to unseen languages without additional training.

cs.CV

Rotation Control Unlearning: Quantifying and Controlling Continuous Unlearning for LLM with The Cognitive Rotation Space

As Large Language Models (LLMs) become increasingly prevalent, their security vulnerabilities have already drawn attention. Machine unlearning is introduced to seek to mitigate these risks by removing the influence of undesirable data. However, existing methods not only rely on the retained dataset to preserve model utility, but also suffer from cumulative catastrophic utility loss under continuous unlearning requests. To solve this dilemma, we propose a novel method, called Rotation Control Unlearning (RCU), which leverages the rotational salience weight of RCU to quantify and control the unlearning degree in the continuous unlearning process. The skew symmetric loss is designed to construct the existence of the cognitive rotation space, where the changes of rotational angle can simulate the continuous unlearning process. Furthermore, we design an orthogonal rotation axes regularization to enforce mutually perpendicular rotation directions for continuous unlearning requests, effectively minimizing interference and addressing cumulative catastrophic utility loss. Experiments on multiple datasets confirm that our method without retained dataset achieves SOTA performance.

cs.LG

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents

Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Although existing Large Language Model (LLM)-based approaches have advanced long video understanding, they remain bottlenecked by task-agnostic, fixed-granularity perception pipelines and suffer from vision-language hallucinations. Inspired by human adaptive perception and active verification, we propose CogniGPT, a framework leveraging an interactive loop between a Multi-Granular Perception Agent (MPA) and an Active Verification Agent (AVA). Specifically, instead of predetermined heuristics, MPA adaptively determines the optimal perception granularity and strategy based on the evolving context, while AVA actively mines multi-perspective visual evidence to cross-verify key observations and eliminate hallucinations. This interaction allows CogniGPT to efficiently identify a minimal set of reliable task-related clues. Extensive experiments on EgoSchema, Video-MME, NExT-QA, and MovieChat demonstrate its superiority in accuracy and efficiency. Notably, on EgoSchema, it surpasses existing training-free methods using only 11.2 frames and achieves performance comparable to Gemini 1.5-Pro.

cs.CV

Deformation-Recovery Diffusion Model (DRDM): Instance Deformation for Image Manipulation and Synthesis

In medical imaging, the diffusion models have shown great potential for synthetic image generation tasks. However, these approaches often lack the interpretable connections between the generated and real images and can create anatomically implausible structures or illusions. To address these limitations, we propose the Deformation-Recovery Diffusion Model (DRDM), a novel diffusion-based generative model that emphasises morphological transformation through deformation fields rather than direct image synthesis. DRDM introduces a topology-preserving deformation field generation strategy, which randomly samples and integrates multi-scale Deformation Velocity Fields (DVFs). DRDM is trained to learn to recover unrealistic deformation components, thus restoring randomly deformed images to a realistic distribution. This formulation enables the generation of diverse yet anatomically plausible deformations that preserve structural integrity, thereby improving data augmentation and synthesis for downstream tasks such as few-shot learning and image registration. Experiments on cardiac Magnetic Resonance Imaging and pulmonary Computed Tomography show that DRDM is capable of creating diverse, large-scale deformations, while maintaining anatomical plausibility of deformation fields. Additional evaluations on 2D image segmentation and 3D image registration tasks indicate notable performance gains, underscoring DRDM's potential to enhance both image manipulation and generative modelling in medical imaging applications. Project page: https://jianqingzheng.github.io/def_diff_rec/

eess.IV

Multimodal Deformable Image Registration for Long-COVID Analysis Based on Progressive Alignment and Multi-perspective Loss

Long COVID is characterized by persistent symptoms, particularly pulmonary impairment, which necessitates advanced imaging for accurate diagnosis. Hyperpolarised Xenon-129 MRI (XeMRI) offers a promising avenue by visualising lung ventilation, perfusion, as well as gas transfer. Integrating functional data from XeMRI with structural data from Computed Tomography (CT) is crucial for comprehensive analysis and effective treatment strategies in long COVID, requiring precise data alignment from those complementary imaging modalities. To this end, CT-MRI registration is an essential intermediate step, given the significant challenges posed by the direct alignment of CT and Xe-MRI. Therefore, we proposed an end-to-end multimodal deformable image registration method that achieves superior performance for aligning long-COVID lung CT and proton density MRI (pMRI) data. Moreover, our method incorporates a novel Multi-perspective Loss (MPL) function, enhancing state-of-the-art deep learning methods for monomodal registration by making them adaptable for multimodal tasks. The registration results achieve a Dice coefficient score of 0.913, indicating a substantial improvement over the state-of-the-art multimodal image registration techniques. Since the XeMRI and pMRI images are acquired in the same sessions and can be roughly aligned, our results facilitate subsequent registration between XeMRI and CT, thereby potentially enhancing clinical decision-making for long COVID management.

eess.IV

Highly nonclassical phonon emission statistics through two-phonon loss of van der Pol oscillator

The ability to produce nonclassical wave in a system is essential for advances in quantum communication and computation. Here, we propose a scheme to generate highly nonclassical phonon emission statistics -- antibunched wave in a quantum van der Pol (vdP) oscillator subject to an external driving, both single- and two-phonon losses. It is found that phonon antibunching depends significantly on the nonlinear two-phonon loss of the vdP oscillator, where the degree of the antibunching increases monotonically with the two-phonon loss and the distinguished parameter regimes with optimal antibunching and single-phonon emission are identified clearly. In addition, we give an in-depth insight into strong antibunching in the emitted phonon statistics by analytical calculations using a three-oscillator-level model, which agree well with the full numerical simulations employing a master-equation approach and a Schrodinger-equation approach at weak driving. In turn, the fluorescence phonon emission spectra of the vdP oscillator, given by the power spectral density, are also evaluated. We further show that high phonon emission amplitudes, simultaneously accompanied by strong phonon antibunching, are attainable in the vdP system, which are beneficial to the correlation measurement in practical experiments. Our approach only requires a single vdP oscillator, without the need for reconfiguring the two coupled nonlinear resonators or the complex nanophotonic structures as compared to the previous unconventional blockade schemes. The present scheme could inspire methods to achieve antibunching in other systems.

physics.optics

Quantum Interference of Stored Coherent Spin-wave Excitations in a Two-channel Memory

Quantum memories are essential elements in long-distance quantum networks and quantum computation. Significant advances have been achieved in demonstrating relative long-lived single-channel memory at single-photon level in cold atomic media. However, the qubit memory corresponding to store two-channel spin-wave excitations (SWEs) still faces challenges, including the limitations resulting from Larmor procession, fluctuating ambient magnetic field, and manipulation/measurement of the relative phase between the two channels. Here, we demonstrate a two-channel memory scheme in an ideal tripod atomic system, in which the total readout signal exhibits either constructive or destructive interference when the two-channel SWEs are retrieved by two reading beams with a controllable relative phase. Experimental result indicates quantum coherence between the stored SWEs. Based on such phase-sensitive storage/retrieval scheme, measurements of the relative phase between the two SWEs and Rabi oscillation, as well as elimination of the collapse and revival of the readout signal, are experimentally demonstrated.

quant-ph

Correlation and entanglement of two-component Bose-Einstein condensates in a double well

We consider a novel system of two-component atomic Bose-Einstein condensate in a double-well potential. Based on the well-known two-mode approximation, we demonstrate that there are obvious avoided level-crossings when both interspecies and intraspecies interactions of two species are increased. The quantum dynamics of the system exhibits revised oscillating behaviors compared with a single component condensate. We also examine the entanglement of two species. Our numerical calculations show the onset of entanglement can be signed as a violation of Cauchy-Schwarz inequality of second-order cross correlation function. Consequently, we use Von Neumann entropy to quantity the degree of entanglement.

cond-mat.other

Deformed Two-Mode Quadrature Operators in Noncommutative Space

Starting from noncommutative quantum mechanics algebra, we investigate the variances of the deformed two-mode quadrature operators under the evolution of three types of two-mode squeezed states in noncommutative space. A novel conclusion can be found and it may associate the checking of the variances in noncommutative space with homodyne detecting technology. Moreover, we analyze the influence of the scaling parameter on the degree of squeezing for the deformed level and the corresponding consequences.

hep-th