SearcharxivSearch

arXiv subjects

Mingjian Zhu

Publications and source records attributed to Mingjian Zhu.

At least 19 recordsLinked to original sources

Dissipative Phase Transitions in an Open Quantum Rabi Model with Two-Photon Processes

We demonstrate a novel mechanism driving complex critical phenomena in open light-atom interacting systems by investigating a parametrically amplified quantum Rabi model (QRM) subject to both single- and two-photon decay. In the classical oscillator limit, four composite phases emerge, arising from the possible normal or superradiant regimes across the upper and lower spin branches. A mean-field analysis reveals that the two-photon decay activates the intrinsic nonlinearity of the QRM. The synergy of the coherent and dissipative two-photon processes, together with the spin-boson coupling, constitutes an ``inverted" regime where superradiance emerges exclusively at weak coupling. This regime features first- and second-order superradiant DPTs separated by a tricritical point. Utilizing an adiabatic approach and the semi-classical Langevin formalism, we further study the steady-state structure beyond the mean-field level. The universality classes of the DPTs are identified, with the corresponding critical and finite-size scaling exponents derived and a scaling ansatz proposed to describe the critical behavior.

quant-ph

Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression

Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering history as images, but the resulting modality shift creates a marked capability gap. Through controlled evaluations of history recovery, matched-state decisions, and complete trajectories, we show that this gap cannot be explained by OCR quality alone. Visual-history agents exhibit systematic drift in action selection, query formulation, stopping, and evidence use, revealing an agentic policy gap. We introduce \textbf{CAPS}, a two-stage \textbf{C}ross-modal \textbf{A}gentic \textbf{P}olicy \textbf{S}elf-distillation framework that uses the same model's stronger text-history policy to supervise its visual-history counterpart. Offline trajectory self-distillation transfers successful text-policy behavior to visual-history inputs, while online policy self-distillation provides dense supervision on states visited by the visual-history policy during reinforcement learning. On SearchQA, CAPS improves over AgentOCR by 5.0\% and 3.4\% with 3B and 7B backbones, respectively. On full-history ALFWorld, the corresponding gains are 15.6\% and 14.5\%. Across settings, CAPS reduces average memory-context cost by up to 63.3\% and peak cost by up to 83.4\% relative to matched text-history policies. These results show that explicit cross-modal policy self-distillation can preserve agent capability under vision--text compression. Our code will be made publicly available in a future release.

cs.AI

C-MOP: Integrating Momentum and Boundary-Aware Clustering for Enhanced Prompt Evolution

Automatic prompt optimization is a promising direction to boost the performance of Large Language Models (LLMs). However, existing methods often suffer from noisy and conflicting update signals. In this research, we propose C-MOP (Cluster-based Momentum Optimized Prompting), a framework that stabilizes optimization via Boundary-Aware Contrastive Sampling (BACS) and Momentum-Guided Semantic Clustering (MGSC). Specifically, BACS utilizes batch-level information to mine tripartite features--Hard Negatives, Anchors, and Boundary Pairs--to precisely characterize the typical representation and decision boundaries of positive and negative prompt samples. To resolve semantic conflicts, MGSC introduces a textual momentum mechanism with temporal decay that distills persistent consensus from fluctuating gradients across iterations. Extensive experiments demonstrate that C-MOP consistently outperforms SOTA baselines like PromptWizard and ProTeGi, yielding average gains of 1.58% and 3.35%. Notably, C-MOP enables a general LLM with 3B activated parameters to surpass a 70B domain-specific dense LLM, highlighting its effectiveness in driving precise prompt evolution. The code is available at https://github.com/huawei-noah/noah-research/tree/master/C-MOP.

cs.CL

Experimental Realization of Thermal Reservoirs with Tunable Temperature in a Trapped-Ion Spin-Boson Simulator

We propose and demonstrate an experimental scheme to engineer thermal baths with independently tunable temperatures and dissipation rates for the motional modes of a trapped-ion system. This approach enables robust thermal-state preparation and quantum simulations of open-system dynamics in bosonic and spin-boson models at well-controlled finite temperatures. We benchmark our protocol by experimentally realizing out-of-equilibrium dynamics of a charge-transfer model at different temperatures. We observe that, when the process occurs at a higher temperature, the transfer rate spectrum broadens, with reduced rates at small donor-acceptor energy gaps and enhanced rates at large gaps. We then employ our scheme to study local-temperature effects in a two-mode vibrationally assisted exciton transfer system, where we observe thermally activated interference pathways for excitation transfer.

quant-ph

Quantum Simulation of Charge and Exciton Transfer in Multi-mode Models using Engineered Reservoirs

Quantum simulation offers a route to study open-system molecular dynamics in non-perturbative regimes by programming the interactions among electronic, vibrational, and environmental degrees of freedom on similar energy scales. Trapped-ion systems possess this capability, with their native spins, phonons, and tunable dissipation integrated within a single platform. Here, we demonstrate an open-system quantum simulation of charge and exciton transfer in a multi-mode linear vibronic coupling model. Employing tailored spin-phonon interactions alongside reservoir engineering techniques, we emulate a system with two dissipative vibrational modes coupled to donor and acceptor electronic sites and follow its non-equilibrium dynamics. We continuously tune the system from the charge transfer (CT) regime to the vibrationally assisted exciton transfer (VAET) regime by controlling the vibronic coupling strengths. We find that degenerate modes enhance CT and VAET rates at large energy gaps, while non-degenerate modes activate slow-mode pathways that reduce the energy-gap dependence, thus enlarging the window for efficient transfer. These results show that the presence of one additional vibration introduces interfering vibrationally assisted pathways and reshapes non-perturbative quantum excitation transfer. Our work establishes a scalable and hardware-efficient route to simulating chemically relevant, many-mode vibronic processes with engineered environments, guiding the design of next-generation organic photovoltaics and molecular electronics.

quant-ph

Dissipation-Assisted Steady-State Entanglement Engineering based on Electron Transfer Models

We propose a series of dissipation-assisted entanglement generation protocols that can be implemented on a trapped-ion quantum simulator. Our approach builds on the single-site molecular electron transfer (ET) model recently realized in the experiment [So et al. Sci. Adv. 10, eads8011 (2024)]. This model leverages spin-dependent boson displacement and dissipation controlled by sympathetic cooling. We show that, when coupled to external degrees of freedom, the ET model can be used as a dissipative quantum control mechanism, enabling the precise tailoring of both spin and phonon steady state of a target sub-system. We derive simplified analytical formalisms that offer intuitive insights into the dissipative dynamics. Using realistic interactions in a trapped-ion system, we develop a protocol for generating $N$-qubit and $N$-boson $W$ states. Additionally, we generalize this protocol to realize generic $N$-qubit Dicke states with tunable excitation numbers. Finally, we outline a realistic experimental setup to implement our schemes in the presence of noise sources.

quant-ph

Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition

This work presents Pangu Embedded, an efficient Large Language Model (LLM) reasoner developed on Ascend Neural Processing Units (NPUs), featuring flexible fast and slow thinking capabilities. Pangu Embedded addresses the significant computational costs and inference latency challenges prevalent in existing reasoning-optimized LLMs. We propose a two-stage training framework for its construction. In Stage 1, the model is finetuned via an iterative distillation process, incorporating inter-iteration model merging to effectively aggregate complementary knowledge. This is followed by reinforcement learning on Ascend clusters, optimized by a latency-tolerant scheduler that combines stale synchronous parallelism with prioritized data queues. The RL process is guided by a Multi-source Adaptive Reward System (MARS), which generates dynamic, task-specific reward signals using deterministic metrics and lightweight LLM evaluators for mathematics, coding, and general problem-solving tasks. Stage 2 introduces a dual-system framework, endowing Pangu Embedded with a "fast" mode for routine queries and a deeper "slow" mode for complex inference. This framework offers both manual mode switching for user control and an automatic, complexity-aware mode selection mechanism that dynamically allocates computational resources to balance latency and reasoning depth. Experimental results on benchmarks including AIME 2024, GPQA, and LiveCodeBench demonstrate that Pangu Embedded with 7B parameters, outperforms similar-size models like Qwen3-8B and GLM4-9B. It delivers rapid responses and state-of-the-art reasoning quality within a single, unified model architecture, highlighting a promising direction for developing powerful yet practically deployable LLM reasoners.

cs.CL

Monogamy-of-entanglement-inspired protocol to quantify bipartite entanglement using spin squeezing

Quantum entanglement is an essential resource for quantum science and technology. However, entanglement detection and quantification, via typical entanglement measures such as linear entanglement entropy or negativity, can be a very challenging task. Here we propose a protocol to detect bipartite entanglement in a system of $N$ qubits inspired by the concept of monogamy of entanglement, where, given a total system in a pure state with some bipartite entanglement between two subsystems, subsequent unitary evolution and measurement of one of the subsystems may be used to quantify the entanglement between the two. To address the difficulty of detection, we propose to use spin squeezing to quantify the entanglement within the individual subsystem. Knowing that the relation between spin squeezing and some entanglement measures is not one-to-one, we give some suggestions on how a judicious choice of squeezing Hamiltonian can lead to better results in our protocol. For systems with a small number of qubits, we derive analytical results and show how our protocol can work optimally for GHZ states. For larger systems, we show how the accuracy of the protocol can be improved by a proper choice of the squeezing Hamiltonian. Our protocol presents an alternative for entanglement detection in platforms where state tomography is inaccessible or hard to perform. Additionally, the ideas presented here can be extended beyond spin-only systems to expand their applicability.

quant-ph

Saliency-driven Dynamic Token Pruning for Large Language Models

Despite the recent success of large language models (LLMs), LLMs are particularly challenging in long-sequence inference scenarios due to the quadratic computational complexity of the attention mechanism. Inspired by the interpretability theory of feature attribution in neural network models, we observe that not all tokens have the same contribution. Based on this observation, we propose a novel token pruning framework, namely Saliency-driven Dynamic Token Pruning (SDTP), to gradually and dynamically prune redundant tokens based on the input context. Specifically, a lightweight saliency-driven prediction module is designed to estimate the importance score of each token with its hidden state, which is added to different layers of the LLM to hierarchically prune redundant tokens. Furthermore, a ranking-based optimization strategy is proposed to minimize the ranking divergence of the saliency score and the predicted importance score. Extensive experiments have shown that our framework is generalizable to various models and datasets. By hierarchically pruning 65\% of the input tokens, our method greatly reduces 33\% $\sim$ 47\% FLOPs and achieves speedup up to 1.75$\times$ during inference, while maintaining comparable performance. We further demonstrate that SDTP can be combined with KV cache compression method for further compression.

cs.CL

GIM: A Million-scale Benchmark for Generative Image Manipulation Detection and Localization

The extraordinary ability of generative models emerges as a new trend in image editing and generating realistic images, posing a serious threat to the trustworthiness of multimedia data and driving the research of image manipulation detection and location (IMDL). However, the lack of a large-scale data foundation makes the IMDL task unattainable. In this paper, we build a local manipulation data generation pipeline that integrates the powerful capabilities of SAM, LLM, and generative models. Upon this basis, we propose the GIM dataset, which has the following advantages: 1) Large scale, GIM includes over one million pairs of AI-manipulated images and real images. 2) Rich image content, GIM encompasses a broad range of image classes. 3) Diverse generative manipulation, the images are manipulated images with state-of-the-art generators and various manipulation tasks. The aforementioned advantages allow for a more comprehensive evaluation of IMDL methods, extending their applicability to diverse images. We introduce the GIM benchmark with two settings to evaluate existing IMDL methods. In addition, we propose a novel IMDL framework, termed GIMFormer, which consists of a ShadowTracer, Frequency-Spatial block (FSB), and a Multi-Window Anomalous Modeling (MWAM) module. Extensive experiments on the GIM demonstrate that GIMFormer surpasses the previous state-of-the-art approach on two different benchmarks.

cs.CV

Trapped-Ion Quantum Simulation of Electron Transfer Models with Tunable Dissipation

Electron transfer is at the heart of many fundamental physical, chemical, and biochemical processes essential for life. The exact simulation of these reactions is often hindered by the large number of degrees of freedom and by the essential role of quantum effects. Here, we experimentally simulate a paradigmatic model of molecular electron transfer using a multispecies trapped-ion crystal, where the donor-acceptor gap, the electronic and vibronic couplings, and the bath relaxation dynamics can all be controlled independently. By manipulating both the ground-state and optical qubits, we observe the real-time dynamics of the spin excitation, measuring the transfer rate in several regimes of adiabaticity and relaxation dynamics. Our results provide a testing ground for increasingly rich models of molecular excitation transfer processes that are relevant for molecular electronics and light-harvesting systems.

quant-ph

GenDet: Towards Good Generalizations for AI-Generated Image Detection

The misuse of AI imagery can have harmful societal effects, prompting the creation of detectors to combat issues like the spread of fake news. Existing methods can effectively detect images generated by seen generators, but it is challenging to detect those generated by unseen generators. They do not concentrate on amplifying the output discrepancy when detectors process real versus fake images. This results in a close output distribution of real and fake samples, increasing classification difficulty in detecting unseen generators. This paper addresses the unseen-generator detection problem by considering this task from the perspective of anomaly detection and proposes an adversarial teacher-student discrepancy-aware framework. Our method encourages smaller output discrepancies between the student and the teacher models for real images while aiming for larger discrepancies for fake images. We employ adversarial learning to train a feature augmenter, which promotes smaller discrepancies between teacher and student networks when the inputs are fake images. Our method has achieved state-of-the-art on public benchmarks, and the visualization results show that a large output discrepancy is maintained when faced with various types of generators.

cs.CV

GenImage: A Million-Scale Benchmark for Detecting AI-Generated Image

The extraordinary ability of generative models to generate photographic images has intensified concerns about the spread of disinformation, thereby leading to the demand for detectors capable of distinguishing between AI-generated fake images and real images. However, the lack of large datasets containing images from the most advanced image generators poses an obstacle to the development of such detectors. In this paper, we introduce the GenImage dataset, which has the following advantages: 1) Plenty of Images, including over one million pairs of AI-generated fake images and collected real images. 2) Rich Image Content, encompassing a broad range of image classes. 3) State-of-the-art Generators, synthesizing images with advanced diffusion models and GANs. The aforementioned advantages allow the detectors trained on GenImage to undergo a thorough evaluation and demonstrate strong applicability to diverse images. We conduct a comprehensive analysis of the dataset and propose two tasks for evaluating the detection method in resembling real-world scenarios. The cross-generator image classification task measures the performance of a detector trained on one generator when tested on the others. The degraded image classification task assesses the capability of the detectors in handling degraded images such as low-resolution, blurred, and compressed images. With the GenImage dataset, researchers can effectively expedite the development and evaluation of superior AI-generated image detectors in comparison to prevailing methodologies.

cs.CV

Dynamic Resolution Network

Deep convolutional neural networks (CNNs) are often of sophisticated design with numerous learnable parameters for the accuracy reason. To alleviate the expensive costs of deploying them on mobile devices, recent works have made huge efforts for excavating redundancy in pre-defined architectures. Nevertheless, the redundancy on the input resolution of modern CNNs has not been fully investigated, i.e., the resolution of input image is fixed. In this paper, we observe that the smallest resolution for accurately predicting the given image is different using the same neural network. To this end, we propose a novel dynamic-resolution network (DRNet) in which the input resolution is determined dynamically based on each input sample. Wherein, a resolution predictor with negligible computational costs is explored and optimized jointly with the desired network. Specifically, the predictor learns the smallest resolution that can retain and even exceed the original recognition accuracy for each image. During the inference, each input image will be resized to its predicted resolution for minimizing the overall computation burden. We then conduct extensive experiments on several benchmark networks and datasets. The results show that our DRNet can be embedded in any off-the-shelf network architecture to obtain a considerable reduction in computational complexity. For instance, DR-ResNet-50 achieves similar performance with an about 34% computation reduction, while gaining 1.4% accuracy increase with 10% computation reduction compared to the original ResNet-50 on ImageNet.

cs.CV

Vision Transformer Pruning

Vision transformer has achieved competitive performance on a variety of computer vision applications. However, their storage, run-time memory, and computational demands are hindering the deployment to mobile devices. Here we present a vision transformer pruning approach, which identifies the impacts of dimensions in each layer of transformer and then executes pruning accordingly. By encouraging dimension-wise sparsity in the transformer, important dimensions automatically emerge. A great number of dimensions with small importance scores can be discarded to achieve a high pruning ratio without significantly compromising accuracy. The pipeline for vision transformer pruning is as follows: 1) training with sparsity regularization; 2) pruning dimensions of linear projections; 3) fine-tuning. The reduced parameters and FLOPs ratios of the proposed algorithm are well evaluated and analyzed on ImageNet dataset to demonstrate the effectiveness of our proposed method.

cs.CV

Dynamic Feature Pyramid Networks for Object Detection

Feature pyramid network (FPN) is a critical component in modern object detection frameworks. The performance gain in most of the existing FPN variants is mainly attributed to the increase of computational burden. An attempt to enhance the FPN is enriching the spatial information by expanding the receptive fields, which is promising to largely improve the detection accuracy. In this paper, we first investigate how expanding the receptive fields affect the accuracy and computational costs of FPN. We explore a baseline model called inception FPN in which each lateral connection contains convolution filters with different kernel sizes. Moreover, we point out that not all objects need such a complicated calculation and propose a new dynamic FPN (DyFPN). The output features of DyFPN will be calculated by using the adaptively selected branch according to a dynamic gating operation. Therefore, the proposed method can provide a more efficient dynamic inference for achieving a better trade-off between accuracy and computational cost. Extensive experiments conducted on MS-COCO benchmark demonstrate that the proposed DyFPN significantly improves performance with the optimal allocation of computation resources. For instance, replacing inception FPN with DyFPN reduces about 40% of its FLOPs while maintaining similar high performance.

cs.CV

Video Captioning in Compressed Video

Existing approaches in video captioning concentrate on exploring global frame features in the uncompressed videos, while the free of charge and critical saliency information already encoded in the compressed videos is generally neglected. We propose a video captioning method which operates directly on the stored compressed videos. To learn a discriminative visual representation for video captioning, we design a residuals-assisted encoder (RAE), which spots regions of interest in I-frames under the assistance of the residuals frames. First, we obtain the spatial attention weights by extracting features of residuals as the saliency value of each location in I-frame and design a spatial attention module to refine the attention weights. We further propose a temporal gate module to determine how much the attended features contribute to the caption generation, which enables the model to resist the disturbance of some noisy signals in the compressed videos. Finally, Long Short-Term Memory is utilized to decode the visual representations into descriptions. We evaluate our method on two benchmark datasets and demonstrate the effectiveness of our approach.

cs.CV

Attribute-Aware Attention Model for Fine-grained Representation Learning

How to learn a discriminative fine-grained representation is a key point in many computer vision applications, such as person re-identification, fine-grained classification, fine-grained image retrieval, etc. Most of the previous methods focus on learning metrics or ensemble to derive better global representation, which are usually lack of local information. Based on the considerations above, we propose a novel Attribute-Aware Attention Model ($A^3M$), which can learn local attribute representation and global category representation simultaneously in an end-to-end manner. The proposed model contains two attention models: attribute-guided attention module uses attribute information to help select category features in different regions, at the same time, category-guided attention module selects local features of different attributes with the help of category cues. Through this attribute-category reciprocal process, local and global features benefit from each other. Finally, the resulting feature contains more intrinsic information for image recognition instead of the noisy and irrelevant features. Extensive experiments conducted on Market-1501, CompCars, CUB-200-2011 and CARS196 demonstrate the effectiveness of our $A^3M$. Code is available at https://github.com/iamhankai/attribute-aware-attention.

cs.CV