Searcharxiv⌕ Search

arXiv subjects

Heng Fan

Publications and source records attributed to Heng Fan.

At least 145 records · Page 8Linked to original sources

Revealing inherent quantum interference and entanglement of a Dirac particle

Although originally predicted in relativistic quantum mechanics, Zitterbewegung can also appear in some classical systems, which leads to the important question of whether Zitterbewegung of Dirac particles is underlain by a more fundamental and universal interference behavior without classical analogs. We here reveal such an interference pattern in phase space, which underlies but goes beyond Zitterbewegung, and whose nonclassicality is manifested by the negativity of the phase space quasiprobability distribution, and the associated pseudospin-momentum entanglement. We confirm this discovery by numerical simulation and an on-chip experiment, where a superconducting qubit and a quantized microwave field respectively emulate the internal and external degrees of freedom of a Dirac particle. The measured quasiprobability negativities agree well with the numerical simulation. Besides being of fundamental importance, the demonstrated nonclassical effects are useful in quantum technology.

quant-ph↗

Two-level approximation of transmons in quantum quench experiments

Quantum quench is a typical protocol in the study of nonequilibrium dynamics of quantum many-body systems. Recently, a number of experiments with superconducting transmon qubits are reported, in which the spin and hard-core boson models with two energy levels on individual sites are used. The transmons are a multilevel system and the coupled qubits are governed by the Bose-Hubbard model. How well they can be approximated by a two-level system has been discussed and analysed in different ways for specific experiments in the literature. Here, we numerically investigate the accuracy and validity of the two-level approximation for the multilevel transmons based on the concept of Loschmidt echo. Using this method, we are able to calculate the fidelity decay (i.e., the time-dependent overlap of evolving wave functions) due to the state leakage to transmon high energy levels. We present the results for different system Hamiltonians with various initial states, qubit coupling strength, and external driving, and for two kinds of quantum quench experiments with time reversal and time evolution in one direction. We show quantitatively the extent to which the fidelity decays with time for changing coupling strength (or on-site interaction over coupling strength) and filled particle number or locations in the initial states under specific system Hamiltonians, which may serve as a way for assessing the two-level approximation of transmons. Finally, we compare our results with the reported experiments using transmon qubits.

quant-ph↗

Local Compressed Video Stream Learning for Generic Event Boundary Detection

Generic event boundary detection aims to localize the generic, taxonomy-free event boundaries that segment videos into chunks. Existing methods typically require video frames to be decoded before feeding into the network, which contains significant spatio-temporal redundancy and demands considerable computational power and storage space. To remedy these issues, we propose a novel compressed video representation learning method for event boundary detection that is fully end-to-end leveraging rich information in the compressed domain, i.e., RGB, motion vectors, residuals, and the internal group of pictures (GOP) structure, without fully decoding the video. Specifically, we use lightweight ConvNets to extract features of the P-frames in the GOPs and spatial-channel attention module (SCAM) is designed to refine the feature representations of the P-frames based on the compressed information with bidirectional information flow. To learn a suitable representation for boundary detection, we construct the local frames bag for each candidate frame and use the long short-term memory (LSTM) module to capture temporal relationships. We then compute frame differences with group similarities in the temporal domain. This module is only applied within a local window, which is critical for event boundary detection. Finally a simple classifier is used to determine the event boundaries of video sequences based on the learned feature representation. To remedy the ambiguities of annotations and speed up the training process, we use the Gaussian kernel to preprocess the ground-truth event boundaries. Extensive experiments conducted on the Kinetics-GEBD and TAPOS datasets demonstrate that the proposed method achieves considerable improvements compared to previous end-to-end approach while running at the same speed. The code is available at https://github.com/GX77/LCVSL.

cs.CV↗

Collaborative Three-Stream Transformers for Video Captioning

As the most critical components in a sentence, subject, predicate and object require special attention in the video captioning task. To implement this idea, we design a novel framework, named COllaborative three-Stream Transformers (COST), to model the three parts separately and complement each other for better representation. Specifically, COST is formed by three branches of transformers to exploit the visual-linguistic interactions of different granularities in spatial-temporal domain between videos and text, detected objects and text, and actions and text. Meanwhile, we propose a cross-granularity attention module to align the interactions modeled by the three branches of transformers, then the three branches of transformers can support each other to exploit the most discriminative semantic information of different granularities for accurate predictions of captions. The whole model is trained in an end-to-end fashion. Extensive experiments conducted on three large-scale challenging datasets, i.e., YouCookII, ActivityNet Captions and MSVD, demonstrate that the proposed method performs favorably against the state-of-the-art methods.

cs.CV↗

Observation of a superradiant phase transition with emergent cat states

Superradiant phase transitions (SPTs) are important for understanding light-matter interactions at the quantum level, and play a central role in criticality-enhanced quantum sensing. So far, SPTs have been observed in driven-dissipative systems, but the emergent light fields did not show any nonclassical characteristic due to the presence of strong dissipation. Here we report an experimental demonstration of the SPT featuring the emergence of a highly nonclassical photonic field, realized with a resonator coupled to a superconducting qubit, implementing the quantum Rabi model. We fully characterize the light-matter state by Wigner matrix tomography. The measured matrix elements exhibit quantum interference intrinsic of a photonic mesoscopic superposition, and reveal light-matter entanglement

quant-ph↗

Simulating Chern insulators on a superconducting quantum processor

The quantum Hall effect, fundamental in modern condensed matter physics, continuously inspires new theories and predicts emergent phases of matter. Here we experimentally demonstrate three types of Chern insulators with synthetic dimensions on a programable 30-qubit-ladder superconducting processor. We directly measure the band structures of the 2D Chern insulator along synthetic dimensions with various configurations of Aubry-André-Harper chains and observe dynamical localisation of edge excitations. With these two signatures of topology, our experiments implement the bulk-edge correspondence in the synthetic 2D Chern insulator. Moreover, we simulate two different bilayer Chern insulators on the ladder-type superconducting processor. With the same and opposite periodically modulated on-site potentials for two coupled chains, we simulate topologically nontrivial edge states with zero Hall conductivity and a Chern insulator with higher Chern numbers, respectively. Our work shows the potential of using superconducting qubits for investigating different intriguing topological phases of quantum matter.

quant-ph↗

Two Birds, One Stone: A Unified Framework for Joint Learning of Image and Video Style Transfers

Current arbitrary style transfer models are limited to either image or video domains. In order to achieve satisfying image and video style transfers, two different models are inevitably required with separate training processes on image and video domains, respectively. In this paper, we show that this can be precluded by introducing UniST, a Unified Style Transfer framework for both images and videos. At the core of UniST is a domain interaction transformer (DIT), which first explores context information within the specific domain and then interacts contextualized domain information for joint learning. In particular, DIT enables exploration of temporal information from videos for the image style transfer task and meanwhile allows rich appearance texture from images for video style transfer, thus leading to mutual benefits. Considering heavy computation of traditional multi-head self-attention, we present a simple yet effective axial multi-head self-attention (AMSA) for DIT, which improves computational efficiency while maintains style transfer performance. To verify the effectiveness of UniST, we conduct extensive experiments on both image and video style transfer tasks and show that UniST performs favorably against state-of-the-art approaches on both tasks. Code is available at https://github.com/NevSNev/UniST.

cs.CV↗

Observation of multiple steady states with engineered dissipation

Simulating the dynamics of open quantum systems is essential in achieving practical quantum computation and understanding novel nonequilibrium behaviors. However, quantum simulation of a many-body system coupled to an engineered reservoir has yet to be fully explored in present-day experiment platforms. In this work, we introduce engineered noise into a one-dimensional ten-qubit superconducting quantum processor to emulate a generic many-body open quantum system. Our approach originates from the stochastic unravellings of the master equation. By measuring the end-to-end correlation, we identify multiple steady states stemmed from a strong symmetry, which is established on the modified Hamiltonian via Floquet engineering. Furthermore, we find that the information saved in the initial state maintains in the steady state driven by the continuous dissipation on a five-qubit chain. Our work provides a manageable and hardware-efficient strategy for the open-system quantum simulation.

quant-ph↗

Unsupervised Domain Adaptive Detection with Network Stability Analysis

Domain adaptive detection aims to improve the generality of a detector, learned from the labeled source domain, on the unlabeled target domain. In this work, drawing inspiration from the concept of stability from the control theory that a robust system requires to remain consistent both externally and internally regardless of disturbances, we propose a novel framework that achieves unsupervised domain adaptive detection through stability analysis. In specific, we treat discrepancies between images and regions from different domains as disturbances, and introduce a novel simple but effective Network Stability Analysis (NSA) framework that considers various disturbances for domain adaptation. Particularly, we explore three types of perturbations including heavy and light image-level disturbances and instancelevel disturbance. For each type, NSA performs external consistency analysis on the outputs from raw and perturbed images and/or internal consistency analysis on their features, using teacher-student models. By integrating NSA into Faster R-CNN, we immediately achieve state-of-the-art results. In particular, we set a new record of 52.7% mAP on Cityscapes-to-FoggyCityscapes, showing the potential of NSA for domain adaptive detection. It is worth noticing, our NSA is designed for general purpose, and thus applicable to one-stage detection model (e.g., FCOS) besides the adopted one, as shown by experiments. https://github.com/tiankongzhang/NSA.

cs.CV↗

ICAFusion: Iterative Cross-Attention Guided Feature Fusion for Multispectral Object Detection

Effective feature fusion of multispectral images plays a crucial role in multi-spectral object detection. Previous studies have demonstrated the effectiveness of feature fusion using convolutional neural networks, but these methods are sensitive to image misalignment due to the inherent deffciency in local-range feature interaction resulting in the performance degradation. To address this issue, a novel feature fusion framework of dual cross-attention transformers is proposed to model global feature interaction and capture complementary information across modalities simultaneously. This framework enhances the discriminability of object features through the query-guided cross-attention mechanism, leading to improved performance. However, stacking multiple transformer blocks for feature enhancement incurs a large number of parameters and high spatial complexity. To handle this, inspired by the human process of reviewing knowledge, an iterative interaction mechanism is proposed to share parameters among block-wise multimodal transformers, reducing model complexity and computation cost. The proposed method is general and effective to be integrated into different detection frameworks and used with different backbones. Experimental results on KAIST, FLIR, and VEDAI datasets show that the proposed method achieves superior performance and faster inference, making it suitable for various practical scenarios. Code will be available at https://github.com/chanchanchan97/ICAFusion.

cs.CV↗

AttMOT: Improving Multiple-Object Tracking by Introducing Auxiliary Pedestrian Attributes

Multi-object tracking (MOT) is a fundamental problem in computer vision with numerous applications, such as intelligent surveillance and automated driving. Despite the significant progress made in MOT, pedestrian attributes, such as gender, hairstyle, body shape, and clothing features, which contain rich and high-level information, have been less explored. To address this gap, we propose a simple, effective, and generic method to predict pedestrian attributes to support general Re-ID embedding. We first introduce AttMOT, a large, highly enriched synthetic dataset for pedestrian tracking, containing over 80k frames and 6 million pedestrian IDs with different time, weather conditions, and scenarios. To the best of our knowledge, AttMOT is the first MOT dataset with semantic attributes. Subsequently, we explore different approaches to fuse Re-ID embedding and pedestrian attributes, including attention mechanisms, which we hope will stimulate the development of attribute-assisted MOT. The proposed method AAM demonstrates its effectiveness and generality on several representative pedestrian multi-object tracking benchmarks, including MOT17 and MOT20, through experiments on the AttMOT dataset. When applied to state-of-the-art trackers, AAM achieves consistent improvements in MOTA, HOTA, AssA, IDs, and IDF1 scores. For instance, on MOT17, the proposed method yields a +1.1 MOTA, +1.7 HOTA, and +1.8 IDF1 improvement when used with FairMOT. To encourage further research on attribute-assisted MOT, we will release the AttMOT dataset.

cs.CV↗

Efficient Quantum Mixed-State Tomography with Unsupervised Tensor Network Machine Learning

Quantum state tomography (QST) is plagued by the ``curse of dimensionality'' due to the exponentially-scaled complexity in measurement and data post-processing. Efficient QST schemes for large-scale mixed states are currently missing. In this work, we propose an efficient and robust mixed-state tomography scheme based on the locally purified state ansatz. We demonstrate the efficiency and robustness of our scheme on various randomly initiated states with different purities. High tomography fidelity is achieved with much smaller numbers of positive-operator-valued measurement (POVM) bases than the conventional least-square (LS) method. On the superconducting quantum experimental circuit [Phys. Rev. Lett. 119, 180511 (2017)], our scheme accurately reconstructs the Greenberger-Horne-Zeilinger (GHZ) state and exhibits robustness to experimental noises. Specifically, we achieve the fidelity $F \simeq 0.92$ for the 10-qubit GHZ state with just $N_m = 500$ POVM bases, which far outperforms the fidelity $F \simeq 0.85$ by the LS method using the full $N_m = 3^{10} = 59049$ bases. Our work reveals the prospects of applying tensor network state ansatz and the machine learning approaches for efficient QST of many-body states.

quant-ph↗

Sequential sharing of two-qudit entanglement based on the entropic uncertainty relation

Entanglement and uncertainty relation are two focuses of quantum theory. We relate entanglement sharing to the entropic uncertainty relation in a $(d\times d)$-dimensional system via weak measurements with different pointers. We consider both the scenarios of one-sided sequential measurements in which the entangled pair is distributed to multiple Alices and one Bob and two-sided sequential measurements in which the entangled pair is distributed to multiple Alices and Bobs. It is found that the maximum number of observers sharing the entanglement strongly depends on the measurement scenarios, the pointer states of the apparatus, and the local dimension $d$ of each subsystem, while the required minimum measurement precision to achieve entanglement sharing decreases to its asymptotic value with the increase of $d$. The maximum number of observers remain unaltered even when the state is not maximally entangled but has strong-enough entanglement.

quant-ph↗

Divert More Attention to Vision-Language Object Tracking

Multimodal vision-language (VL) learning has noticeably pushed the tendency toward generic intelligence owing to emerging large foundation models. However, tracking, as a fundamental vision problem, surprisingly enjoys less bonus from recent flourishing VL learning. We argue that the reasons are two-fold: the lack of large-scale vision-language annotated videos and ineffective vision-language interaction learning of current works. These nuisances motivate us to design more effective vision-language representation for tracking, meanwhile constructing a large database with language annotation for model learning. Particularly, in this paper, we first propose a general attribute annotation strategy to decorate videos in six popular tracking benchmarks, which contributes a large-scale vision-language tracking database with more than 23,000 videos. We then introduce a novel framework to improve tracking by learning a unified-adaptive VL representation, where the cores are the proposed asymmetric architecture search and modality mixer (ModaMixer). To further improve VL representation, we introduce a contrastive loss to align different modalities. To thoroughly evidence the effectiveness of our method, we integrate the proposed framework on three tracking methods with different designs, i.e., the CNN-based SiamCAR, the Transformer-based OSTrack, and the hybrid structure TransT. The experiments demonstrate that our framework can significantly improve all baselines on six benchmarks. Besides empirical results, we theoretically analyze our approach to show its rationality. By revealing the potential of VL representation, we expect the community to divert more attention to VL tracking and hope to open more possibilities for future tracking with diversified multimodal messages.

cs.CV↗

Deficiency-Aware Masked Transformer for Video Inpainting

Recent video inpainting methods have made remarkable progress by utilizing explicit guidance, such as optical flow, to propagate cross-frame pixels. However, there are cases where cross-frame recurrence of the masked video is not available, resulting in a deficiency. In such situation, instead of borrowing pixels from other frames, the focus of the model shifts towards addressing the inverse problem. In this paper, we introduce a dual-modality-compatible inpainting framework called Deficiency-aware Masked Transformer (DMT), which offers three key advantages. Firstly, we pretrain a image inpainting model DMT_img serve as a prior for distilling the video model DMT_vid, thereby benefiting the hallucination of deficiency cases. Secondly, the self-attention module selectively incorporates spatiotemporal tokens to accelerate inference and remove noise signals. Thirdly, a simple yet effective Receptive Field Contextualizer is integrated into DMT, further improving performance. Extensive experiments conducted on YouTube-VOS and DAVIS datasets demonstrate that DMT_vid significantly outperforms previous solutions. The code and video demonstrations can be found at github.com/yeates/DMT.

cs.CV↗

Quantum simulation of topological zero modes on a 41-qubit superconducting processor

Quantum simulation of different exotic topological phases of quantum matter on a noisy intermediate-scale quantum (NISQ) processor is attracting growing interest. Here, we develop a one-dimensional 43-qubit superconducting quantum processor, named as Chuang-tzu, to simulate and characterize emergent topological states. By engineering diagonal Aubry-Andr$\acute{\mathrm{e}}$-Harper (AAH) models, we experimentally demonstrate the Hofstadter butterfly energy spectrum. Using Floquet engineering, we verify the existence of the topological zero modes in the commensurate off-diagonal AAH models, which have never been experimentally realized before. Remarkably, the qubit number over 40 in our quantum processor is large enough to capture the substantial topological features of a quantum system from its complex band structure, including Dirac points, the energy gap's closing, the difference between even and odd number of sites, and the distinction between edge and bulk states. Our results establish a versatile hybrid quantum simulation approach to exploring quantum topological systems in the NISQ era.

quant-ph↗

CCTV-Gun: Benchmarking Handgun Detection in CCTV Images

Gun violence is a critical security problem, and it is imperative for the computer vision community to develop effective gun detection algorithms for real-world scenarios, particularly in Closed Circuit Television (CCTV) surveillance data. Despite significant progress in visual object detection, detecting guns in real-world CCTV images remains a challenging and under-explored task. Firearms, especially handguns, are typically very small in size, non-salient in appearance, and often severely occluded or indistinguishable from other small objects. Additionally, the lack of principled benchmarks and difficulty collecting relevant datasets further hinder algorithmic development. In this paper, we present a meticulously crafted and annotated benchmark, called \textbf{CCTV-Gun}, which addresses the challenges of detecting handguns in real-world CCTV images. Our contribution is three-fold. Firstly, we carefully select and analyze real-world CCTV images from three datasets, manually annotate handguns and their holders, and assign each image with relevant challenge factors such as blur and occlusion. Secondly, we propose a new cross-dataset evaluation protocol in addition to the standard intra-dataset protocol, which is vital for gun detection in practical settings. Finally, we comprehensively evaluate both classical and state-of-the-art object detection algorithms, providing an in-depth analysis of their generalizing abilities. The benchmark will facilitate further research and development on this topic and ultimately enhance security. Code, annotations, and trained models are available at https://github.com/srikarym/CCTV-Gun.

cs.CV↗

Tunable Coupling Architectures with Capacitively Connecting Pads for Large-Scale Superconducting Multi-Qubit Processors

We have proposed and experimentally verified a tunable inter-qubit coupling scheme for large-scale integration of superconducting qubits. The key feature of the scheme is the insertion of connecting pads between qubit and tunable coupling element. In such a way, the distance between two qubits can be increased considerably to a few millimeters, leaving enough space for arranging control lines, readout resonators and other necessary structures. The increased inter-qubit distance provides more wiring space for flip-chip process and reduces crosstalk between qubits and from control lines to qubits. We use the term Tunable Coupler with Capacitively Connecting Pad (TCCP) to name the tunable coupling part that consists of a transmon coupler and capacitively connecting pads. With the different placement of connecting pads, different TCCP architectures can be realized. We have designed and fabricated a few multi-qubit devices in which TCCP is used for coupling. The measured results show that the performance of the qubits coupled by the TCCP, such as $T_1$ and $T_2$, was similar to that of the traditional transmon qubits without TCCP. Meanwhile, our TCCP also exhibited a wide tunable range of the effective coupling strength and a low residual ZZ interaction between the qubits by properly tuning the parameters on the design. Finally, we successfully implemented an adiabatic CZ gate with TCCP. Furthermore, by introducing TCCP, we also discuss the realization of the flip-chip process and tunable coupling qubits between different chips.

quant-ph↗