Searcharxiv⌕ Search

arXiv subjects

Heng Fan

Publications and source records attributed to Heng Fan.

At least 109 records · Page 6Linked to original sources

Unlocking Cross-Lingual Sentiment Analysis through Emoji Interpretation: A Multimodal Generative AI Approach

Emojis have become ubiquitous in online communication, serving as a universal medium to convey emotions and decorative elements. Their widespread use transcends language and cultural barriers, enhancing understanding and fostering more inclusive interactions. While existing work gained valuable insight into emojis understanding, exploring emojis' capability to serve as a universal sentiment indicator leveraging large language models (LLMs) has not been thoroughly examined. Our study aims to investigate the capacity of emojis to serve as reliable sentiment markers through LLMs across languages and cultures. We leveraged the multimodal capabilities of ChatGPT to explore the sentiments of various representations of emojis and evaluated how well emoji-conveyed sentiment aligned with text sentiment on a multi-lingual dataset collected from 32 countries. Our analysis reveals that the accuracy of LLM-based emoji-conveyed sentiment is 81.43%, underscoring emojis' significant potential to serve as a universal sentiment marker. We also found a consistent trend that the accuracy of sentiment conveyed by emojis increased as the number of emojis grew in text. The results reinforce the potential of emojis to serve as global sentiment indicators, offering insight into fields such as cross-lingual and cross-cultural sentiment analysis on social media platforms. Code: https://github.com/ResponsibleAILab/emoji-universal-sentiment.

cs.CL↗

Topological eigenvalues braiding and quantum state transfer near a third-order exceptional point

Non-Hermitian systems exhibit a variety of unique features rooted in the presence of exceptional points (EP). The distinct topological structure in the proximity of an EP gives rise to counterintuitive behaviors absent in Hermitian systems, which emerge after encircling the EP either quasistatically or dynamically. However, experimental exploration of EP encirclement in quantum systems, particularly those involving high-order EPs, remains challenging due to the difficulty of coherently controlling more degrees of freedom. In this work, we experimentally investigate the eigenvalues braiding and state transfer arising from the encirclement of EP in a three-dimensional non-Hermitian quantum system using superconducting circuits. We characterize the second- and third-order EPs through the coalescence of eigenvalues. Then we reveal the topological structure near the EP3 by quasistatically encircling it along various paths with three independent parameters, which yields the eigenvalues braiding described by the braid group $B_3$. Additionally, we observe chiral state transfer between three eigenstates under a fast driving scheme when no EPs are enclosed, while time-symmetric behavior occurs when at least one EP is encircled. Our findings offer insights into understanding non-Hermitian topological structures and the manipulation of quantum states through dynamic operations.

quant-ph↗

GSOT3D: Towards Generic 3D Single Object Tracking in the Wild

In this paper, we present a novel benchmark, GSOT3D, that aims at facilitating development of generic 3D single object tracking (SOT) in the wild. Specifically, GSOT3D offers 620 sequences with 123K frames, and covers a wide selection of 54 object categories. Each sequence is offered with multiple modalities, including the point cloud (PC), RGB image, and depth. This allows GSOT3D to support various 3D tracking tasks, such as single-modal 3D SOT on PC and multi-modal 3D SOT on RGB-PC or RGB-D, and thus greatly broadens research directions for 3D object tracking. To provide highquality per-frame 3D annotations, all sequences are labeled manually with multiple rounds of meticulous inspection and refinement. To our best knowledge, GSOT3D is the largest benchmark dedicated to various generic 3D object tracking tasks. To understand how existing 3D trackers perform and to provide comparisons for future research on GSOT3D, we assess eight representative point cloud-based tracking models. Our evaluation results exhibit that these models heavily degrade on GSOT3D, and more efforts are required for robust and generic 3D object tracking. Besides, to encourage future research, we present a simple yet effective generic 3D tracker, named PROT3D, that localizes the target object via a progressive spatial-temporal network and outperforms all current solutions by a large margin. By releasing GSOT3D, we expect to advance further 3D tracking in future research and applications. Our benchmark and model as well as the evaluation results will be publicly released at our webpage https://github.com/ailovejinx/GSOT3D.

cs.CV↗

Simulation of the massless Dirac field in 1+1D curved spacetime

Simulating the nature of quantum fields in diverse spacetime backgrounds offers valuable insights for the fundamental comprehension of quantum mechanics and general relativity. Here we introduce a novel method for mapping the massless Dirac equation in 1+1D curved spacetime to a controllable quantum simulation model, applicable to various observers' perspectives. We perform numerical simulations of Simpson spacetime and calculate tunneling rates in Painleve and Schwarzschild coordinates, which align closely with theoretical predictions of Hawking radiation. Additionally, we show the transition of Simpson spacetime from a regular black hole to a wormhole as the parameter $"a > r_s"$. This method facilitates the study of spacetime from various coordinate perspectives (observers), providing deeper insights and understanding.

gr-qc↗

Multipartite Entanglement in Crossing the Quantum Critical Point

We investigate the multipartite entanglement for a slow quantum quench crossing a critical point. We consider the quantum Ising model and the Lipkin-Meshkov-Glick model, which are local and full-connected quantum systems, respectively. The multipartite entanglement is quantified by quantum Fisher information with the generator defined as the operator of the ferromagnetic order parameter. The quench dynamics begins with a ground state in a paramagnetic phase, and then the transverse field is driven slowly to cross a quantum critical point, and ends with a zero transverse field. For the quantum Ising model, based on methods of matrix product states, we calculate the quantum Fisher information density of the final state. Numerical results of both linear and nonlinear quenches show that the quantum Fisher information density of the final state scales as a power law of the quench rate, which overall conforms to the prediction of the Kibble-Zurek mechanism with a small correction. We show that this correction results from the long-range behaviors. We also calculate the quantum Fisher information density in the Lipkin-Meshkov-Glick model. The results show that the scaling of quantum Fisher information in this full-connected system conforms to the Kibble-Zurek mechanism better, since the long-range physics cannot be defined in this nonlocal system. Our results reveal that the multipartite entanglement provides an alternative viewpoint to understand the dynamics of quantum phase transitions, specifically, the nontrivial long-range physics.

quant-ph↗

Bounded light cone and robust topological order out of equilibrium

The ground state degeneracy of topologically ordered gapped Hamiltonians is the bedrock for self-correcting quantum memories, which are unfortunately not stable away from equilibrium even at zero temperature. This plague precludes practical robust self-correction since stability at zero temperature is a prerequisite for finite-temperature robustness. In this work, we show that the emergence of a bounded light cone renders the unitary time evolution a quasi-adiabatic continuation that preserves topological order, with the initial ground space retaining its macroscopic distance at all times as a quantum code. We also show how bounded light cones can emerge through suitable perturbations in Kitaev's toric code and honeycomb model. Our results suggest that topological orders and self-correcting quantum memories can be dynamically robust at zero temperature.

quant-ph↗

High-precision pulse calibration of tunable couplers for high-fidelity two-qubit gates in superconducting quantum processors

For superconducting quantum processors, stable high-fidelity two-qubit operations depend on precise flux control of the tunable coupler. However, the pulse distortion poses a significant challenge to the control precision. Current calibration methods, which often rely on microwave crosstalk or additional readout resonators for coupler excitation and readout, tend to be cumbersome and inefficient, especially when couplers only have flux control. Here, we introduce and experimentally validate a novel pulse calibration scheme that exploits the strong coupling between qubits and couplers, eliminating the need for extra coupler readout and excitation. Our method directly measures the short-time and long-time step responses of the coupler flux pulse transient, enabling us to apply predistortion to subsequent signals using fast Fourier transformation and deconvolution. This approach not only simplifies the calibration process but also significantly improves the precision and stability of the flux control. We demonstrate the efficacy of our method through the implementation of diabatic CZ and iSWAP gates with fidelities of $99.61\pm0.04\%$ and $99.82\pm0.02\%$, respectively, as well as a series of diabatic CPhase gates with high fidelities characterized by cross-entropy benchmarking. The consistency and robustness of our technique are further validated by the reduction in pulse distortion and phase error observed across multilayer CZ gates. These results underscore the potential of our calibration and predistortion method to enhance the performance of two-qubit gates in superconducting quantum processors.

quant-ph↗

Optically Coherent Nitrogen-Vacancy Centers in HPHT Treated Diamonds

As a point defect with unique spin and optical properties, nitrogen-vacancy (NV) center in diamond has attracted much attention in the fields of quantum sensing, quantum simulation, and quantum networks. The optical properties of an NV center are crucial for all these quantum applications. However, NV centers fabricated by destructive methods such as electron irradiation or ion implantation usually exhibit poor optical coherence. In this work, we demonstrate a non-destructive method to fabricate optically coherent NV centers. High-purity single crystal diamonds are annealed under high pressure and high temperature (1700 $^{\circ}$C, 5.5 GPa), and individually resolvable NV centers with narrow PLE linewidth (<100 MHz) are produced. The high-pressure condition prevents the conversion of diamond to graphite during high-temperature annealing, significantly expanding the parameter space for creating high-performance artificial defects for quantum information science. These findings deepen our understanding of NV center formation in diamond and have implications for the optimization of color centers in solids, including silicon carbide and hexagonal boron nitride.

quant-ph↗

Dynamics of quantum coherence in many-body localized systems

We demonstrate that the dynamics of quantum coherence serves as an effective probe for identifying dephasing, which is a distinctive signature of many-body localization (MBL). Quantum coherence can be utilized to measure both the local coherence of specific subsystems and the total coherence of the whole system in a consistent manner. Our results reveal that the local coherence of small subsystems decays over time following a power law in the MBL phase, while it reaches a stable value within the same time window in the Anderson localized (AL) phase. In contrast, the total coherence of the whole system exhibits logarithmic growth during the MBL phase and reaches a stable value in the AL phase. Notably, this dynamic characteristic of quantum coherence remains robust even with weak interactions and displays unbounded behavior in infinite systems. Our results provide insights into understanding many-body dephasing phenomena in MBL systems and propose a novel feasible method for identifying and characterizing MBL phases in experiments.

quant-ph↗

Distributed Quantum Computation via Entanglement Forging and Teleportation

Distributed quantum computation is a practical method for large-scale quantum computation on quantum processors with limited size. It can be realized by direct quantum channels in flying qubits. Moreover, the pre-established quantum entanglements can also play the role of quantum channels with local operations and classical channels. However, without quantum correlations like quantum channels and entanglements, the entanglement forging technique allows us to classically forge the entangled states with local operations and classical channels only. In this paper, we demonstrate the methods to implement a nonlocal quantum circuit on two quantum processors without any quantum correlations, which is based on the fact that teleportation with classically forged Bell states is equivalent to quantum state tomography. In compensation, the overhead of single-shot measurement will increase, and several auxiliary qubits are required. Our results extend the possibility of integrating quantum processors. We expect that our methods will complement the toolbox of distributed quantum computation, and facilitate the extension of the scale of quantum computations.

quant-ph↗

Purity-Assisted Zero-Noise Extrapolation for Quantum Error Mitigation

Quantum error mitigation aims to reduce errors in quantum systems and improve accuracy. Zero-noise extrapolation (ZNE) is a commonly used method, where noise is amplified, and the target expectation is extrapolated to a noise-free point. However, ZNE relies on assumptions about error rates based on the error model. In this study, a purity-assisted zero-noise extrapolation (pZNE) method is utilized to address limitations in error rate assumptions and enhance the extrapolation process. The pZNE is based on the Pauli diagonal error model implemented using the Pauli twirling technique. Although this method does not significantly reduce the bias of routine ZNE, it extends its effectiveness to a wider range of error rates where routine ZNE may face limitations. In addition, the practicality of the pZNE method is verified through numerical simulations and experiments on the online quantum computation platform, Quafu. Comparisons with routine ZNE and virtual distillation methods show that biases in extrapolation methods increase with error rates and may become divergent at high error rates. The bias of pZNE is slightly lower than routine ZNE, while its error rate threshold surpasses that of routine ZNE. Furthermore, for full density matrix information, the pZNE method is more efficient than the routine ZNE.

quant-ph↗

Probing spin hydrodynamics on a superconducting quantum simulator

Characterizing the nature of hydrodynamical transport properties in quantum dynamics provides valuable insights into the fundamental understanding of exotic non-equilibrium phases of matter. Experimentally simulating infinite-temperature transport on large-scale complex quantum systems is of considerable interest. Here, using a controllable and coherent superconducting quantum simulator, we experimentally realize the analog quantum circuit, which can efficiently prepare the Haar-random states, and probe spin transport at infinite temperature. We observe diffusive spin transport during the unitary evolution of the ladder-type quantum simulator with ergodic dynamics. Moreover, we explore the transport properties of the systems subjected to strong disorder or a tilted potential, revealing signatures of anomalous subdiffusion in accompany with the breakdown of thermalization. Our work demonstrates a scalable method of probing infinite-temperature spin transport on analog quantum simulators, which paves the way to study other intriguing out-of-equilibrium phenomena from the perspective of transport.

quant-ph↗

SiCP: Simultaneous Individual and Cooperative Perception for 3D Object Detection in Connected and Automated Vehicles

Cooperative perception for connected and automated vehicles is traditionally achieved through the fusion of feature maps from two or more vehicles. However, the absence of feature maps shared from other vehicles can lead to a significant decline in 3D object detection performance for cooperative perception models compared to standalone 3D detection models. This drawback impedes the adoption of cooperative perception as vehicle resources are often insufficient to concurrently employ two perception models. To tackle this issue, we present Simultaneous Individual and Cooperative Perception (SiCP), a generic framework that supports a wide range of the state-of-the-art standalone perception backbones and enhances them with a novel Dual-Perception Network (DP-Net) designed to facilitate both individual and cooperative perception. In addition to its lightweight nature with only 0.13M parameters, DP-Net is robust and retains crucial gradient information during feature map fusion. As demonstrated in a comprehensive evaluation on the V2V4Real and OPV2V datasets, thanks to DP-Net, SiCP surpasses state-of-the-art cooperative perception solutions while preserving the performance of standalone perception solutions.

cs.CV↗

Beyond MOT: Semantic Multi-Object Tracking

Current multi-object tracking (MOT) aims to predict trajectories of targets (i.e., ''where'') in videos. Yet, knowing merely ''where'' is insufficient in many crucial applications. In comparison, semantic understanding such as fine-grained behaviors, interactions, and overall summarized captions (i.e., ''what'') from videos, associated with ''where'', is highly-desired for comprehensive video analysis. Thus motivated, we introduce Semantic Multi-Object Tracking (SMOT), that aims to estimate object trajectories and meanwhile understand semantic details of associated trajectories including instance captions, instance interactions, and overall video captions, integrating ''where'' and ''what'' for tracking. In order to foster the exploration of SMOT, we propose BenSMOT, a large-scale Benchmark for Semantic MOT. Specifically, BenSMOT comprises 3,292 videos with 151K frames, covering various scenarios for semantic tracking of humans. BenSMOT provides annotations for the trajectories of targets, along with associated instance captions in natural language, instance interactions, and overall caption for each video sequence. To our best knowledge, BenSMOT is the first publicly available benchmark for SMOT. Besides, to encourage future research, we present a novel tracker named SMOTer, which is specially designed and end-to-end trained for SMOT, showing promising performance. By releasing BenSMOT, we expect to go beyond conventional MOT by predicting ''where'' and ''what'' for SMOT, opening up a new direction in tracking for video understanding. We will release BenSMOT and SMOTer at https://github.com/Nathan-Li123/SMOTer.

cs.CV↗

Tracking Meets LoRA: Faster Training, Larger Model, Stronger Performance

Motivated by the Parameter-Efficient Fine-Tuning (PEFT) in large language models, we propose LoRAT, a method that unveils the power of large ViT model for tracking within laboratory-level resources. The essence of our work lies in adapting LoRA, a technique that fine-tunes a small subset of model parameters without adding inference latency, to the domain of visual tracking. However, unique challenges and potential domain gaps make this transfer not as easy as the first intuition. Firstly, a transformer-based tracker constructs unshared position embedding for template and search image. This poses a challenge for the transfer of LoRA, usually requiring consistency in the design when applied to the pre-trained backbone, to downstream tasks. Secondly, the inductive bias inherent in convolutional heads diminishes the effectiveness of parameter-efficient fine-tuning in tracking models. To overcome these limitations, we first decouple the position embeddings in transformer-based trackers into shared spatial ones and independent type ones. The shared embeddings, which describe the absolute coordinates of multi-resolution images (namely, the template and search images), are inherited from the pre-trained backbones. In contrast, the independent embeddings indicate the sources of each token and are learned from scratch. Furthermore, we design an anchor-free head solely based on MLP to adapt PETR, enabling better performance with less computational overhead. With our design, 1) it becomes practical to train trackers with the ViT-g backbone on GPUs with only memory of 25.8GB (batch size of 16); 2) we reduce the training time of the L-224 variant from 35.0 to 10.8 GPU hours; 3) we improve the LaSOT SUC score from 0.703 to 0.742 with the L-224 variant; 4) we fast the inference speed of the L-224 variant from 52 to 119 FPS. Code and models are available at https://github.com/LitingLin/LoRAT.

cs.CV↗

Recovery of damaged information via scrambling in indefinite casual order

Scrambling prevents the access to local information with local operators and therefore can be used to protect quantum information from damage caused by local perturbations. Even though partial quantum information can be recovered if the type of the damage is known, the initial target state cannot be completely recovered, because the obtained state is a mixture of the initial state and a maximally mixed state. Here, we demonstrate an improved scheme to recover damaged quantum information via scrambling in indefinite causal order. We show that scheme with indefinite causal order can record information of the damage and distill the initial state from the damaged state simultaneously. It allows us to retrieve initial information versus any damage. Moreover, by iterating the schemes, the initial quantum state can be completely recovered. In addition, we experimentally demonstrate our schemes on the cloud-based quantum computer, named as Quafu. Our work proposes a feasible scheme to protect whole quantum information from damage, which is also compatible with other techniques such as quantum error corrections and entanglement purification protocols. We expect that our scheme will be useful in the both quantum information recovery from the damage and systems bench-marking.

quant-ph↗

Cyclic Refiner: Object-Aware Temporal Representation Learning for Multi-View 3D Detection and Tracking

We propose a unified object-aware temporal learning framework for multi-view 3D detection and tracking tasks. Having observed that the efficacy of the temporal fusion strategy in recent multi-view perception methods may be weakened by distractors and background clutters in historical frames, we propose a cyclic learning mechanism to improve the robustness of multi-view representation learning. The essence is constructing a backward bridge to propagate information from model predictions (e.g., object locations and sizes) to image and BEV features, which forms a circle with regular inference. After backward refinement, the responses of target-irrelevant regions in historical frames would be suppressed, decreasing the risk of polluting future frames and improving the object awareness ability of temporal fusion. We further tailor an object-aware association strategy for tracking based on the cyclic learning model. The cyclic learning model not only provides refined features, but also delivers finer clues (e.g., scale level) for tracklet association. The proposed cycle learning method and association module together contribute a novel and unified multi-task framework. Experiments on nuScenes show that the proposed model achieves consistent performance gains over baselines of different designs (i.e., dense query-based BEVFormer, sparse query-based SparseBEV and LSS-based BEVDet4D) on both detection and tracking evaluation.

cs.CV↗

LaMOT: Language-Guided Multi-Object Tracking

Vision-Language MOT is a crucial tracking problem and has drawn increasing attention recently. It aims to track objects based on human language commands, replacing the traditional use of templates or pre-set information from training sets in conventional tracking tasks. Despite various efforts, a key challenge lies in the lack of a clear understanding of why language is used for tracking, which hinders further development in this field. In this paper, we address this challenge by introducing Language-Guided MOT, a unified task framework, along with a corresponding large-scale benchmark, termed LaMOT, which encompasses diverse scenarios and language descriptions. Specially, LaMOT comprises 1,660 sequences from 4 different datasets and aims to unify various Vision-Language MOT tasks while providing a standardized evaluation platform. To ensure high-quality annotations, we manually assign appropriate descriptive texts to each target in every video and conduct careful inspection and correction. To the best of our knowledge, LaMOT is the first benchmark dedicated to Language-Guided MOT. Additionally, we propose a simple yet effective tracker, termed LaMOTer. By establishing a unified task framework, providing challenging benchmarks, and offering insights for future algorithm design and evaluation, we expect to contribute to the advancement of research in Vision-Language MOT. We will release the data at https://github.com/Nathan-Li123/LaMOT.

cs.CV↗