Searcharxiv⌕ Search

arXiv subjects

Heng Fan

Publications and source records attributed to Heng Fan.

At least 55 records · Page 3Linked to original sources

Towards Visual Query Segmentation in the Wild

In this paper, we introduce visual query segmentation (VQS), a new paradigm of visual query localization (VQL) that aims to segment all pixel-level occurrences of an object of interest in an untrimmed video, given an external visual query. Compared to existing VQL locating only the last appearance of a target using bounding boxes, VQS enables more comprehensive (i.e., all object occurrences) and precise (i.e., pixel-level masks) localization, making it more practical for real-world scenarios. To foster research on this task, we present VQS-4K, a large-scale benchmark dedicated to VQS. Specifically, VQS-4K contains 4,111 videos with more than 1.3 million frames and covers a diverse set of 222 object categories. Each video is paired with a visual query defined by a frame outside the search video and its target mask, and annotated with spatial-temporal masklets corresponding to the queried target. To ensure high quality, all videos in VQS-4K are manually labeled with meticulous inspection and iterative refinement. To the best of our knowledge, VQS-4K is the first benchmark specifically designed for VQS. Furthermore, to stimulate future research, we present a simple yet effective method, named VQ-SAM, which extends SAM 2 by leveraging target-specific and background distractor cues from the video to progressively evolve the memory through a novel multi-stage framework with an adaptive memory generation (AMG) module for VQS, significantly improving the performance. In our extensive experiments on VQS-4K, VQ-SAM achieves promising results and surpasses all existing approaches, demonstrating its effectiveness. With the proposed VQS-4K and VQ-SAM, we expect to go beyond the current VQL paradigm and inspire more future research and practical applications on VQS. Our benchmark, code, and results will be made publicly available.

cs.CV↗

DMTrack: Spatio-Temporal Multimodal Tracking via Dual-Adapter

In this paper, we explore adapter tuning and introduce a novel dual-adapter architecture for spatio-temporal multimodal tracking, dubbed DMTrack. The key of our DMTrack lies in two simple yet effective modules, including a spatio-temporal modality adapter (STMA) and a progressive modality complementary adapter (PMCA) module. The former, applied to each modality alone, aims to adjust spatio-temporal features extracted from a frozen backbone by self-prompting, which to some extent can bridge the gap between different modalities and thus allows better cross-modality fusion. The latter seeks to facilitate cross-modality prompting progressively with two specially designed pixel-wise shallow and deep adapters. The shallow adapter employs shared parameters between the two modalities, aiming to bridge the information flow between the two modality branches, thereby laying the foundation for following modality fusion, while the deep adapter modulates the preliminarily fused information flow with pixel-wise inner-modal attention and further generates modality-aware prompts through pixel-wise inter-modal attention. With such designs, DMTrack achieves promising spatio-temporal multimodal tracking performance with merely 0.93M trainable parameters. Extensive experiments on five benchmarks demonstrate that DMTrack achieves state-of-the-art results. Our code and models will be available at https://github.com/Nightwatch-Fox11/DMTrack.

cs.CV↗

TensorCircuit-NG: A Universal, Composable, and Scalable Platform for Quantum Computing and Quantum Simulation

We present TensorCircuit-NG, a next-generation quantum software platform designed to bridge the gap between quantum physics, artificial intelligence, and high-performance computing. Moving beyond the scope of traditional circuit simulators, TensorCircuit-NG establishes a unified, tensor-native programming paradigm where quantum circuits, tensor networks, and neural networks fuse into a single, end-to-end differentiable computational graph. Built upon industry-standard machine learning backends (JAX, TensorFlow, PyTorch), the framework introduces comprehensive capabilities for approximate circuit simulation, analog dynamics, fermion Gaussian states, qudit systems, and scalable noise modeling. To tackle the exponential complexity of deep quantum circuits, TensorCircuit-NG implements advanced distributed computing strategies, including automated data parallelism and model-parallel tensor network slicing. We validate these capabilities on GPU clusters, demonstrating a near-linear speedup in distributed variational quantum algorithms. TensorCircuit-NG enables flagship applications, including end-to-end QML for CIFAR-100 computer vision, efficient pipelines from quantum states to neural networks via classical shadows, and differentiable optimization of tensor network states for many-body physics.

quant-ph↗

Digital Quantum Simulation of the Lindblad Master Equation and Its Nonlinear Extensions via Quantum Trajectory Averaging

Since precisely controlling dissipation in realistic environments is challenging, digital simulation of the Lindblad master equation (LME) is of great significance for understanding nonequilibrium dynamics in open quantum systems. However, achieving long-time simulations for complex systems with multiple dissipation channels remains a major challenge, both theoretically and experimentally. Here, we propose a 1-dilation digital scheme for simulating the LME based on quantum trajectory averaging without postselection. By rigorously matching the stochasticity inherent in quantum trajectories with the probabilistic outcomes of quantum measurements, our method effectively translates the classically established quantum jump algorithm into executable quantum circuits. A key advantage of our method is that it overcomes the exponential suppression of success probability seen in some existing postselection-dependent schemes, especially for long-time evolution or systems with numerous jump operators. Moreover, the scheme can be extended to a 2-dilation framework for the nonlinear LME with postselection, bridging the full LME and non-Hermitian Hamiltonian dynamics. This extended scheme provides a digital approach for exploring the interplay between non-Hermitian Hamiltonians and dissipative terms within a monitored quantum dynamics framework.

quant-ph↗

QSteed: A Resource-Virtualized and Hardware-Aware Quantum Compilation Framework for Real Quantum Computing Processors

As quantum computing systems continue to scale up and become more clustered, efficiently compiling user quantum programs into high fidelity executable sequences on real hardware remains a key challenge for current quantum compilation systems. In this study, we introduce a system software framework that integrates resource virtualization and hardware aware compilation for real quantum computing processors, termed QSteed. QSteed virtualizes quantum processors through a four layer abstraction hierarchy comprising the Real Quantum Processing Unit (QPU), Standard QPU (StdQPU), Substructure of the QPU (SubQPU), and Virtual QPU (VQPU). These abstractions, together with calibration data, device topology, and noise descriptors, are maintained in a dedicated database to enable unified and fine grained management across superconducting quantum platforms. At run time, the modular compiler queries the database to match each incoming circuit with the most suitable VQPU, after which it confines layout, routing, gate resynthesis, and noise adaptive optimizations to that virtual subregion. The complete stack has been deployed on the Quafu superconducting cluster, where experimental runs confirm the correctness of the virtualization model and the efficacy of the compiler without requiring modifications to user code. By integrating resource virtualization with a select-then-compile workflow, QSteed demonstrates a robust architecture for compiling programs on noisy superconducting processors. This architectural approach offers a promising path towards efficient compilation needs across various superconducting quantum computing platforms in the noisy intermediate scale quantum (NISQ) era.

quant-ph↗

Markov Gap and Bound Entanglement in Haar Random State

Bound entanglement refers to entangled states that cannot be distilled into maximally entangled states and therefore cannot directly be used in many quantum information processing protocols. We identify a relationship between bound entanglement and the Markov gap, which is introduced within holography via the entanglement wedge cross section and is related to the fidelity of the partial Markov recovery problem. We prove that a bound entangled state must have a nonzero Markov gap. Conversely, for sufficiently large systems, a state with a weakly nonzero Markov gap typically has a bound entangled or separable marginal state, where entanglement is undistillable. Furthermore, this implies that the transition from a bound entangled to a separable state originates from the properties of states with a weakly nonzero Markov gap, which may be dual to non-perturbative effects from a holographic perspective. Our results shed light on the investigation of the Markov gap and enhance interdisciplinary applications of quantum information.

quant-ph↗

Mathieu Control of the Effective Coupling in Superconducting Qubits

A common challenge in superconducting quantum circuits is the trade-off between strong coupling and computational subspace integrity. We present Mathieu control, which uses a non-resonant two-photon drive to create a selective nonlinear frequency shift. This shift modifies interactions while preserving qubit states, enabling continuous tuning of the ZZ coupling, including full suppression, and integrating single- and two-qubit gates with low leakage. For a qubit-coupler-qubit device, it allows independent ZZ control, facilitating a programmable Heisenberg (XXZ) Hamiltonian. Extended to a five-qubit chain, the system can be reconfigured to simulate dynamics of quantum magnetic phases. Mathieu control thus provides a framework for high-fidelity quantum logic and programmable simulation.

quant-ph↗

OLC-WA: Drift Aware Tuning-Free Online Classification with Weighted Average

Real-world data sets often exhibit temporal dynamics characterized by evolving data distributions. Disregarding this phenomenon, commonly referred to as concept drift, can significantly diminish a model's predictive accuracy. Furthermore, the presence of hyperparameters in online models exacerbates this issue. These parameters are typically fixed and cannot be dynamically adjusted by the user in response to the evolving data distribution. This paper introduces Online Classification with Weighted Average (OLC-WA), an adaptive, hyperparameter-free online classification model equipped with an automated optimization mechanism. OLC-WA operates by blending incoming data streams with an existing base model. This blending is facilitated by an exponentially weighted moving average. Furthermore, an integrated optimization mechanism dynamically detects concept drift, quantifies its magnitude, and adjusts the model based on the observed data stream characteristics. This approach empowers the model to effectively adapt to evolving data distributions within streaming environments. Rigorous empirical evaluation across diverse benchmark datasets shows that OLC-WA achieves performance comparable to batch models in stationary environments, maintaining accuracy within 1-3%, and surpasses leading online baselines by 10-25% under drift, demonstrating its effectiveness in adapting to dynamic data streams.

cs.LG↗

Structured Context Learning for Generic Event Boundary Detection

Generic Event Boundary Detection (GEBD) aims to identify moments in videos that humans perceive as event boundaries. This paper proposes a novel method for addressing this task, called Structured Context Learning, which introduces the Structured Partition of Sequence (SPoS) to provide a structured context for learning temporal information. Our approach is end-to-end trainable and flexible, not restricted to specific temporal models like GRU, LSTM, and Transformers. This flexibility enables our method to achieve a better speed-accuracy trade-off. Specifically, we apply SPoS to partition the input frame sequence and provide a structured context for the subsequent temporal model. Notably, SPoS's overall computational complexity is linear with respect to the video length. We next calculate group similarities to capture differences between frames, and a lightweight fully convolutional network is utilized to determine the event boundaries based on the grouped similarity maps. To remedy the ambiguities of boundary annotations, we adapt the Gaussian kernel to preprocess the ground-truth event boundaries. Our proposed method has been extensively evaluated on the challenging Kinetics-GEBD, TAPOS, and shot transition detection datasets, demonstrating its superiority over existing state-of-the-art methods.

cs.CV↗

Stable and Efficient Charging of Superconducting Capacitively Shunted Flux Quantum Batteries

Quantum batteries, as miniature energy storage devices, have sparked significant research interest in recent years. However, achieving rapid and stable energy transfer in quantum batteries while obeying quantum speed limits remains a critical challenge. In this work, we experimentally optimize the charging process by leveraging the unique energy level structure of a superconducting capacitively-shunted flux qubit, using counterdiabatic pulses in the stimulated Raman adiabatic passage. Compared to previous studies, we impose two different norm constraints on the driving Hamiltonian, achieving optimal charging without exceeding the overall driving strength. Furthermore, we experimentally demonstrate a charging process that achieves the quantum speed limit. In addition, we introduce a dimensionless parameter $\mathcal{S}$ to unify charging speed and stability, offering a universal metric for performance optimization. In contrast to metrics such as charging power and thermodynamic efficiency, the $\mathcal{S}$ criterion quantitatively captures the stability of ergentropy while also considering the charging speed. Our results highlight the potential of the capacitively-shunted qubit platform as an ideal candidate for realizing three-level quantum batteries and deliver novel strategies for optimizing energy transfer protocols.

quant-ph↗

PlanarTrack: A high-quality and challenging benchmark for large-scale planar object tracking

Planar tracking has drawn increasing interest owing to its key roles in robotics and augmented reality. Despite recent great advancement, further development of planar tracking, particularly in the deep learning era, is largely limited compared to generic tracking due to the lack of large-scale platforms. To mitigate this, we propose PlanarTrack, a large-scale high-quality and challenging benchmark for planar tracking. Specifically, PlanarTrack consists of 1,150 sequences with over 733K frames, including 1,000 short-term and 150 new long-term videos, which enables comprehensive evaluation of short- and long-term tracking performance. All videos in PlanarTrack are recorded in unconstrained conditions from the wild, which makes PlanarTrack challenging but more realistic for real-world applications. To ensure high-quality annotations, each video frame is manually annotated by four corner points with multi-round meticulous inspection and refinement. To enhance target diversity of PlanarTrack, we only capture a unique target in one sequence, which is different from existing benchmarks. To our best knowledge, PlanarTrack is by far the largest and most diverse and challenging dataset dedicated to planar tracking. To understand performance of existing methods on PlanarTrack and to provide a comparison for future research, we evaluate 10 representative planar trackers with extensive comparison and in-depth analysis. Our evaluation reveals that, unsurprisingly, the top planar trackers heavily degrade on the challenging PlanarTrack, which indicates more efforts are required for improving planar tracking. Our data and results will be released at https://github.com/HengLan/PlanarTrack

cs.CV↗

All You Need is One: Capsule Prompt Tuning with a Single Vector

Prompt-based learning has emerged as a parameter-efficient finetuning (PEFT) approach to facilitate Large Language Model (LLM) adaptation to downstream tasks by conditioning generation with task-aware guidance. Despite its successes, current prompt-based learning methods heavily rely on laborious grid searching for optimal prompt length and typically require considerable number of prompts, introducing additional computational burden. Worse yet, our pioneer findings indicate that the task-aware prompt design is inherently limited by its absence of instance-aware information, leading to a subtle attention interplay with the input sequence. In contrast, simply incorporating instance-aware information as a part of the guidance can enhance the prompt-tuned model performance without additional fine-tuning. Moreover, we find an interesting phenomenon, namely "attention anchor", that incorporating instance-aware tokens at the earliest position of the sequence can successfully preserve strong attention to critical structural information and exhibit more active attention interaction with all input tokens. In light of our observation, we introduce Capsule Prompt-Tuning (CaPT), an efficient and effective solution that leverages off-the-shelf, informative instance semantics into prompt-based learning. Our approach innovatively integrates both instance-aware and task-aware information in a nearly parameter-free manner (i.e., one single capsule prompt). Empirical results demonstrate that our method can exhibit superior performance across various language tasks (e.g., 84.03\% average accuracy on T5-Large), serving as an "attention anchor," while enjoying high parameter efficiency (e.g., 0.003\% of model parameters on Llama3.2-1B).

cs.CL↗

Microwave-activated high-fidelity three-qubit gate scheme for fixed-frequency superconducting qubits

Scalable superconducting quantum processors require balancing critical constraints in coherence, control complexity, and spectral crowding. Fixed-frequency architectures suppress flux noise and simplify control via all-microwave operations but remain limited by residual ZZ crosstalk. Here we propose a microwave-activated three-qubit gate protocol for fixed-frequency transmon qubits in the large-detuning regime ($|Δ| \gg g$), leveraging the third-order nonlinear interaction to coherently exchange $\ket{001} \leftrightarrow \ket{110}$ states. By incorporating a phase-compensated optimization protocol, numerical simulations demonstrate a high average gate fidelity exceeding $99.9\%$. Systematic error analysis identifies static long-range ZZ coupling as the dominant error source in multi-qubit systems, which can be suppressed via operations in the large-detuning regime ($\sim 1$ GHz). The protocol maintains process fidelities exceeding $98\%$ under decoherence, while demonstrating intrinsic robustness to fabrication-induced parameter variations and compatibility with existing all-microwave two-qubit gate architectures. This hardware-efficient strategy advances scalable quantum computing systems by improving coherence properties, reducing spectral congestion, and expanding the experimental toolkit for error-resilient quantum operations in the noisy intermediate-scale quantum era.

quant-ph↗

Experimental Extraction of Coherent Ergotropy and Its Energetic Cost in a Superconducting Qubit

Quantum coherence, encoded in the off-diagonal elements of a system's density matrix, is a key resource in quantum thermodynamics, fundamentally limiting the maximum extractable work known as ergotropy. While previous experiments have isolated coherence-related contributions to work extraction, it remains unclear how coherence can be harnessed in a controllable and energy-efficient manner. Here, we experimentally investigate the role of initial-state coherence in work extraction from a superconducting transmon qubit. By preparing a variety of pure states and implementing three complementary extraction protocols, we reveal how coherence governs the partitioning of ergotropy. We find that the choice of initial state depends on the dominant decoherence channel-energy relaxation or dephasing. By further accounting for thermodynamic costs, we identify optimal initial states that maximize the efficiency. Our results demonstrate that the initial-state design provides a scalable approach to coherence control and advances the development of efficient quantum thermodynamic devices.

quant-ph↗

DP-GTR: Differentially Private Prompt Protection via Group Text Rewriting

Prompt privacy is crucial, especially when using online large language models (LLMs), due to the sensitive information often contained within prompts. While LLMs can enhance prompt privacy through text rewriting, existing methods primarily focus on document-level rewriting, neglecting the rich, multi-granular representations of text. This limitation restricts LLM utilization to specific tasks, overlooking their generalization and in-context learning capabilities, thus hindering practical application. To address this gap, we introduce DP-GTR, a novel three-stage framework that leverages local differential privacy (DP) and the composition theorem via group text rewriting. DP-GTR is the first framework to integrate both document-level and word-level information while exploiting in-context learning to simultaneously improve privacy and utility, effectively bridging local and global DP mechanisms at the individual data point level. Experiments on CommonSense QA and DocVQA demonstrate that DP-GTR outperforms existing approaches, achieving a superior privacy-utility trade-off. Furthermore, our framework is compatible with existing rewriting techniques, serving as a plug-in to enhance privacy protection. Our code is publicly available at github.com/ResponsibleAILab/DP-GTR.

cs.CL↗

Hybrid Quantum-Classical Neural Networks for Few-Shot Credit Risk Assessment

Quantum Machine Learning (QML) offers a new paradigm for addressing complex financial problems intractable for classical methods. This work specifically tackles the challenge of few-shot credit risk assessment, a critical issue in inclusive finance where data scarcity and imbalance limit the effectiveness of conventional models. To address this, we design and implement a novel hybrid quantum-classical workflow. The methodology first employs an ensemble of classical machine learning models (Logistic Regression, Random Forest, XGBoost) for intelligent feature engineering and dimensionality reduction. Subsequently, a Quantum Neural Network (QNN), trained via the parameter-shift rule, serves as the core classifier. This framework was evaluated through numerical simulations and deployed on the Quafu Quantum Cloud Platform's ScQ-P21 superconducting processor. On a real-world credit dataset of 279 samples, our QNN achieved a robust average AUC of 0.852 +/- 0.027 in simulations and yielded an impressive AUC of 0.88 in the hardware experiment. This performance surpasses a suite of classical benchmarks, with a particularly strong result on the recall metric. This study provides a pragmatic blueprint for applying quantum computing to data-constrained financial scenarios in the NISQ era and offers valuable empirical evidence supporting its potential in high-stakes applications like inclusive finance.

cs.LG↗

IRDFusion: Iterative Relation-Map Difference guided Feature Fusion for Multispectral Object Detection

Current multispectral object detection methods often retain extraneous background or noise during feature fusion, limiting perceptual performance. To address this, we propose an innovative feature fusion framework based on cross-modal feature contrastive and screening strategy, diverging from conventional approaches. The proposed method adaptively enhances salient structures by fusing object-aware complementary cross-modal features while suppressing shared background interference. Our solution centers on two novel, specially designed modules: the Mutual Feature Refinement Module (MFRM) and the Differential Feature Feedback Module (DFFM). The MFRM enhances intra- and inter-modal feature representations by modeling their relationships, thereby improving cross-modal alignment and discriminative power. Inspired by feedback differential amplifiers, the DFFM dynamically computes inter-modal differential features as guidance signals and feeds them back to the MFRM, enabling adaptive fusion of complementary information while suppressing common-mode noise across modalities. To enable robust feature learning, the MFRM and DFFM are integrated into a unified framework, which is formally formulated as an Iterative Relation-Map Differential Guided Feature Fusion mechanism, termed IRDFusion. IRDFusion enables high-quality cross-modal fusion by progressively amplifying salient relational signals through iterative feedback, while suppressing feature noise, leading to significant performance gains. In extensive experiments on FLIR, LLVIP and M$^3$FD datasets, IRDFusion achieves state-of-the-art performance and consistently outperforms existing methods across diverse challenging scenarios, demonstrating its robustness and effectiveness. Code will be available at https://github.com/61s61min/IRDFusion.git.

cs.CV↗

G3CN: Gaussian Topology Refinement Gated Graph Convolutional Network for Skeleton-Based Action Recognition

Graph Convolutional Networks (GCNs) have proven to be highly effective for skeleton-based action recognition, primarily due to their ability to leverage graph topology for feature aggregation, a key factor in extracting meaningful representations. However, despite their success, GCNs often struggle to effectively distinguish between ambiguous actions, revealing limitations in the representation of learned topological and spatial features. To address this challenge, we propose a novel approach, Gaussian Topology Refinement Gated Graph Convolution (G$^{3}$CN), to address the challenge of distinguishing ambiguous actions in skeleton-based action recognition. G$^{3}$CN incorporates a Gaussian filter to refine the skeleton topology graph, improving the representation of ambiguous actions. Additionally, Gated Recurrent Units (GRUs) are integrated into the GCN framework to enhance information propagation between skeleton points. Our method shows strong generalization across various GCN backbones. Extensive experiments on NTU RGB+D, NTU RGB+D 120, and NW-UCLA benchmarks demonstrate that G$^{3}$CN effectively improves action recognition, particularly for ambiguous samples.

cs.CV↗