Searcharxiv⌕ Search

arXiv subjects

Heng Fan

Publications and source records attributed to Heng Fan.

At least 37 records · Page 2Linked to original sources

Towards Long-Form Spatio-Temporal Video Grounding

In real scenarios, videos can span several minutes or even hours. However, existing research on spatio-temporal video grounding (STVG), given a textual query, mainly focuses on localizing targets in short videos of tens of seconds, typically less than one minute, which limits real-world applications. In this paper, we explore Long-Form STVG (LF-STVG), which aims to locate targets in long-term videos. Compared with short videos, long-term videos contain much longer temporal spans and more irrelevant information, making it difficult for existing STVG methods that process all frames at once. To address this challenge, we propose an AutoRegressive Transformer architecture for LF-STVG, termed ART-STVG. Unlike conventional STVG methods that require the entire video sequence to make predictions at once, ART-STVG treats the video as streaming input and processes frames sequentially, enabling efficient handling of long videos. To model spatio-temporal context, we design spatial and temporal memory banks and apply them to the decoders. Since memories from different moments are not always relevant to the current frame, we introduce simple yet effective memory selection strategies to provide more relevant information to the decoders, significantly improving performance. Furthermore, instead of parallel spatial and temporal localization, we propose a cascaded spatio-temporal design that connects the spatial decoder to the temporal decoder, allowing fine-grained spatial cues to assist complex temporal localization in long videos. Experiments on newly extended LF-STVG datasets show that ART-STVG significantly outperforms state-of-the-art methods, while achieving competitive performance on conventional short-form STVG. Our code is at: https://github.com/HengLan/ART-STVG.

cs.CV↗

Operation Mpemba effect: Breakdown of resource-Markovianity of free dynamics

The Mpemba effect refers to faster relaxation of states that are initially farther from equilibrium, yet its characterization is often tied to a chosen distance or resource measure. We introduce resource-Markovianity, an extended concept of quantum Markovianity to quantum resource theories, and formulate the resource Mpemba effect operationally as the breaking of resource-Markovianity by a relaxation operation. This yields a measure-independent operational characterization of resource Mpemba effects in general resource theories, together with quantitative characterizations based on resource-non-Markovianity measures. We illustrate the framework with the Mpemba effect for distinguishability of states, due to its relation to quantum Markovianity, and with the thermomajorization Mpemba effect from an operational perspective. These results reveal a deep interplay between quantum resources, non-Markovianity, and the Mpemba effect.

quant-ph↗

Programmable spectral symmetries in an anisotropic quantum Rabi simulator

The quantum Rabi model captures fundamental aspects of light--matter interaction, where symmetry dictates both spectra and dynamics. Over the past years, experiments have explored many of its nonperturbative properties, but have mostly focused on the isotropic limit, where rotating and counterrotating processes are locked together, leaving the broader symmetry landscape largely unexplored. Here we realize a programmable anisotropic quantum Rabi model in a superconducting processor, with independent control of the rotating and counterrotating couplings $(g_1,g_2)$ and of a transverse bias $\varepsilon$. Continuous anisotropy tuning, combined with a duality mapping, gives access to the full parameter space from the Jaynes-Cummings to the anti-Jaynes-Cummings limits. In the deep-strong-coupling regime, we show that anisotropy reconstructs the spectrum and turns complete collapse-revival dynamics into incomplete revivals even near degeneracy. With adiabatic state preparation and joint tomography, we resolve an anisotropy-induced ground-state parity switch, a crossing that has no analogue in the isotropic model. We further observe selective tunnelling associated with hidden symmetry in biased Rabi models and track its anisotropic displacement within the same device. These results establish a controllable route to engineering nonperturbative light--matter Hamiltonians, where symmetry, spectrum, and dynamics can be programmed independently.

quant-ph↗

Off-resonant preservation and generation of imaginarity in distributed scenarios

We study the nonlocal advantage of quantum imaginarity (NAQI) and distillable imaginarity of assistance (DIA), which treat imaginarity as a resource in distributed scenarios. For two qubits interacting with a lossy cavity, it is shown that both the NAQI and DIA can be well preserved for long times in the presence of large and symmetric detuning between the qubits and the cavity. Moreover, the off-resonant interaction generates a high degree of NAQI and DIA from the initial product states of two qubits having the same detunings and unequal couplings to the cavity. Based on the effective coupling of the qubits induced by the cavity mode, we explain the physical mechanism underlying the validity of this strategy. Our findings shed light on the role that off-resonant interactions have in the efficient control of imaginarity in distributed scenarios.

quant-ph↗

Observation and Modulation of the Quantum Mpemba Effect on a Superconducting Quantum Processor

In non-equilibrium quantum systems, the quantum Mpemba effect (QME) emerges as a counterintuitive phenomenon: systems exhibiting greater initial symmetry breaking restore symmetry faster. It has been attracting broad interest in studying QME dynamics and potential applications in quantum information science. While theoretical exploration of QME has surged, experimental studies, specifically on its flexible modulation, remain limited. Here, we report the observation and modulation of QME using a superconducting processor featuring an all-to-all connected, tunable-coupling architecture that enables precise control from short- to long-range interactions. This platform allows independent manipulation of coupling regimes, on-site potentials, and initial states, enabling us to elucidate their roles in QME. To quantify symmetry restoration, we employ entanglement asymmetry (EA), derived from the reconstructed density matrix via quantum state tomography, as a sensitive probe. In strong short-range coupling regimes, EA crossovers during quenches from tilted Néel states confirm the presence of QME. In contrast, in intermediate coupling regimes, synchronized EA and entanglement entropy dynamics reveal the suppression of QME. Remarkably, QME reemerges with the introduction of on-site linear potentials or quenches from tilted ferromagnetic states, the latter proving robust against on-site disorder. Our study demonstrates flexible QME modulation on a superconducting platform with multiple controllable parameters, shedding light on quantum many-body non-equilibrium dynamics and opening avenues for quantum information applications.

quant-ph↗

LoRe: Adaptive Interaction-Evaluation Routing with Per-Step Interaction Budgets for Iterative Graph Solvers

Diffusion-based neural solvers for combinatorial optimization repeatedly re-evaluate dense edge/factor interactions, making inference expensive in wall-clock time and often memory-bound at scale. Inspired by the computational methodologies of many-body physics, we introduce LoRe, a training-free, inference-time drop-in wrapper that enforces per-step interaction-evaluation budgeting: at each iteration, it evaluates only a fixed fraction of interactions by dynamically routing computation to high-conflict or high-uncertainty interactions, instead of using a fixed sparsification (e.g., static kNN graphs or static masks). Under fully inclusive end-to-end wall-clock accounting, LoRe substantially improves scalability on the Maximum Independent Set (MIS) problem, extending feasible inference more than $3\times$ beyond the baseline's out-of-memory limit, delivering a $\sim 8\times$ speedup and a $\sim 12\times$ peak-memory reduction, with solution quality preserved in this regime. Demonstrating cross-task generality on the large-scale Traveling Salesperson Problem (TSP) and zero-shot robustness to topology shifts, LoRe achieves a $\sim 15\times$ speedup at $n=1000$ with a $44\times$ memory reduction and competitive tour quality.

cs.LG↗

Dissipation-Selected Resonant Fronts in a Driven-Dissipative Bose-Hubbard Lattice

Spatially structured dissipation organizes driven quantum matter beyond Hamiltonian control. We show that a dissipation gradient combined with a Stark-induced detuning ramp selects a nonlinear resonance slice in a two-dimensional driven-dissipative Bose-Hubbard lattice, producing a pinned density front in generalized Gross-Pitaevskii simulations. The underlying resonance condition fixes the front position, while its Airy-like profile obeys a width scaling set by tunneling stiffness and the effective detuning slope. Treating the front as an emergent interface explains how tuning the selected resonance toward the minimum-loss side yields Peierls-Nabarro depinning steps, discrete transverse pattern locking, spatiotemporal chaos, and minimum-loss localization. Center-of-mass and generalized-imbalance diagnostics map these outcomes into a dynamical phase diagram as detuning-ramp slope and dissipation-gradient strength vary. The results suggest structured dissipation as a mechanism for reconfigurable transport barriers and nonequilibrium interfaces in programmable bosonic lattices.

cond-mat.quant-gas↗

Interference-Induced Suppression of Doublon Transport and Prethermalization in the Extended Bose-Hubbard Model

The coherent mobility of doublons, arising from second-order virtual dissociation-recombination processes, fundamentally limits their use as information carriers in the strongly interacting Bose-Hubbard model. We propose a disorder-free suppression mechanism by introducing an optimized nearest-neighbor pair-hopping term that destructively interferes with the dominant virtual hopping channel. Using the third-order Schrieffer-Wolff transformation, we derive an analytical optimal condition that accounts for lattice geometry corrections. Exact numerical simulations demonstrate that this optimized scheme achieves near-complete dynamical arrest and entanglement preservation in one-dimensional chains, while in two-dimensional square lattices, it significantly suppresses ballistic spreading yet permits a slow residual expansion. Furthermore, in the many-body regime, finite-size scaling analysis identifies the observed long-lived density-wave order as a prethermal plateau emerging from the dramatic separation of microscopic and thermalization timescales.

cond-mat.quant-gas↗

A Qudit-native Framework for Discrete Time Crystals

We introduce a qudit-native framework for engineering rich and robust discrete time crystals (DTCs) by leveraging their internal multilevel structure. Unlike in qubit systems, qudit-based DTCs exhibit distinct dynamical mechanisms that arise only in multilevel systems, as supported by a dressed normal-form analysis in the heating-suppression regime. These mechanisms are manifested in representative systems: we show that subspace-selective embedded kicks stabilize higher-order subharmonic responses and suppress thermalization, as demonstrated in spin-1 chains; in spin-3/2 systems, extending embedded kicks to more levels enables different level partitions and reveals that DTC robustness is dictated by the symmetry of the partition; and in spin-2 platforms, we realize concurrent 2T and 3T DTCs under a unified drive. These findings establish a systematic, hardware-efficient methodology for designing stable and multifunctional Floquet phases of matter on modern qudit-based quantum processors.

quant-ph↗

Probe of Generic Quantum Contextuality and Nonlocal Resources for Qubits

We reveal that the entropic uncertainty relation with a quantum memory is able to intrinsically connect local generic contextuality addressed in the pioneering work by Spekkens and nonlocal quantum resources such as entanglement and Bell nonlocality. Based on the constructed optimal set for any given single-qubit state, we prove rigorously a faithful criterion to witness the generic contextuality in the scenario of local quantum state preparation. Furthermore, within the framework of quantum resource distribution, it is proved that there exist quantitative trade-off relations between local preparation contextuality and bipartite entanglement or Bell nonlocality in a shared quantum system, which are captured by two inequalities where the local and nonlocal quantum resources can coexist. The faithful criterion and quantitative inequalities are all experimentally testable, which are verified through two independent well-designed experiments on the Quafu quantum cloud platform.

quant-ph↗

Multimode Purcell Filter for Superconducting-Qubit Reset and Readout with Intrinsic Purcell Protection

Efficient qubit reset and leakage reduction are essential for scalable superconducting quantum computing, particularly in the context of quantum error correction. However, such operations often require additional on-chip components. Here, we propose and experimentally demonstrate a hardware-efficient approach to qubit reset and readout using a multi-mode Purcell filter in a superconducting quantum circuit. We exploit the inherent multi-mode structure of a coplanar waveguide resonator, using its fundamental and second-order modes for qubit reset and readout, respectively, thereby avoiding additional components. Implemented in a flip-chip architecture, our device achieves unconditional reset with residual excitation below 1\% in 220 ns, and a leakage reduction unit that selectively resets the second excited state within 62 ns with a residual $|f\rangle$ population of 6.1\%, accounting for the readout error. Despite the qubits being directly coupled to the filter in our configuration, the measured relaxation times are not degraded owing to intrinsic Purcell protection provided by an auxiliary mode. To our knowledge, this is the first experimental trial that exploits different-order modes of a microwave resonator for distinct qubit operations, representing a new direction toward scalable, hardware-efficient quantum processor design.

quant-ph↗

Generalized Entanglement of Purification Criteria for 2-Producible States in Multipartite Systems

Multipartite entanglement has a much more complex structure than bipartite entanglement. A state that lacks generic multipartite entanglement is 2-producible, i.e. it can be written as a tensor product of at most 2-partite entangled states. Recently, it has been proved that a tripartite pure state is 2-producible if and only if the gap between the entanglement of purification and its lower bound vanishes. Here, we show that the entanglement of purification gap is insufficient to detect more than tripartite entanglement in 4-partite stabilizer states. We then generalize entanglement of purification to the multipartite cases, and demonstrate that a multipartite pure state is 2-producible if and only if all the generalized entanglement of purification gaps vanish. The generalized entanglement of purification gap quantifies the quantum communication cost for redistributing one part of the system to the others, and also relates to the local recoverability of a multipartite state and the relative entropy between that state and 2-producible states. Moreover, we calculate the generalized entanglement of purification for states satisfying the general Schmidt decomposition, which implies that 4-partite stabilizer states do not necessarily have a general Schmidt decomposition. Our results provide a quantitative characterization of multipartite entanglement in multipartite system, which will promote further investigations and understanding of multipartite entanglement.

quant-ph↗

Exploring Hilbert-Space Fragmentation on a Superconducting Processor

Isolated interacting quantum systems generally thermalize, yet there are several examples for the breakdown of ergodicity, such as many-body localization and quantum scars. Recently, ergodicity breaking has been observed in systems subjected to linear potentials, termed Stark many-body localization. This phenomenon is closely associated with Hilbert-space fragmentation, characterized by a strong dependence of dynamics on initial conditions. Here, we explore initial-state dependent dynamics using a ladder-type superconducting processor with up to 24 qubits, which enables precise control of the qubit frequency and initial state preparation. In systems with linear potentials, we experimentally observe distinct non-equilibrium dynamics for initial states with the same quantum numbers and energy, but with varying domain wall numbers. Accompanied by the numerical simulation for systems with larger sizes, we reveal that this distinction becomes increasingly pronounced as the system size grows, in contrast with weakly disordered interacting systems. Our results provide convincing experimental evidence of the fragmentation in Stark systems, enriching our understanding of the weak breakdown of ergodicity.

quant-ph↗

Robust Ego-Exo Correspondence with Long-Term Memory

Establishing object-level correspondence between egocentric and exocentric views is essential for intelligent assistants to deliver precise and intuitive visual guidance. However, this task faces numerous challenges, including extreme viewpoint variations, occlusions, and the presence of small objects. Existing approaches usually borrow solutions from video object segmentation models, but still suffer from the aforementioned challenges. Recently, the Segment Anything Model 2 (SAM 2) has shown strong generalization capabilities and excellent performance in video object segmentation. Yet, when simply applied to the ego-exo correspondence (EEC) task, SAM 2 encounters severe difficulties due to ineffective ego-exo feature fusion and limited long-term memory capacity, especially for long videos. Addressing these problems, we propose a novel EEC framework based on SAM 2 with long-term memories by presenting a dual-memory architecture and an adaptive feature routing module inspired by Mixture-of-Experts (MoE). Compared to SAM 2, our approach features (i) a Memory-View MoE module which consists of a dual-branch routing mechanism to adaptively assign contribution weights to each expert feature along both channel and spatial dimensions, and (ii) a dual-memory bank system with a simple yet effective compression strategy to retain critical long-term information while eliminating redundancy. In the extensive experiments on the challenging EgoExo4D benchmark, our method, dubbed LM-EEC, achieves new state-of-the-art results and significantly outperforms existing methods and the SAM 2 baseline, showcasing its strong generalization across diverse scenarios. Our code and model are available at https://github.com/juneyeeHu/LM-EEC.

cs.CV↗

Observing Quantum Correlation Dynamics in Tunable Superconducting Bose-Hubbard Simulators

The dynamics of quantum correlations are central to understanding many physical properties of quantum systems. Here we experimentally study the correlation dynamics via two-particle quantum walks in superconducting Bose-Hubbard qutrit arrays, with tunable on-site interaction $U$ realized by Floquet engineering. Quantum walks show the characteristic change from bosonic bunching to fermionic antibunching with increasing $U$. The two-site entanglement and quantum correlation dynamics, as measured by negativity and quantum discord, are investigated. We find that depending on the initial state, the propagation of entanglement can be strongly suppressed with increasing $U$, while that of quantum discord exhibits considerably larger amplitude; or both of them appear insensitive to $U$. Furthermore, the forms of entanglement are found to persist throughout particle walks for $U =$ 0 and it is generally not the case when $U$ increases. Our work highlights the role of interaction in shaping quantum dynamics and extends the realm of simulating correlated quantum systems with superconducting circuits.

quant-ph↗

DiveUp: Learning Feature Upsampling from Diverse Vision Foundation Models

Recently, feature upsampling has gained increasing attention owing to its effectiveness in enhancing vision foundation models (VFMs) for pixel-level understanding tasks. Existing methods typically rely on high-resolution features from the same foundation model to achieve upsampling via self-reconstruction. However, relying solely on intra-model features forces the upsampler to overfit to the source model's inherent location misalignment and high-norm artifacts. To address this fundamental limitation, we propose DiveUp, a novel framework that breaks away from single-model dependency by introducing multi-VFM relational guidance. Instead of naive feature fusion, DiveUp leverages diverse VFMs as a panel of experts, utilizing their structural consensus to regularize the upsampler's learning process, effectively preventing the propagation of inaccurate spatial structures from the source model. To reconcile the unaligned feature spaces across different VFMs, we propose a universal relational feature representation, formulated as a local center-of-mass (COM) field, that extracts intrinsic geometric structures, enabling seamless cross-model interaction. Furthermore, we introduce a spikiness-aware selection strategy that evaluates the spatial reliability of each VFM, effectively filtering out high-norm artifacts to aggregate guidance from only the most reliable expert at each local region. DiveUp is a unified, encoder-agnostic framework; a jointly-trained model can universally upsample features from diverse VFMs without requiring per-model retraining. Extensive experiments demonstrate that DiveUp achieves state-of-the-art performance across various downstream dense prediction tasks, validating the efficacy of multi-expert relational guidance. Our code and models are available at: https://github.com/Xiaoqiong-Liu/DiveUp

cs.CV↗

Demonstration of High-Fidelity Gates in a Strongly Anharmonic with Long-Coherence C-Shunt Flux Qubit

We demonstrate high-fidelity single-qubit gates on a C-shunt flux qubit that simultaneously combines a large anharmonicity ($\mathcal{A}/2π=848~\mathrm{MHz}$) with long relaxation time ($T_1 = 23~μ\text{s}$). The large anharmonicity significantly suppresses leakage to higher energy levels, enabling fast and precise microwave control. Using DRAG pulses and randomized benchmarking, the qubit achieves gate fidelities exceeding 99.9\%, highlighting the capability of C-shunt flux qubits for robust and high-performance quantum operations. These results establish them as a promising platform for scalable quantum information processing.

quant-ph↗

Towards Visual Query Segmentation in the Wild

In this paper, we introduce visual query segmentation (VQS), a new paradigm of visual query localization (VQL) that aims to segment all pixel-level occurrences of an object of interest in an untrimmed video, given an external visual query. Compared to existing VQL locating only the last appearance of a target using bounding boxes, VQS enables more comprehensive (i.e., all object occurrences) and precise (i.e., pixel-level masks) localization, making it more practical for real-world scenarios. To foster research on this task, we present VQS-4K, a large-scale benchmark dedicated to VQS. Specifically, VQS-4K contains 4,111 videos with more than 1.3 million frames and covers a diverse set of 222 object categories. Each video is paired with a visual query defined by a frame outside the search video and its target mask, and annotated with spatial-temporal masklets corresponding to the queried target. To ensure high quality, all videos in VQS-4K are manually labeled with meticulous inspection and iterative refinement. To the best of our knowledge, VQS-4K is the first benchmark specifically designed for VQS. Furthermore, to stimulate future research, we present a simple yet effective method, named VQ-SAM, which extends SAM 2 by leveraging target-specific and background distractor cues from the video to progressively evolve the memory through a novel multi-stage framework with an adaptive memory generation (AMG) module for VQS, significantly improving the performance. In our extensive experiments on VQS-4K, VQ-SAM achieves promising results and surpasses all existing approaches, demonstrating its effectiveness. With the proposed VQS-4K and VQ-SAM, we expect to go beyond the current VQL paradigm and inspire more future research and practical applications on VQS. Our benchmark, code, and results will be made publicly available.

cs.CV↗