SearcharxivSearch

arXiv subjects

Zixu Zhao

Publications and source records attributed to Zixu Zhao.

At least 19 recordsLinked to original sources

Entanglement dynamics of accelerated atoms with environment-induced interactions

We investigate the influence of environment-induced interactions on the entanglement dynamics of two uniformly accelerated atoms coupled to a fluctuating massless scalar field with a reflecting boundary. The two atoms are aligned vertically to the boundary. The entanglement behaviors are influenced by the competition between the environment-induced interatomic and the environment-induced atom-plate interactions, which can be characterized by certain critical values. The maximum of concurrence generated during evolution decreases non-monotonically with the acceleration, which implies the anti-Unruh phenomenon can exist for some situations even when both environment-induced interatomic and the environment-induced atom-plate interactions are considered.

quant-ph

TRACE: Tourism Recommendation with Accountable Citation Evidence

Tourism is a high-stakes setting for conversational recommender systems (CRS): a plausible-sounding suggestion can waste real money and trip time once a traveler acts on it. Existing CRS benchmarks primarily evaluate systems with a single Recall@k score over entity mentions, and tourism-specific resources add spatial or knowledge-graph context, yet none of them couple multi-turn recommendation with verbatim review-span evidence and rejection recovery. This leaves an evaluation gap for tourism recommendation that is simultaneously trustworthy, verifiable, and adaptive: recommend the right point of interest (POI) for multi-aspect preferences (such as cuisine, price, atmosphere, walking distance), justify each suggestion with verifiable evidence from prior visitors so the traveler can act without trial and error, and recover when the first recommendation is rejected mid-dialogue. We introduce TRACE, where each item is a multi-turn tourism recommendation dialogue with review-span citations and explicit rejection turns: 10,000 dialogues over 2,400 Yelp POIs and 34,208 reviews across eight U.S. cities, paired with 14 retrieval, planning, and LLM baselines, along with 25 metrics organized under Accuracy, Grounding, and Recovery. Across these baselines, TRACE reveals the Three-Competency Gap: LLM Zero-Shot leads in closed-set Recall@1 and rejection recovery but cites less densely than retrievers; non-LLM retrievers achieve surface-verbatim grounding but with low accuracy; Multi-Review Synthesis fails at recovery. The Grounding Score agrees with human citation precision (Spearman rho=+0.80, p<10^-20), and paired t-tests reproduce the per-baseline ranking (p<0.01 on the dominant contrasts). TRACE reframes accountable tourism recommendation as a joint target (right POI, verifiable evidence, adaptive repair) rather than a single-axis leaderboard.

cs.IR

Mutual information harvesting for circularly accelerated detectors

We investigate the mutual information harvesting of two circularly accelerated detectors that interact with the massless scalar fields near a reflecting boundary. We consider that the two detectors share a common rotational axis with the same acceleration and trajectory radius. As the interdetector separation increases, the mutual information may exhibit oscillatory behavior at large acceleration and small radius. For a fixed radius, a larger acceleration leads to a larger peak value of the mutual information. Near the boundary, the mutual information may oscillate and the maximum can be obtained. As the acceleration increases, the mutual information in a small interdetector separation first increases and then decreases. For an intermediate interdetector separation, the mutual information may oscillate with the increase of acceleration. For a not large interdetector separation, when we take large acceleration and small radius, as the energy gap increases, the mutual information first decreases, then oscillates, and finally goes to zero. The combination of large acceleration and small radius corresponds to the fast rotation, which significantly modifies the vacuum fluctuations of the field, leading to the oscillatory behavior. Furthermore, the oscillation intensifies near the boundary, which indicates that it is related to the coherent superposition of boundary reflections.

quant-ph

TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval

Unmanned aerial vehicles (UAVs) have become powerful platforms for real-time, high-resolution data collection, producing massive volumes of aerial videos. Efficient retrieval of relevant content from these videos is crucial for applications in urban management, emergency response, security, and disaster relief. While text-video retrieval has advanced in natural video domains, the UAV domain remains underexplored due to limitations in existing datasets, such as coarse and redundant captions. Thus, in this work, we construct the Drone Video-Text Match Dataset (DVTMD), which contains 2,864 videos and 14,320 fine-grained, semantically diverse captions. The annotations capture multiple complementary aspects, including human actions, objects, background settings, environmental conditions, and visual style, thereby enhancing text-video correspondence and reducing redundancy. Building on this dataset, we propose the Text-Conditioned Multi-granularity Alignment (TCMA) framework, which integrates global video-sentence alignment, sentence-guided frame aggregation, and word-guided patch alignment. To further refine local alignment, we design a Word and Patch Selection module that filters irrelevant content, as well as a Text-Adaptive Dynamic Temperature Mechanism that adapts attention sharpness to text type. Extensive experiments on DVTMD and CapERA establish the first complete benchmark for drone text-video retrieval. Our TCMA achieves state-of-the-art performance, including 45.5% R@1 in text-to-video and 42.8% R@1 in video-to-text retrieval, demonstrating the effectiveness of our dataset and method. The code and dataset will be released.

cs.CV

Effect of environment-induced interatomic interaction on entanglement generation for uniformly accelerated atoms with a boundary

Considering environment-induced interatomic interaction, we study the entanglement dynamics of two uniformly accelerated atoms that interact with fluctuating massless scalar fields in the Minkowski vacuum in the presence of a reflecting boundary. The two atoms are initially prepared in a state such that one is in the ground state and the other is in the excited state, which is separable. When the acceleration is small, the rate of entanglement generation at the initial time and the maximum of concurrence generated during evolution oscillate with the distance between the atoms and the boundary before reaching a stable value, and may decrease non-monotonically with the acceleration, which means the anti-Unruh phenomenon can exist for some situations even when environmental considerations are taken into account. The results show that there exists the competition of the vacuum fluctuations caused by the boundary and the acceleration. In addition, the time evolution of concurrence will not be affected by the environment-induced interatomic interaction under certain conditions. For a larger acceleration, when the environment-induced interatomic interaction is considered, the concurrence may disappear later compared with the result when the environment-induced interatomic interaction is neglected.

quant-ph

Entanglement harvesting of circularly accelerated detectors with a reflecting boundary

We study the properties of the transition probability for a circularly accelerated detector that interacts with the massless scalar fields in the presence of a reflecting boundary. As the trajectory radius increases, the transition probability may exhibit some peaks under specific conditions, which can lead to the possibility of the identical result for different trajectory radius with the same acceleration and energy gap. These behaviors can be characterized by certain critical values. Furthermore, we analyze the entanglement harvesting phenomenon for two circularly accelerated detectors with a boundary. We consider that the two detectors are rotating around a common axis with the same acceleration, trajectory radius and angular velocity. When the detectors are close to the boundary, there may exist two peaks in entanglement harvesting. Interestingly, as trajectory radius increases, entanglement harvesting in some situations first decreases to zero, then remains zero, and finally increases to a stable value. Our results show that the features observed properties of the detectors are closely related to the vacuum fluctuations of the field and states of the detectors. Under appropriate conditions, circular motion can be used to simulate the results of uniform acceleration.

quant-ph

Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation

Recent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion and video editing. A specialized area within generative modeling is layout-to-image (L2I) generation, where predefined layouts of objects guide the generative process. In this study, we introduce a novel regional cross-attention module tailored to enrich layout-to-image generation. This module notably improves the representation of layout regions, particularly in scenarios where existing methods struggle with highly complex and detailed textual descriptions. Moreover, while current open-vocabulary L2I methods are trained in an open-set setting, their evaluations often occur in closed-set environments. To bridge this gap, we propose two metrics to assess L2I performance in open-vocabulary scenarios. Additionally, we conduct a comprehensive user study to validate the consistency of these metrics with human preferences.

cs.CV

VideoSAM: Open-World Video Segmentation

Video segmentation is essential for advancing robotics and autonomous driving, particularly in open-world settings where continuous perception and object association across video frames are critical. While the Segment Anything Model (SAM) has excelled in static image segmentation, extending its capabilities to video segmentation poses significant challenges. We tackle two major hurdles: a) SAM's embedding limitations in associating objects across frames, and b) granularity inconsistencies in object segmentation. To this end, we introduce VideoSAM, an end-to-end framework designed to address these challenges by improving object tracking and segmentation consistency in dynamic environments. VideoSAM integrates an agglomerated backbone, RADIO, enabling object association through similarity metrics and introduces Cycle-ack-Pairs Propagation with a memory mechanism for stable object tracking. Additionally, we incorporate an autoregressive object-token mechanism within the SAM decoder to maintain consistent granularity across frames. Our method is extensively evaluated on the UVO and BURST benchmarks, and robotic videos from RoboTAP, demonstrating its effectiveness and robustness in real-world scenarios. All codes will be available.

cs.CV

Unsupervised Open-Vocabulary Object Localization in Videos

In this paper, we show that recent advances in video representation learning and pre-trained vision-language models allow for substantial improvements in self-supervised video object localization. We propose a method that first localizes objects in videos via an object-centric approach with slot attention and then assigns text to the obtained slots. The latter is achieved by an unsupervised way to read localized semantic information from the pre-trained CLIP model. The resulting video object localization is entirely unsupervised apart from the implicit annotation contained in CLIP, and it is effectively the first unsupervised approach that yields good results on regular video benchmarks.

cs.CV

Object-Centric Multiple Object Tracking

Unsupervised object-centric learning methods allow the partitioning of scenes into entities without additional localization information and are excellent candidates for reducing the annotation burden of multiple-object tracking (MOT) pipelines. Unfortunately, they lack two key properties: objects are often split into parts and are not consistently tracked over time. In fact, state-of-the-art models achieve pixel-level accuracy and temporal consistency by relying on supervised object detection with additional ID labels for the association through time. This paper proposes a video object-centric model for MOT. It consists of an index-merge module that adapts the object-centric slots into detection outputs and an object memory module that builds complete object prototypes to handle occlusions. Benefited from object-centric learning, we only require sparse detection labels (0%-6.25%) for object localization and feature binding. Relying on our self-supervised Expectation-Maximization-inspired loss for object association, our approach requires no ID labels. Our experiments significantly narrow the gap between the existing object-centric model and the fully supervised state-of-the-art and outperform several unsupervised trackers.

cs.CV

Masked Vision and Language Pre-training with Unimodal and Multimodal Contrastive Losses for Medical Visual Question Answering

Medical visual question answering (VQA) is a challenging task that requires answering clinical questions of a given medical image, by taking consider of both visual and language information. However, due to the small scale of training data for medical VQA, pre-training fine-tuning paradigms have been a commonly used solution to improve model generalization performance. In this paper, we present a novel self-supervised approach that learns unimodal and multimodal feature representations of input images and text using medical image caption datasets, by leveraging both unimodal and multimodal contrastive losses, along with masked language modeling and image text matching as pretraining objectives. The pre-trained model is then transferred to downstream medical VQA tasks. The proposed approach achieves state-of-the-art (SOTA) performance on three publicly available medical VQA datasets with significant accuracy improvements of 2.2%, 14.7%, and 1.7% respectively. Besides, we conduct a comprehensive analysis to validate the effectiveness of different components of the approach and study different pre-training settings. Our codes and models are available at https://github.com/pengfeiliHEU/MUMC.

cs.CV

Thermodynamic phase transition of Euler-Heisenberg-AdS black hole on free energy landscape

We study the first order phase transition of Euler-Heisenberg-AdS black hole based on free energy landscape. By solving the Fokker-Planck equation, we research the probability distribution of the system states. The small (large) black hole can have the chance to switch to the large (small) black hole due to the change of the temperature $T$ or Euler-Heisenberg parameter $a$. A higher (lower) $T$ corresponds to a larger probability for a large (small) black hole. The coexistent small and large black hole states can be acquired for some conditions. For $0< a\leq \frac{32}{7} Q^2 $, the small-large black hole phase transition can be acquired with a small $a$. The probability of small (large) black holes will decrease to zero for a large $a$. For a small $a$, a higher peak of the first passage time can be acquired for higher (lower) $T$ or smaller (larger) $a$ with the initial small (large) black hole state. For $a<0$, a smaller (larger) $a$ corresponds to a larger probability for a large (small) black hole. A higher peak of the first passage time can also be obtained for higher (lower) $T$ or smaller (larger) $a$ with initial small (large) black hole state.

gr-qc

PointPatchMix: Point Cloud Mixing with Patch Scoring

Data augmentation is an effective regularization strategy for mitigating overfitting in deep neural networks, and it plays a crucial role in 3D vision tasks, where the point cloud data is relatively limited. While mixing-based augmentation has shown promise for point clouds, previous methods mix point clouds either on block level or point level, which has constrained their ability to strike a balance between generating diverse training samples and preserving the local characteristics of point clouds. Additionally, the varying importance of each part of the point clouds has not been fully considered, cause not all parts contribute equally to the classification task, and some parts may contain unimportant or redundant information. To overcome these challenges, we propose PointPatchMix, a novel approach that mixes point clouds at the patch level and integrates a patch scoring module to generate content-based targets for mixed point clouds. Our approach preserves local features at the patch level, while the patch scoring module assigns targets based on the content-based significance score from a pre-trained teacher model. We evaluate PointPatchMix on two benchmark datasets, ModelNet40 and ScanObjectNN, and demonstrate significant improvements over various baselines in both synthetic and real-world datasets, as well as few-shot settings. With Point-MAE as our baseline, our model surpasses previous methods by a significant margin, achieving 86.3% accuracy on ScanObjectNN and 94.1% accuracy on ModelNet40. Furthermore, our approach shows strong generalization across multiple architectures and enhances the robustness of the baseline model.

cs.CV

Estimation precision of the acceleration for a two-level atom coupled to fluctuating vacuum electromagnetic fields

In open quantum systems, we study the quantum Fisher information of acceleration for a uniformly accelerated two-level atom coupled to fluctuating electromagnetic fields in the Minkowski vacuum. With the time evolution, for the initial atom state parameter $θ\neqπ$, the quantum Fisher information can exist a maximum value and a local minimum value before reaching a stable value. In addition, in a short time, the quantum Fisher information varies with the initial state parameter, and the quantum Fisher information can take a maximum value at $θ=0$. The quantum Fisher information may exist two peak values at a certain moment. These features are different from the massless scalar fields case. With the time evolution, $F_{max}$ firstly increases, then decreases, and finally, reaches the same value. However, $F_{max}$ will arrive at a stable maximum value for the case of the massless scalar fields. Although the atom response to the vacuum fluctuation electromagnetic fields is different from the case of massless scalar fields, the quantum Fisher information eventually reaches a stable value.

quant-ph

Geometric phases acquired for a two-level atom coupled to fluctuating vacuum scalar fields due to linear acceleration and circular motion

In open quantum systems, we study the geometric phases acquired for a two-level atom coupled to a bath of fluctuating vacuum massless scalar fields due to linear acceleration and circular motion without and with a boundary. In free space, as we amplify acceleration, the geometric phase acquired purely due to linear acceleration case firstly is smaller than the circular acceleration case in the ultrarelativistic limit for the initial atomic state $θ\in(0,\fracπ{2})\cup(\fracπ{2},π)$, then equals to the circular acceleration case in a certain acceleration, and finally, is larger than the circular acceleration case. The spontaneous transition rates show a similar feature. This result is different from the case of a bath of fluctuating vacuum electromagnetic fields that has been studied. Considering the initial atomic state $θ\in(0,π)$, we find that the geometric phase acquired purely due to linear acceleration always equals to the circular acceleration case for the certain acceleration. The feature implies that, in a certain condition, one can simulate the case of the uniformly accelerated two-level atom by studying the properties of the two-level atom in circular motion. Adding a reflecting boundary, we observe that a larger value of a geometric phase can be obtained compared to the absence of a boundary. Besides, the geometric phase fluctuates along $z$, and the maximum of geometric phase is closer to the boundary for a larger acceleration. We also find that geometric phases can be acquired purely due to the linear acceleration case and circular acceleration case with $θ\in(0,π)$ for a smaller $z$.

quant-ph

Pseudo-label Guided Cross-video Pixel Contrast for Robotic Surgical Scene Segmentation with Limited Annotations

Surgical scene segmentation is fundamentally crucial for prompting cognitive assistance in robotic surgery. However, pixel-wise annotating surgical video in a frame-by-frame manner is expensive and time consuming. To greatly reduce the labeling burden, in this work, we study semi-supervised scene segmentation from robotic surgical video, which is practically essential yet rarely explored before. We consider a clinically suitable annotation situation under the equidistant sampling. We then propose PGV-CL, a novel pseudo-label guided cross-video contrast learning method to boost scene segmentation. It effectively leverages unlabeled data for a trusty and global model regularization that produces more discriminative feature representation. Concretely, for trusty representation learning, we propose to incorporate pseudo labels to instruct the pair selection, obtaining more reliable representation pairs for pixel contrast. Moreover, we expand the representation learning space from previous image-level to cross-video, which can capture the global semantics to benefit the learning process. We extensively evaluate our method on a public robotic surgery dataset EndoVis18 and a public cataract dataset CaDIS. Experimental results demonstrate the effectiveness of our method, consistently outperforming the state-of-the-art semi-supervised methods under different labeling ratios, and even surpassing fully supervised training on EndoVis18 with 10.1% labeling.

cs.CV

Exploring Intra- and Inter-Video Relation for Surgical Semantic Scene Segmentation

Automatic surgical scene segmentation is fundamental for facilitating cognitive intelligence in the modern operating theatre. Previous works rely on conventional aggregation modules (e.g., dilated convolution, convolutional LSTM), which only make use of the local context. In this paper, we propose a novel framework STswinCL that explores the complementary intra- and inter-video relations to boost segmentation performance, by progressively capturing the global context. We firstly develop a hierarchy Transformer to capture intra-video relation that includes richer spatial and temporal cues from neighbor pixels and previous frames. A joint space-time window shift scheme is proposed to efficiently aggregate these two cues into each pixel embedding. Then, we explore inter-video relation via pixel-to-pixel contrastive learning, which well structures the global embedding space. A multi-source contrast training objective is developed to group the pixel embeddings across videos with the ground-truth guidance, which is crucial for learning the global property of the whole data. We extensively validate our approach on two public surgical video benchmarks, including EndoVis18 Challenge and CaDIS dataset. Experimental results demonstrate the promising performance of our method, which consistently exceeds previous state-of-the-art approaches. Code is available at https://github.com/YuemingJin/STswinCL.

cs.CV

Excited states of holographic superconductors with hyperscaling violation

We employ the numerical and analytical methods to study the effects of the hyperscaling violation on the ground and excited states of holographic superconductors for $d=2,~z=2$. For both the holographic s-wave and p-wave models with the hyperscaling violation, we observe that the excited state has a lower critical temperature than the corresponding ground state, which is similar to the relativistic case, and the difference of the dimensionless critical chemical potential between the consecutive states decreases as the hyperscaling violation increases. Interestingly, as we amplify the hyperscaling violation in the s-wave model, the critical temperature of the ground state first decreases and then increases, but that of the excited states always decreases. In the p-wave model, regardless of the ground state or the excited states, the critical temperature always decreases with increasing the hyperscaling violation. In addition, we find that the hyperscaling violation affects the conductivity $σ$ which has $n$ peaks for the $n$-th excited state, and changes the relation in the gap frequency for the excited states in both s-wave and p-wave models.

hep-th