Searcharxiv⌕ Search

arXiv subjects

Jun Feng

Publications and source records attributed to Jun Feng.

At least 37 records · Page 2Linked to original sources

Quantum thermodynamics in a rotating BTZ black hole spacetime

We address the problem of the thermalization process for an Unruh-DeWitt (UDW) detector outside a BTZ black hole, from a perspective of quantum thermodynamics. In the context of an open quantum system, we derive the complete dynamics of the detector, which encodes a complicated response to scalar background fields. Using various information theory tools, such as quantum relative entropy, quantum heat, coherence, quantum Fisher information, and quantum speed of evolution, we examined three quantum thermodynamic laws for the UDW detector, where the influences from BTZ angular momentum and Hawking radiation are investigated. In particular, based on information geometry theory, we find an intrinsic asymmetry in the detector's thermolization process as it undergoes Hawking radiation from the BTZ black hole. In particular, we find that the detector consistently heats faster than it cools, analogous to the quantum Mpemba effect for nonequilibrium systems. Moreover, we demonstrate that the spin of a black hole significantly influences the magnitude of the asymmetry, while preserving the dominance of heating over cooling.

hep-th↗

PASC-Net:Plug-and-play Shape Self-learning Convolutions Network with Hierarchical Topology Constraints for Vessel Segmentation

Accurate vessel segmentation is crucial to assist in clinical diagnosis by medical experts. However, the intricate tree-like tubular structure of blood vessels poses significant challenges for existing segmentation algorithms. Small vascular branches are often overlooked due to their low contrast compared to surrounding tissues, leading to incomplete vessel segmentation. Furthermore, the complex vascular topology prevents the model from accurately capturing and reconstructing vascular structure, resulting in incorrect topology, such as breakpoints at the bifurcation of the vascular tree. To overcome these challenges, we propose a novel vessel segmentation framework called PASC Net. It includes two key modules: a plug-and-play shape self-learning convolutional (SSL) module that optimizes convolution kernel design, and a hierarchical topological constraint (HTC) module that ensures vascular connectivity through topological constraints. Specifically, the SSL module enhances adaptability to vascular structures by optimizing conventional convolutions into learnable strip convolutions, which improves the network's ability to perceive fine-grained features of tubular anatomies. Furthermore, to better preserve the coherence and integrity of vascular topology, the HTC module incorporates hierarchical topological constraints-spanning linear, planar, and volumetric levels-which serve to regularize the network's representation of vascular continuity and structural consistency. We replaced the standard convolutional layers in U-Net, FCN, U-Mamba, and nnUNet with SSL convolutions, leading to consistent performance improvements across all architectures. Furthermore, when integrated into the nnUNet framework, our method outperformed other methods on multiple metrics, achieving state-of-the-art vascular segmentation performance.

eess.IV↗

TABNet: A Triplet Augmentation Self-Recovery Framework with Boundary-Aware Pseudo-Labels for Medical Image Segmentation

Background and objective: Medical image segmentation is a core task in various clinical applications. However, acquiring large-scale, fully annotated medical image datasets is both time-consuming and costly. Scribble annotations, as a form of sparse labeling, provide an efficient and cost-effective alternative for medical image segmentation. However, the sparsity of scribble annotations limits the feature learning of the target region and lacks sufficient boundary supervision, which poses significant challenges for training segmentation networks. Methods: We propose TAB Net, a novel weakly-supervised medical image segmentation framework, consisting of two key components: the triplet augmentation self-recovery (TAS) module and the boundary-aware pseudo-label supervision (BAP) module. The TAS module enhances feature learning through three complementary augmentation strategies: intensity transformation improves the model's sensitivity to texture and contrast variations, cutout forces the network to capture local anatomical structures by masking key regions, and jigsaw augmentation strengthens the modeling of global anatomical layout by disrupting spatial continuity. By guiding the network to recover complete masks from diverse augmented inputs, TAS promotes a deeper semantic understanding of medical images under sparse supervision. The BAP module enhances pseudo-supervision accuracy and boundary modeling by fusing dual-branch predictions into a loss-weighted pseudo-label and introducing a boundary-aware loss for fine-grained contour refinement. Results: Experimental evaluations on two public datasets, ACDC and MSCMR seg, demonstrate that TAB Net significantly outperforms state-of-the-art methods for scribble-based weakly supervised segmentation. Moreover, it achieves performance comparable to that of fully supervised methods.

cs.CV↗

Quantum Fisher information of a cosmic qubit undergoing non-Markovian de Sitter evolution

We revisit the problem of thermalization process for an Unruh-DeWitt (UDW) detector in de Sitter space. We derive the full dynamics of the detector in the context of open quantum system, neither using Markovian or RWA approximations. We utilize quantum Fisher information (QFI) for Hubble parameter estimation, as a process function to distinguish the thermalization paths in detector Hilbert space, determined by its local properties, e.g., detector energy gap and its initial state preparation, or global spacetime geometry. We find that the non-Markovian contribution in general reduces the QFI comparing with Markovian approximated solution. Regarding to arbitrary initial states, the late-time QFI would converge to an asymptotic value. In particular, we are interested in the background field in the one parameter family of $α$-vacua in de Sitter space. We show that for general $α$-vacuum choices, the asymptotic values of converged QFI are significantly suppressed, comparing to previous known results for Bunch-Davies vacuum.

hep-th↗

SHIELD : An Evaluation Benchmark for Face Spoofing and Forgery Detection with Multimodal Large Language Models

Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-related tasks, capitalizing on their visual semantic comprehension and reasoning capabilities. However, their ability to detect subtle visual spoofing and forgery clues in face attack detection tasks remains underexplored. In this paper, we introduce a benchmark, SHIELD, to evaluate MLLMs for face spoofing and forgery detection. Specifically, we design true/false and multiple-choice questions to assess MLLM performance on multimodal face data across two tasks. For the face anti-spoofing task, we evaluate three modalities (i.e., RGB, infrared, and depth) under six attack types. For the face forgery detection task, we evaluate GAN-based and diffusion-based data, incorporating visual and acoustic modalities. We conduct zero-shot and few-shot evaluations in standard and chain of thought (COT) settings. Additionally, we propose a novel multi-attribute chain of thought (MA-COT) paradigm for describing and judging various task-specific and task-irrelevant attributes of face images. The findings of this study demonstrate that MLLMs exhibit strong potential for addressing the challenges associated with the security of facial recognition technology applications.

cs.CV↗

Adversary-Aware DPO: Enhancing Safety Alignment in Vision Language Models via Adversarial Training

Safety alignment is critical in pre-training large language models (LLMs) to generate responses aligned with human values and refuse harmful queries. Unlike LLM, the current safety alignment of VLMs is often achieved with post-hoc safety fine-tuning. However, these methods are less effective to white-box attacks. To address this, we propose $\textit{Adversary-aware DPO (ADPO)}$, a novel training framework that explicitly considers adversarial. $\textit{Adversary-aware DPO (ADPO)}$ integrates adversarial training into DPO to enhance the safety alignment of VLMs under worst-case adversarial perturbations. $\textit{ADPO}$ introduces two key components: (1) an adversarial-trained reference model that generates human-preferred responses under worst-case perturbations, and (2) an adversarial-aware DPO loss that generates winner-loser pairs accounting for adversarial distortions. By combining these innovations, $\textit{ADPO}$ ensures that VLMs remain robust and reliable even in the presence of sophisticated jailbreak attacks. Extensive experiments demonstrate that $\textit{ADPO}$ outperforms baselines in the safety alignment and general utility of VLMs.

cs.CR↗

Relative entropy formulation of thermalization process in a Schwarzschild spacetime

We revisit the problem of the thermalization process in an entropic formulation for the Unruh-DeWitt (UDW) detector outside a Schwarzschild black hole. We derive the late-time dynamics of the detector in the context of open quantum system, and capture the path distinguishability and thermodynamic irreversibility of detector thermalization process by using quantum relative entropy (QRE). We find that beyond the Planckian transition rate, the refined thermalization process in detector Hilbert space can be distinguished by the time behavior of the related QRE. We show that the exotic position-dependent behaviors of the QRE emerge corresponding to different choices of black hole vacua (i.e., the Boulware, Hartle-Hawking, and Unruh vacua). Finally, from a perspective of quantum thermodynamics, we recast the free energy change of the UDW detector undergoing Hawking radiation into an entropic combination form, where the classical Kullback-Leibler divergence and quantum coherence are presented in specific QRE-like forms. With growing Hawking temperature, we find that the consumption rate of quantum coherence is larger than that of its classical counterpart.

hep-th↗

Coherence revival under the Unruh effect and its metrological advantage

In this paper, we investigate the quantum coherence extraction {between} two accelerating Unruh-DeWitt detectors, coupling to a scalar field in $(3+1)$-dimensional Minkowski spacetime. We find that quantum coherence as a nonclassical correlation can be generated through the Markovian evolution of the {detector} system, just like quantum entanglement. However, with growing Unruh temperature, in contrast to monotonously degrading entanglement, we find that quantum coherence exhibits a striking revival phenomenon. For certain detectors' initial state choices, {the} coherence measure will reduce to zero at first {and} then grow to an asymptotic value. We verify such coherence revival by inspecting its metrological advantage on the quantum Fisher information (QFI) enhancement. Since the maximal QFI {bounds} the accuracy of quantum parameter estimation, we conclude that the extracted coherence can be utilized as a physical resource in quantum metrology.

hep-th↗

HELPNet: Hierarchical Perturbations Consistency and Entropy-guided Ensemble for Scribble Supervised Medical Image Segmentation

Creating fully annotated labels for medical image segmentation is prohibitively time-intensive and costly, emphasizing the necessity for innovative approaches that minimize reliance on detailed annotations. Scribble annotations offer a more cost-effective alternative, significantly reducing the expenses associated with full annotations. However, scribble annotations offer limited and imprecise information, failing to capture the detailed structural and boundary characteristics necessary for accurate organ delineation. To address these challenges, we propose HELPNet, a novel scribble-based weakly supervised segmentation framework, designed to bridge the gap between annotation efficiency and segmentation performance. HELPNet integrates three modules. The Hierarchical perturbations consistency (HPC) module enhances feature learning by employing density-controlled jigsaw perturbations across global, local, and focal views, enabling robust modeling of multi-scale structural representations. Building on this, the Entropy-guided pseudo-label (EGPL) module evaluates the confidence of segmentation predictions using entropy, generating high-quality pseudo-labels. Finally, the structural prior refinement (SPR) module incorporates connectivity and bounded priors to enhance the precision and reliability and pseudo-labels. Experimental results on three public datasets ACDC, MSCMRseg, and CHAOS show that HELPNet significantly outperforms state-of-the-art methods for scribble-based weakly supervised segmentation and achieves performance comparable to fully supervised methods. The code is available at https://github.com/IPMI-NWU/HELPNet.

cs.CV↗

APS-LSTM: Exploiting Multi-Periodicity and Diverse Spatial Dependencies for Flood Forecasting

Accurate flood prediction is crucial for disaster prevention and mitigation. Hydrological data exhibit highly nonlinear temporal patterns and encompass complex spatial relationships between rainfall and flow. Existing flood prediction models struggle to capture these intricate temporal features and spatial dependencies. This paper presents an adaptive periodic and spatial self-attention method based on LSTM (APS-LSTM) to address these challenges. The APS-LSTM learns temporal features from a multi-periodicity perspective and captures diverse spatial dependencies from different period divisions. The APS-LSTM consists of three main stages, (i) Multi-Period Division, that utilizes Fast Fourier Transform (FFT) to divide various periodic patterns; (ii) Spatio-Temporal Information Extraction, that performs periodic and spatial self-attention focusing on intra- and inter-periodic temporal patterns and spatial dependencies; (iii) Adaptive Aggregation, that relies on amplitude strength to aggregate the computational results from each periodic division. The abundant experiments on two real-world datasets demonstrate the superiority of APS-LSTM. The code is available: https://github.com/oopcmd/APS-LSTM.

cs.LG↗

Gaze-directed Vision GNN for Mitigating Shortcut Learning in Medical Image

Deep neural networks have demonstrated remarkable performance in medical image analysis. However, its susceptibility to spurious correlations due to shortcut learning raises concerns about network interpretability and reliability. Furthermore, shortcut learning is exacerbated in medical contexts where disease indicators are often subtle and sparse. In this paper, we propose a novel gaze-directed Vision GNN (called GD-ViG) to leverage the visual patterns of radiologists from gaze as expert knowledge, directing the network toward disease-relevant regions, and thereby mitigating shortcut learning. GD-ViG consists of a gaze map generator (GMG) and a gaze-directed classifier (GDC). Combining the global modelling ability of GNNs with the locality of CNNs, GMG generates the gaze map based on radiologists' visual patterns. Notably, it eliminates the need for real gaze data during inference, enhancing the network's practical applicability. Utilizing gaze as the expert knowledge, the GDC directs the construction of graph structures by incorporating both feature distances and gaze distances, enabling the network to focus on disease-relevant foregrounds. Thereby avoiding shortcut learning and improving the network's interpretability. The experiments on two public medical image datasets demonstrate that GD-ViG outperforms the state-of-the-art methods, and effectively mitigates shortcut learning. Our code is available at https://github.com/SX-SS/GD-ViG.

cs.CV↗

Prompt-Guided Generation of Structured Chest X-Ray Report Using a Pre-trained LLM

Medical report generation automates radiology descriptions from images, easing the burden on physicians and minimizing errors. However, current methods lack structured outputs and physician interactivity for clear, clinically relevant reports. Our method introduces a prompt-guided approach to generate structured chest X-ray reports using a pre-trained large language model (LLM). First, we identify anatomical regions in chest X-rays to generate focused sentences that center on key visual elements, thereby establishing a structured report foundation with anatomy-based sentences. We also convert the detected anatomy into textual prompts conveying anatomical comprehension to the LLM. Additionally, the clinical context prompts guide the LLM to emphasize interactivity and clinical requirements. By integrating anatomy-focused sentences and anatomy/clinical prompts, the pre-trained LLM can generate structured chest X-ray reports tailored to prompted anatomical regions and clinical contexts. We evaluate using language generation and clinical effectiveness metrics, demonstrating strong performance.

cs.AI↗

Scalable Normalizing Flows Enable Boltzmann Generators for Macromolecules

The Boltzmann distribution of a protein provides a roadmap to all of its functional states. Normalizing flows are a promising tool for modeling this distribution, but current methods are intractable for typical pharmacological targets; they become computationally intractable due to the size of the system, heterogeneity of intra-molecular potential energy, and long-range interactions. To remedy these issues, we present a novel flow architecture that utilizes split channels and gated attention to efficiently learn the conformational distribution of proteins defined by internal coordinates. We show that by utilizing a 2-Wasserstein loss, one can smooth the transition from maximum likelihood training to energy-based training, enabling the training of Boltzmann Generators for macromolecules. We evaluate our model and training strategy on villin headpiece HP35(nle-nle), a 35-residue subdomain, and protein G, a 56-residue protein. We demonstrate that standard architectures and training strategies, such as maximum likelihood alone, fail while our novel architecture and multi-stage training strategy are able to model the conformational distributions of protein G and HP35.

cs.LG↗

Cross-Corpus Multilingual Speech Emotion Recognition: Amharic vs. Other Languages

In a conventional Speech emotion recognition (SER) task, a classifier for a given language is trained on a pre-existing dataset for that same language. However, where training data for a language does not exist, data from other languages can be used instead. We experiment with cross-lingual and multilingual SER, working with Amharic, English, German and URDU. For Amharic, we use our own publicly-available Amharic Speech Emotion Dataset (ASED). For English, German and Urdu we use the existing RAVDESS, EMO-DB and URDU datasets. We followed previous research in mapping labels for all datasets to just two classes, positive and negative. Thus we can compare performance on different languages directly, and combine languages for training and testing. In Experiment 1, monolingual SER trials were carried out using three classifiers, AlexNet, VGGE (a proposed variant of VGG), and ResNet50. Results averaged for the three models were very similar for ASED and RAVDESS, suggesting that Amharic and English SER are equally difficult. Similarly, German SER is more difficult, and Urdu SER is easier. In Experiment 2, we trained on one language and tested on another, in both directions for each pair: Amharic<->German, Amharic<->English, and Amharic<->Urdu. Results with Amharic as target suggested that using English or German as source will give the best result. In Experiment 3, we trained on several non-Amharic languages and then tested on Amharic. The best accuracy obtained was several percent greater than the best accuracy in Experiment 2, suggesting that a better result can be obtained when using two or three non-Amharic languages for training than when using just one non-Amharic language. Overall, the results suggest that cross-lingual and multilingual training can be an effective strategy for training a SER classifier when resources for a language are scarce.

cs.CL↗

Invariant Content Synergistic Learning for Domain Generalization of Medical Image Segmentation

While achieving remarkable success for medical image segmentation, deep convolution neural networks (DCNNs) often fail to maintain their robustness when confronting test data with the novel distribution. To address such a drawback, the inductive bias of DCNNs is recently well-recognized. Specifically, DCNNs exhibit an inductive bias towards image style (e.g., superficial texture) rather than invariant content (e.g., object shapes). In this paper, we propose a method, named Invariant Content Synergistic Learning (ICSL), to improve the generalization ability of DCNNs on unseen datasets by controlling the inductive bias. First, ICSL mixes the style of training instances to perturb the training distribution. That is to say, more diverse domains or styles would be made available for training DCNNs. Based on the perturbed distribution, we carefully design a dual-branches invariant content synergistic learning strategy to prevent style-biased predictions and focus more on the invariant content. Extensive experimental results on two typical medical image segmentation tasks show that our approach performs better than state-of-the-art domain generalization methods.

eess.IV↗

Scalar perturbation of gravitating double-kink solutions

In this letter, a two-dimensional (2D) gravity-scalar model is studied. This model supports interesting double-kink solutions, and the corresponding metric solutions can be derived analytically. Depending on a tunable parameter $c$, the metric can be symmetric or asymmetric. The Schrödinger-like equation for normal modes of the physical linear perturbation is derived. As $c$ varies, the effective potential can have one or two singular barriers. If $c$ is larger than a critical value, the zero mode will be normalizable, despite of the appearance of a strong repulsive singularity. The double-kink solution is always stable against linear perturbations.

hep-th↗

Thermality of the Unruh effect with intermediate statistics

Utilizing quantum coherence monotone, we reexamine the thermal nature of the Unruh effect of an accelerating detector. We consider an UDW detector coupling to a n-dimensional conformal field in Minkowski spacetime, whose response spectrum generally exhibits an intermediate statistics of (1+1) anyon field. We find that the thermal nature of the Unruh effect guaranteed by KMS condition is characterized by a vanishing asymptotic quantum coherence. We show that the time-evolution of coherence monotone can distinguish the different thermalizing ways of the detector, which depends on the scaling dimension of the conformal primary field. In particular, for the conformal background with certain scaling dimension, we demonstrate that at fixed proper time a revival of coherence can occur even for growing Unruh decoherence. Finally, we show that coherence monotone has distinct dynamics under the Unruh decoherence and a thermal bath for a static observer.

hep-th↗

Quantum Fisher information as a probe for Unruh thermality

A long-standing debate on Unruh effect is about its obscure thermal nature. In this Letter, we use quantum Fisher information (QFI) as an effective probe to explore the thermal nature of Unruh effect from both local and global perspectives. By resolving the full dynamics of UDW detector, we find that the QFI is a time-evolving function of detector's energy gap, Unruh temperature $T_U$ and particularities of background field, e.g., mass and spacetime dimensionality. We show that the asymptotic QFI whence detector arrives its equilibrium is solely determined by $T_U$, demonstrating the global side of Unruh thermality alluded by the KMS condition. We also show that the local side of Unruh effect, i.e., the different ways for the detector to approach the same thermal equilibrium, is encoded in the corresponding time-evolution of the QFI. In particular, we find that with massless scalar background the QFI has unique monotonicity in $n=3$ dimensional spacetime, and becomes non-monotonous for $n\neq3$ models where a local peak value exists at early time and for finite acceleration, indicating an enhanced precision of estimation on Unruh temperature at a relative low acceleration can be achieved. Once the field acquiring mass, the related QFI becomes significantly robust against the Unruh decoherence in the sense that its local peak sustains for a very long time. While coupling to a more massive background, the persistence can even be strengthened and the QFI possesses a larger maximal value. Such robustness of QFI can surely facilitate any practical quantum estimation task.

hep-th↗