SearcharxivSearch

arXiv subjects

Haishu Tan

Publications and source records attributed to Haishu Tan.

At least 19 recordsLinked to original sources

IRPol-Fuse: Energy-structure coordination for infrared polarization fusion under low visibility

Robust perception under low-visibility conditions requires fused imagery that jointly preserves infrared thermal saliency and polarization-derived structural details. However, existing infrared-polarization image fusion (IPIF) methods often overemphasize dominant infrared responses, causing weak yet informative polarization textures in dark regions to be suppressed. To address this issue, we propose IRPol-Fuse, an energy-structure coordinated IPIF framework for challenging low-visibility scenarios. The proposed framework contains three key modules: Polarization Attention Fusion for adaptive infrared-polarization allocation, Infrared Highlight Injector for highlight-guided infrared preservation, and Polarization Texture Injector for polarization texture restoration and fine-detail recovery. We further construct LI-PI, a dedicated infrared-polarization evaluation dataset for low-visibility and visually concealed scenes. Experiments on LI-PI and the public LDDRS dataset demonstrate that IRPol-Fuse achieves favorable performance in thermal target preservation, structural detail recovery, and visual naturalness. Region-aware evaluation and downstream object detection further verify that the proposed energy-structure coordination strategy effectively preserves both infrared target saliency and polarization-derived structural information. Code is available at https://github.com/1hzf/IRPolar-Fuse .

cs.CV

Multi-modality Image Fusion under Adverse Weather: Mask-Guided Feature Restoration and Interaction

Multi-modality image fusion (MMIF) enhances scene representation by exploiting complementary cues from different modalities. Adverse weather, however, causes significant image degradation, disrupting feature representation and requiring simultaneous feature restoration and cross-modal complementarity. Existing methods often struggle with effective representation learning under such conditions, limiting their practical performance. To address these challenges, we propose a mask-guided MMIF method that integrates feature restoration and interaction. We first introduce "Pseudo Ground Truth" to simplify training, promoting faster and more effective feature learning. Then, we design a mask generation mechanism based on the mapping relationship between the fused result and the source images, quantifying the relative contribution of each modality during the fusion process. By incorporating the proposed mask-guided cross-modal cross-attention mechanism, the network is encouraged to selectively attend to informative features during modality interaction, mitigating the risk of overfitting to the static distribution of the "Pseudo Ground Truth". Additionally, we propose a mask-guided learning strategy and a task-coupled degradation-aware learning strategy to balance feature restoration and interaction. Extensive experiments on synthetic and real-world datasets demonstrate that our method surpasses state-of-the-art approaches in visual quality, quantitative metrics, and downstream tasks. The source code is available at https://github.com/ixilai/AMG-Fuse.

cs.CV

MAVFusion: Efficient Infrared and Visible Video Fusion via Motion-Aware Sparse Interaction

Infrared and visible video fusion combines the object saliency from infrared images with the texture details from visible images to produce semantically rich fusion results. However, most existing methods are designed for static image fusion and cannot effectively handle frame-to-frame motion in videos. Current video fusion methods improve temporal consistency by introducing interactions across frames, but they often require high computational cost. To mitigate these challenges, we propose MAVFusion, an end-to-end video fusion framework featuring a motion-aware sparse interaction mechanism that enhances efficiency while maintaining superior fusion quality. Specifically, we leverage optical flow to identify dynamic regions in multi-modal sequences, adaptively allocating computationally intensive cross-modal attention to these sparse areas to capture salient transitions and facilitate inter-modal information exchange. For static background regions, a lightweight weak interaction module is employed to maintain structural and appearance integrity. By decoupling the processing of dynamic and static regions, MAVFusion simultaneously preserves temporal consistency and fine-grained details while significantly accelerating inference. Extensive experiments demonstrate that MAVFusion achieves state-of-the-art performance on multiple infrared and visible video benchmarks, achieving a speed of 14.16 FPS at $640 \times 480$ resolution. The source code will be available at https://github.com/ixilai/MAVFusion.

cs.CV

CAWM-Mamba: A unified model for infrared-visible image fusion and compound adverse weather restoration

Multimodal Image Fusion (MMIF) integrates complementary information from various modalities to produce clearer and more informative fused images. MMIF under adverse weather is particularly crucial in autonomous driving and UAV monitoring applications. However, existing adverse weather fusion methods generally only tackle single types of degradation such as haze, rain, or snow, and fail when multiple degradations coexist (e.g., haze+rain, rain+snow). To address this challenge, we propose Compound Adverse Weather Mamba (CAWM-Mamba), the first end-to-end framework that jointly performs image fusion and compound weather restoration with unified shared weights. Our network contains three key components: (1) a Weather-Aware Preprocess Module (WAPM) to enhance degraded visible features and extracts global weather embeddings; (2) a Cross-modal Feature Interaction Module (CFIM) to facilitate the alignment of heterogeneous modalities and exchange of complementary features across modalities; and (3) a Wavelet Space State Block (WSSB) that leverages wavelet-domain decomposition to decouple multi-frequency degradations. WSSB includes Freq-SSM, a module that models anisotropic high-frequency degradation without redundancy, and a unified degradation representation mechanism to further improve generalization across complex compound weather conditions. Extensive experiments on the AWMM-100K benchmark and three standard fusion datasets demonstrate that CAWM-Mamba consistently outperforms state-of-the-art methods in both compound and single-weather scenarios. In addition, our fusion results excel in downstream tasks covering semantic segmentation and object detection, confirming the practical value in real-world adverse weather perception. The source code will be available at https://github.com/Feecuin/CAWM-Mamba.

cs.CV

A Luminance-Aware Multi-Scale Network for Polarization Image Fusion with a Multi-Scene Dataset

Polarization image fusion combines S0 and DOLP images to reveal surface roughness and material properties through complementary texture features, which has important applications in camouflage recognition, tissue pathology analysis, surface defect detection and other fields. To intergrate coL-Splementary information from different polarized images in complex luminance environment, we propose a luminance-aware multi-scale network (MLSN). In the encoder stage, we propose a multi-scale spatial weight matrix through a brightness-branch , which dynamically weighted inject the luminance into the feature maps, solving the problem of inherent contrast difference in polarized images. The global-local feature fusion mechanism is designed at the bottleneck layer to perform windowed self-attention computation, to balance the global context and local details through residual linking in the feature dimension restructuring stage. In the decoder stage, to further improve the adaptability to complex lighting, we propose a Brightness-Enhancement module, establishing the mapping relationship between luminance distribution and texture features, realizing the nonlinear luminance correction of the fusion result. We also present MSP, an 1000 pairs of polarized images that covers 17 types of indoor and outdoor complex lighting scenes. MSP provides four-direction polarization raw maps, solving the scarcity of high-quality datasets in polarization image fusion. Extensive experiment on MSP, PIF and GAND datasets verify that the proposed MLSN outperms the state-of-the-art methods in subjective and objective evaluations, and the MS-SSIM and SD metircs are higher than the average values of other methods by 8.57%, 60.64%, 10.26%, 63.53%, 22.21%, and 54.31%, respectively. The source code and dataset is avalable at https://github.com/1hzf/MLS-UNet.

cs.CV

NDLPNet: A Location-Aware Nighttime Deraining Network and a Real-World Benchmark Dataset

Visual degradation caused by rain streak artifacts in low-light conditions significantly hampers the performance of nighttime surveillance and autonomous navigation. Existing image deraining techniques are primarily designed for daytime conditions and perform poorly under nighttime illumination due to the spatial heterogeneity of rain distribution and the impact of light-dependent stripe visibility. In this paper, we propose a novel Nighttime Deraining Location-enhanced Perceptual Network(NDLPNet) that effectively captures the spatial positional information and density distribution of rain streaks in low-light environments. Specifically, we introduce a Position Perception Module (PPM) to capture and leverage spatial contextual information from input data, enhancing the model's capability to identify and recalibrate the importance of different feature channels. The proposed nighttime deraining network can effectively remove the rain streaks as well as preserve the crucial background information. Furthermore, We construct a night scene rainy (NSR) dataset comprising 900 image pairs, all based on real-world nighttime scenes, providing a new benchmark for nighttime deraining task research. Extensive qualitative and quantitative experimental evaluations on both existing datasets and the NSR dataset consistently demonstrate our method outperform the state-of-the-art (SOTA) methods in nighttime deraining tasks. The source code and dataset is available at https://github.com/Feecuin/NDLPNet.

cs.CV

FlexiD-Fuse: Flexible number of inputs multi-modal medical image fusion based on diffusion model

Different modalities of medical images provide unique physiological and anatomical information for diseases. Multi-modal medical image fusion integrates useful information from different complementary medical images with different modalities, producing a fused image that comprehensively and objectively reflects lesion characteristics to assist doctors in clinical diagnosis. However, existing fusion methods can only handle a fixed number of modality inputs, such as accepting only two-modal or tri-modal inputs, and cannot directly process varying input quantities, which hinders their application in clinical settings. To tackle this issue, we introduce FlexiD-Fuse, a diffusion-based image fusion network designed to accommodate flexible quantities of input modalities. It can end-to-end process two-modal and tri-modal medical image fusion under the same weight. FlexiD-Fuse transforms the diffusion fusion problem, which supports only fixed-condition inputs, into a maximum likelihood estimation problem based on the diffusion process and hierarchical Bayesian modeling. By incorporating the Expectation-Maximization algorithm into the diffusion sampling iteration process, FlexiD-Fuse can generate high-quality fused images with cross-modal information from source images, independently of the number of input images. We compared the latest two and tri-modal medical image fusion methods, tested them on Harvard datasets, and evaluated them using nine popular metrics. The experimental results show that our method achieves the best performance in medical image fusion with varying inputs. Meanwhile, we conducted extensive extension experiments on infrared-visible, multi-exposure, and multi-focus image fusion tasks with arbitrary numbers, and compared them with the perspective SOTA methods. The results of the extension experiments consistently demonstrate the effectiveness and superiority of our method.

cs.CV

MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks

The goal of multimodal image fusion is to integrate complementary information from infrared and visible images, generating multimodal fused images for downstream tasks. Existing downstream pre-training models are typically trained on visible images. However, the significant pixel distribution differences between visible and multimodal fusion images can degrade downstream task performance, sometimes even below that of using only visible images. This paper explores adapting multimodal fused images with significant modality differences to object detection and semantic segmentation models trained on visible images. To address this, we propose MambaTrans, a novel multimodal fusion image modality translator. MambaTrans uses descriptions from a multimodal large language model and masks from semantic segmentation models as input. Its core component, the Multi-Model State Space Block, combines mask-image-text cross-attention and a 3D-Selective Scan Module, enhancing pure visual capabilities. By leveraging object detection prior knowledge, MambaTrans minimizes detection loss during training and captures long-term dependencies among text, masks, and images. This enables favorable results in pre-trained models without adjusting their parameters. Experiments on public datasets show that MambaTrans effectively improves multimodal image performance in downstream tasks.

cs.CV

Simultaneous Tri-Modal Medical Image Fusion and Super-Resolution using Conditional Diffusion Model

In clinical practice, tri-modal medical image fusion, compared to the existing dual-modal technique, can provide a more comprehensive view of the lesions, aiding physicians in evaluating the disease's shape, location, and biological activity. However, due to the limitations of imaging equipment and considerations for patient safety, the quality of medical images is usually limited, leading to sub-optimal fusion performance, and affecting the depth of image analysis by the physician. Thus, there is an urgent need for a technology that can both enhance image resolution and integrate multi-modal information. Although current image processing methods can effectively address image fusion and super-resolution individually, solving both problems synchronously remains extremely challenging. In this paper, we propose TFS-Diff, a simultaneously realize tri-modal medical image fusion and super-resolution model. Specially, TFS-Diff is based on the diffusion model generation of a random iterative denoising process. We also develop a simple objective function and the proposed fusion super-resolution loss, effectively evaluates the uncertainty in the fusion and ensures the stability of the optimization process. And the channel attention module is proposed to effectively integrate key information from different modalities for clinical diagnosis, avoiding information loss caused by multiple image processing. Extensive experiments on public Harvard datasets show that TFS-Diff significantly surpass the existing state-of-the-art methods in both quantitative and visual evaluations. Code is available at https://github.com/XylonXu01/TFS-Diff.

eess.IV

TSJNet: A Multi-modality Target and Semantic Awareness Joint-driven Image Fusion Network

This study aims to address the problem of incomplete information in unimodal images for semantic segmentation and object detection tasks. Existing multimodal fusion methods suffer from limited capability in discriminative modeling of multi-scale semantic structures and salient target regions, which further restricts the effective fusion of task-related semantic details and target information across modalities. To tackle these challenges, this paper proposes a novel fusion network termed TSJNet, which leverages the semantic information output by high-level tasks in a joint manner to guide the fusion process. Specifically, we design a multi-dimensional feature extraction module with dual parallel branches to capture multi-scale and salient features. Meanwhile, a data-agnostic spatial attention module embedded in the decoder dynamically calibrates attention allocation across different data domains, significantly enhancing the model's generalization ability. To optimize both fusion and advanced visual tasks, we balance performance by combining fusion loss with semantic losses. Additionally, we have developed a multimodal unmanned aerial vehicle (UAV) dataset covering multiple scenarios (UMS). Extensive experiments demonstrate that TSJNet achieves outstanding performance on five public datasets (MSRS, M\textsuperscript{3}FD, RoadScene, LLVIP, and TNO) and our UMS dataset. The generated fusion results exhibit favorable visual effects, and compared to state-of-the-art methods, the mean average precision (mAP@0.5) and mean intersection over union (mIoU) for object detection and segmentation, respectively, improve by 7.97\% and 10.88\%.The code and the dataset has been publicly released at https://github.com/XylonXu01/TSJNet.

cs.CV

SAMF: Small-Area-Aware Multi-focus Image Fusion for Object Detection

Existing multi-focus image fusion (MFIF) methods often fail to preserve the uncertain transition region and detect small focus areas within large defocused regions accurately. To address this issue, this study proposes a new small-area-aware MFIF algorithm for enhancing object detection capability. First, we enhance the pixel attributes within the small focus and boundary regions, which are subsequently combined with visual saliency detection to obtain the pre-fusion results used to discriminate the distribution of focused pixels. To accurately ensure pixel focus, we consider the source image as a combination of focused, defocused, and uncertain regions and propose a three-region segmentation strategy. Finally, we design an effective pixel selection rule to generate segmentation decision maps and obtain the final fusion results. Experiments demonstrated that the proposed method can accurately detect small and smooth focus areas while improving object detection performance, outperforming existing methods in both subjective and objective evaluations. The source code is available at https://github.com/ixilai/SAMF.

cs.CV

Bridging the Gap between Multi-focus and Multi-modal: A Focused Integration Framework for Multi-modal Image Fusion

Multi-modal image fusion (MMIF) integrates valuable information from different modality images into a fused one. However, the fusion of multiple visible images with different focal regions and infrared images is a unprecedented challenge in real MMIF applications. This is because of the limited depth of the focus of visible optical lenses, which impedes the simultaneous capture of the focal information within the same scene. To address this issue, in this paper, we propose a MMIF framework for joint focused integration and modalities information extraction. Specifically, a semi-sparsity-based smoothing filter is introduced to decompose the images into structure and texture components. Subsequently, a novel multi-scale operator is proposed to fuse the texture components, capable of detecting significant information by considering the pixel focus attributes and relevant data from various modal images. Additionally, to achieve an effective capture of scene luminance and reasonable contrast maintenance, we consider the distribution of energy information in the structural components in terms of multi-directional frequency variance and information entropy. Extensive experiments on existing MMIF datasets, as well as the object detection and depth estimation tasks, consistently demonstrate that the proposed algorithm can surpass the state-of-the-art methods in visual perception and quantitative evaluation. The code is available at https://github.com/ixilai/MFIF-MMIF.

cs.CV

Dressed bound states at chiral exceptional points

Atom-photon dressed states are a basic concept of quantum optics. Here, we demonstrate that the non-Hermiticity of open cavity can be harnessed to form the dressed bound states (DBS) and identify two types of DBS, the vacancy-like DBS and Friedrich-Wintgen DBS, in a microring resonator operating at a chiral exceptional point. With the analytical DBS conditions, we show that the vacancy-like DBS occurs when an atom couples to the standing wave mode that is a node of photonic wave function, and thus is immune to the cavity dissipation and characterized by the null spectral density at cavity resonance. While the Friedrich-Wintgen DBS can be accessed by continuously tuning the system parameters, such as the atom-photon detuning, and evidenced by a vanishing Rabi peak in emission spectrum, an unusual feature in the strong-coupling anticrossing. We also demonstrate the quantum-optics applications of the proposed DBS. Our work exhibits the quantum states control through non-Hermiticity of open quantum system and presents a clear physical picture on DBS at chiral exceptional points, which holds great potential in building high-performance quantum devices for sensing, photon storage, and nonclassical light generation.

quant-ph

Fano effect induced giant and robust enhancement of photon correlations in cavity QED systems

Fano effect arising from interference between two dissipation channels to the radiation continuum enables to tune the photon statistics. Understanding the role of Fano effect and exploiting to achieve strong photon correlations are of both fundamental and applied significance. We present an analytical description of Fano-enhanced photon correlations based on cavity quantum electrodynamics to show that, the Fano effect in atom-cavity systems can improve the degree of antibunching by over four orders of magnitude. The enhancement factors and the optimal conditions are explicitly given, and found related to the Fano parameter $q$. Remarkably, the Fano enhancement manifests robustness against the decoherence processes, and can survive in the weak coupling regime. We expect our work to provide insight in tuning the photon statistics through Fano effect, which offers a new route to enhance the photon correlations, as well as the possibility of generating nonclassical light in a wider diversity of systems without the need of strong light-matter interaction.

quant-ph

Single-photon blockade in quasichiral atom-photon interaction: Simultaneous high purity and high efficiency

We investigate the single-photon blockade (1PB) in the quasichiral regime of atom-photon interaction that mediates via dissipative environment, where the effective atom-photon interaction is asymmetrical but achiral. The synthetic magnetic current in the closed-loop coupling breaks down the reciprocity of atom-photon interaction, resulting in an asymmetrical and even completely unidirectional effective coupling between two selected quantum states. As an example, we couple the single-atom cavity-QED (cQED) system to a strongly dissipative auxiliary cavity. We find that in the quasichiral regime, the unconventional photon blockade (UPB) always incorporates with the conventional photon blockade (CPB) in the condition of maximum chirality. Furthermore, we show that 1PB in the quasichiral regime combines the advantages of UPB and CPB, demonstrating the perfect single-photon purity, higher efficiency, smooth time dynamics as well as lower requirement of modes coupling to achieve UPB. Our work paves the way for 1PB towards practical applications and reveals the intriguing quantum-optics phenomena in the quasichiral light-matter interaction.

physics.optics

Two-dimensional vortex quantum droplets

It was recently found that the Lee-Huang-Yang (LHY) correction to the mean-field Hamiltonian suppresses the collapse and creates stable localized modes (two-component "quantum droplets", QDs) in two and three dimensions. We construct two-dimensional\ self-trapped modes in the form of QDs with vorticity $S$ embedded into each component. The QDs feature a flat-top shape, which expands with the increase of $S$ and norm $N$. An essential finding, produced by a systematic numerical analysis and analytical estimates, is that the vortical QDs are \emph{stable} (which is a critical issue for vortex solitons in nonlinear models) up to $S=5$, for $N$ exceeding a certain threshold value. In the condensate of $^{39}$K atoms, in which QDs with $S=0$ and a quasi-2D shape were created recently, the vortical droplets may have radial size $\lesssim 30$ $\mathrmμ$m, with the number of atoms in the range of $10^{4}-10^{5}$. It is worthy to note that \textit{hidden-vorticity} states in QDs with topological charges $% S_{+}=-S_{-}=1$ in its components, which are prone to strong instability in other settings, have their stability region too, although it may be located beyond applicability limits of the underlying model. Dynamics of elliptically deformed QDs, which form rotating elongated patterns or ones with strong oscillations of the eccentricity, as well as collisions of QDs, are also addressed.

cond-mat.quant-gas

Self-trapping under the two-dimensional spin-orbit-coupling and spatially growing repulsive nonlinearity

We elaborate a method for the creation of two- and one-dimensional (2D and 1D) self-trapped modes in binary spin-orbit (SO)-coupled Bose-Einstein condensates (BECs) with the contact repulsive interaction, whose local strength grows fast enough from the center to periphery. In particular, an exact semi-vortex (SV) solution is found for the anti-Gaussian radial-modulation profile. The exact modes are included in the numerically produced family of SV solitons. Other families, in the form of mixed modes(MMs), as well as excited state of SVs and MMs, are produced too. While the excited states are unstable in all previously studied models, they are partially stable in the present one. In the 1D version of the system, exact solutions for the counterpart of the SVs, namely, \textit{semi-dipole} solitons, are found too. Families of semi-dipoles, as well as the 1D version of MMs, are produced numerically.

cond-mat.quant-gas

Two-dimensional solitons and quantum droplets supported by competing self- and cross-interactions in spin-orbit-coupled condensates

We study two-dimensional (2D) matter-wave solitons in spinor Bose-Einstein condensates (BECs) under the action of the spin-orbit coupling (SOC) and opposite signs of the self- and cross-interactions. Stable 2D two-component solitons of the mixed-mode (MM) type are found if the cross-interaction between the components is attractive, while the self-interaction is repulsive in each component. Stable solitons of the semi-vortex type are formed in the opposite case, under the action of competing self-attraction and cross-repulsion. The solitons exist with the total norm taking values below a collapse threshold. Further, in the case of the repulsive self-interaction and inter-component attraction, stable 2D self-trapped modes, which may be considered as quantum droplets (QDs), are created if the beyond-mean-field Lee-Huang-Yang (LHY) terms are added to the self-repulsion in the underlying system of coupled Gross-Pitaevskii equations. Stable QDs of the MM type, of a large size with an anisotropic density profile, exist with arbitrarily large values of the norm, as the LHY terms eliminate the collapse. The effect of the SOC term on characteristics of the QDs is systematically studied. We also address the existence and stability of QDs in the case of SOC with mixed Rashba and Dresselhaus terms, which makes the density profile of the QD more isotropic. Thus, QDs in the spin-orbit-coupled binary BEC are for the first time studied in the present work.

cond-mat.quant-gas