Searcharxiv⌕ Search

arXiv subjects

Phat Kim Huynh

Publications and source records attributed to Phat Kim Huynh.

3 recordsLinked to original sources

ESC: Emotional Self-Correction for Reliable Vision-Language Models

Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, yet they remain vulnerable to unreliable reasoning. Existing self-correction methods mitigate these issues but typically rely on post-training or carefully engineered feedback, incurring high computational cost. In this work, we revisit this challenge through the lens of emotional cues, asking whether they can activate latent self-correction behaviors in VLMs without additional training. \textbf{We find that emotional signals serve as an effective trigger for self-correction, encouraging more cautious and reflective reasoning}. Motivated by this finding, we propose \escabstract (\textbf{\underline{E}}motional \textbf{\underline{S}}elf-\textbf{\underline{C}}orrection), a training-free self-correction framework. ESC introduces an external verifier that detects potentially incorrect initial responses and injects emotional feedback to encourage model to reflect, and produce a better revised response without additional training. Extensive experiments across safety, hallucination, vision-centric perception, and multimodal reasoning benchmarks show that ESC consistently improves reliability while preserving overall model utility. These results suggest that emotion can function not only as an ability to be recognized, but also as a practical control signal for scalable self-correction in VLMs. \textbf{We therefore believe that ESC provides a strong foundation for a new reliable human-like, emotion-integrated research direction.} Our project is publicly available at \textcolor{red}{https://genai4e.github.io/ESC/}.

cs.CV↗

Modality-Specific Enhancement and Complementary Fusion for Semi-Supervised Multi-Modal Brain Tumor Segmentation

Semi-supervised learning (SSL) has become a promising direction for medical image segmentation, enabling models to learn from limited labeled data alongside abundant unlabeled samples. However, existing SSL approaches for multi-modal medical imaging often struggle to exploit the complementary information between modalities due to semantic discrepancies and misalignment across MRI sequences. To address this, we propose a novel semi-supervised multi-modal framework that explicitly enhances modality-specific representations and facilitates adaptive cross-modal information fusion. Specifically, we introduce a Modality-specific Enhancing Module (MEM) to strengthen semantic cues unique to each modality via channel-wise attention, and a learnable Complementary Information Fusion (CIF) module to adaptively exchange complementary knowledge between modalities. The overall framework is optimized using a hybrid objective combining supervised segmentation loss and cross-modal consistency regularization on unlabeled data. Extensive experiments on the BraTS 2019 (HGG subset) demonstrate that our method consistently outperforms strong semi-supervised and multi-modal baselines under 1\%, 5\%, and 10\% labeled data settings, achieving significant improvements in both Dice and Sensitivity scores. Ablation studies further confirm the complementary effects of our proposed MEM and CIF in bridging cross-modality discrepancies and improving segmentation robustness under scarce supervision.

cs.CV↗

Modal analysis of blood flows in saccular aneurysms

Currently, it is challenging to investigate aneurismal hemodynamics based on current in-vivo data such as Magnetic Resonance Imaging or Computed Tomography due to the limitations in both spatial and temporal resolutions. In this work, we investigate the use of modal analysis at various resolutions to examine its usefulness for analyzing blood flows in brain aneurysms. Two variants of Dynamic Mode Decomposition (DMD): (i) Hankel-DMD; and (ii) Optimized-DMD, are used to extract the time-dependent dynamics of blood flows during one cardiac cycle. First, high-resolution hemodynamic data in patient-specific aneurysms are obtained using Computational Fluid Dynamics. Second, the dynamics modes, along with their spatial amplitudes and temporal magnitudes are calculated using the DMD analysis. Third, an examination of DMD analyses using a range of spatial and temporal resolutions of hemodynamic data to validate the applicability of DMD for low-resolution data, similar to ones in clinical practices. Our results show that DMD is able to characterize the inflow jet dynamics by separating large-scale structures and flow instabilities even at low spatial and temporal resolutions. Its robustness in quantifying the flow dynamics using the energy spectrum is demonstrated across different resolutions in all aneurysms in our study population. Our work indicates that DMD can be used for analyzing blood flow patterns of brain aneurysms and is a promising tool to be explored in in-vivo.

physics.flu-dyn↗