SearcharxivSearch

arXiv subjects

Zijian Gu

Publications and source records attributed to Zijian Gu.

6 recordsLinked to original sources

UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment

Multi-modal learning combining medical images and clinical text is promising for disease diagnosis. However, standard multi-modal training leads to shortcut learning: models exploit the easier modality (e.g., diagnostic cues in text) while neglecting harder-to-learn features (e.g., subtle visual patterns). We propose UniMod, a framework that mitigates shortcut learning by requiring each modality to predict the diagnosis on its own. It supervises image-only, text-only, and multi-modal classification simultaneously, so each modality must extract diagnostic features. We add cross-modality alignment for knowledge transfer and within-modality supervised contrastive alignment over same-diagnosis patients. On Harvard-Glaucoma, UniMod reaches 0.850 AUC, outperforming OGM-GE and Gradient Blending by 1.6-1.8%; on CheXpert Plus, it reaches 0.966 AUC, surpassing them by over 5%. UniMod also extends to 5-class multi-label diagnosis without architectural change, improving mean AUC by 0.097 over CGGM.

cs.CV

AMSB in $Sp(N_c)$ Gauge Theories

We present a careful study of the chiral symmetry breaking minima and other potential minima in supersymmetric symplectic QCD ($Sp(N_c)$ with $N_f$ flavors) perturbed by Anomaly Mediated Supersymmetry Breaking (AMSB). Although the case of $N_f = N_c +1$ requires particular care due to the inherently strongly coupled nature of the quantum modified moduli space, we are able to show that all $Sp(N_c)$ theories to which AMSB can be applied ($N_f < 3(N_c + 1)$) possess stable chiral symmetry breaking minima, which are plausibly continuously connected to the vacua of QCD-like $Sp(N_c)$ theories for large SUSY breaking, and are protected from runaways to incalculable minima.

hep-th

Fairness-Aware Fine-Tuning of Vision-Language Models for Medical Glaucoma Diagnosis

Vision-language models achieve expert-level performance on medical imaging tasks but exhibit significant diagnostic accuracy disparities across demographic groups. We introduce fairness-aware Low-Rank Adaptation for medical VLMs, combining parameter efficiency with explicit fairness optimization. Our key algorithmic contribution is a differentiable MaxAccGap loss that enables end-to-end optimization of accuracy parity across demographic groups. We propose three methods: FR-LoRA integrates MaxAccGap regularization into the training objective, GR-LoRA applies inverse frequency weighting to balance gradient contributions, and Hybrid-LoRA combines both mechanisms. Evaluated on 10,000 glaucoma fundus images, GR-LoRA reduces diagnostic accuracy disparities by 69% while maintaining 53.15% overall accuracy. Ablation studies reveal that strong regularization strength achieves optimal fairness with minimal accuracy trade-off, and race-specific optimization yields 60% disparity reduction. Our approach requires only 0.24% trainable parameters, enabling practical deployment of fair medical AI in resource-constrained healthcare settings.

cs.CV

Fairness in Multi-modal Medical Diagnosis with Demonstration Selection

Multimodal large language models (MLLMs) have shown strong potential for medical image reasoning, yet fairness across demographic groups remains a major concern. Existing debiasing methods often rely on large labeled datasets or fine-tuning, which are impractical for foundation-scale models. We explore In-Context Learning (ICL) as a lightweight, tuning-free alternative for improving fairness. Through systematic analysis, we find that conventional demonstration selection (DS) strategies fail to ensure fairness due to demographic imbalance in selected exemplars. To address this, we propose Fairness-Aware Demonstration Selection (FADS), which builds demographically balanced and semantically relevant demonstrations via clustering-based sampling. Experiments on multiple medical imaging benchmarks show that FADS consistently reduces gender-, race-, and ethnicity-related disparities while maintaining strong accuracy, offering an efficient and scalable path toward fair medical image reasoning. These results highlight the potential of fairness-aware in-context learning as a scalable and data-efficient solution for equitable medical image reasoning.

cs.CV

HGSFusion: Radar-Camera Fusion with Hybrid Generation and Synchronization for 3D Object Detection

Millimeter-wave radar plays a vital role in 3D object detection for autonomous driving due to its all-weather and all-lighting-condition capabilities for perception. However, radar point clouds suffer from pronounced sparsity and unavoidable angle estimation errors. To address these limitations, incorporating a camera may partially help mitigate the shortcomings. Nevertheless, the direct fusion of radar and camera data can lead to negative or even opposite effects due to the lack of depth information in images and low-quality image features under adverse lighting conditions. Hence, in this paper, we present the radar-camera fusion network with Hybrid Generation and Synchronization (HGSFusion), designed to better fuse radar potentials and image features for 3D object detection. Specifically, we propose the Radar Hybrid Generation Module (RHGM), which fully considers the Direction-Of-Arrival (DOA) estimation errors in radar signal processing. This module generates denser radar points through different Probability Density Functions (PDFs) with the assistance of semantic information. Meanwhile, we introduce the Dual Sync Module (DSM), comprising spatial sync and modality sync, to enhance image features with radar positional information and facilitate the fusion of distinct characteristics in different modalities. Extensive experiments demonstrate the effectiveness of our approach, outperforming the state-of-the-art methods in the VoD and TJ4DRadSet datasets by $6.53\%$ and $2.03\%$ in RoI AP and BEV AP, respectively. The code is available at https://github.com/garfield-cpp/HGSFusion.

cs.CV

Designing transformation-induced plasticity and twinning-induced plasticity Cr-Co-Ni medium entropy alloys: theory and experiment

In order to efficiently explore the nearly infinite composition space in multicomponent solid solution alloys, it is important to establish predictive design strategies and use computation-aided methods. In the present work, we demonstrated the density functional theory calculations informed design routes for realizing transformation-induced plasticity (TRIP) and twinning-induced plasticity (TWIP) in Cr-Co-Ni medium entropy alloys (MEAs). We systematically studied the effects of magnetism and chemical composition on the generalized stacking fault energy surface (gamma-surface) and showed that both chemistry and the coupled magnetic state strongly affect the gamma-surface, consequently, the primary deformation modes. Based on the calculated effective energy barriers for the competing deformation modes, we constructed composition and magnetism dependent deformation maps at both room and cryogenic temperatures. Accordingly, we proposed various design routes for achieving desired primary deformation modes in the ternary Cr-Co-Ni alloys. The deformation mechanisms predicted by our theoretical models are in nice agreement with available experimental observations in literature. Furthermore, we fabricated two non-equiatomic Cr-Co-Ni MEAs possessing the designed TWIP and TRIP effects, showing excellent combinations of tensile strength and ductility.

cond-mat.mtrl-sci