SearcharxivSearch

arXiv subjects

Bing Ji

Publications and source records attributed to Bing Ji.

8 recordsLinked to original sources

ELDiff: When Evidential Learning Meets Text-to-Image Diffusion

In multi-object text-to-image (T2I) diffusion, ensuring semantic consistency between textual prompts and generated visual content is crucial for image synthesis. However, such consistency constraint is often underemphasized in the denoising process of diffusion models. Although token supervised diffusion models can mitigate this issue by learning object-wise consistency between the image content and object segmentation maps, it tends to suffer from the problems of segmentation map bias and semantic overlap conflict, especially when involving multiple objects. In this paper, we propose ELDiff, a new evidential learning-supervised T2I diffusion model, which leverages the advantages of uncertainty metric and conflict detection to enhance the fault tolerance of unreliable segmentation maps and suppress semantic conflicts, strengthening object-wise consistency learning. Specifically, a pixel evidence loss is proposed to restrain overconfidence in unreliable labels through evidential regularization, and a token conflict loss is designed to weaken the contradiction between semantics through optimizing a measured conflict factor. Extensive experiments show that our ELDiff outperforms existing training based and train-free based T2I diffusion models on SD v1.4, SD v2.1, SDXL, SD v3.5, and Qwen-Image, without requiring additional inference-time manipulations. Notably, ELDiff can be seamlessly extended to the existing training pipeline of T2I diffusion models. Code can be found at https://github.com/QingtaoPan/ELDiff.

cs.CV

CPS4: Class Prompt driven Semi-Supervised Spine Segmentation with Class-specific Consistency Constraint

Vision Language Model (VLM) has great potential to enhance the quality of pseudo labels in semi-supervised spine segmentation by leveraging textual class prompts to generate segmentation map, but no one has studied it yet. Although promising, it lacks explicit constraints to ensure consistency between spine class prompts and spine unit region, resulting in unsatisfactory performance in multi-class segmentation map generation. In this paper, we propose CPS4, the first text-guided semi-supervised spine segmentation network using class prompts to enhance the quality of spine pseudo labels. Specifically, CPS4 is implemented through two training stages. (i) Class-specific consistency constrained VLM pretraining stage: we propose token- and pixel-level attention loss to optimize the consistency between class prompts and spine units, forcing the textual class prompt to be closely coupled with the target spine unit in the semantic space. (ii) Class Prompt driven semi-supervised spine segmentation stage: using the pretrained vision-text encoder, we derive each class-specific binary segmentation map for the unlabeled spine image and integrate them into an unified multi-class segmentation map, improving the quality of the spine pseudo label generated by the semi-supervised spine segmentation network. Experimental results show that our CPS4 achieves superior spine segmentation performance with Dice of 80.44%, only using 5% labeled data on the public spine segmentation dataset, surpassing popular semi-supervised learning and VLM methods. Our code will be available.

cs.CV

MedLoc-R1: Performance-Aware Curriculum Reward Scheduling for GRPO-Based Medical Visual Grounding

Medical visual grounding serves as a crucial foundation for fine-grained multimodal reasoning and interpretable clinical decision support. Despite recent advances in reinforcement learning (RL) for grounding tasks, existing approaches such as Group Relative Policy Optimization~(GRPO) suffer from severe reward sparsity when directly applied to medical images, primarily due to the inherent difficulty of localizing small or ambiguous regions of interest, which is further exacerbated by the rigid and suboptimal nature of fixed IoU-based reward schemes in RL. This leads to vanishing policy gradients and stagnated optimization, particularly during early training. To address this challenge, we propose MedLoc-R1, a performance-aware reward scheduling framework that progressively tightens the reward criterion in accordance with model readiness. MedLoc-R1 introduces a sliding-window performance tracker and a multi-condition update rule that automatically adjust the reward schedule from dense, easily obtainable signals to stricter, fine-grained localization requirements, while preserving the favorable properties of GRPO without introducing auxiliary networks or additional gradient paths. Experiments on three medical visual grounding benchmarks demonstrate that MedLoc-R1 consistently improves both localization accuracy and training stability over GRPO-based baselines. Our framework offers a general, lightweight, and effective solution for RL-based grounding in high-stakes medical applications. Code \& checkpoints are available at \hyperlink{}{https://github.com/MembrAI/MedLoc-R1}.

cs.CV

Look Closer! An Adversarial Parametric Editing Framework for Hallucination Mitigation in VLMs

While Vision-Language Models (VLMs) have garnered increasing attention in the AI community due to their promising practical applications, they exhibit persistent hallucination issues, generating outputs misaligned with visual inputs. Recent studies attribute these hallucinations to VLMs' over-reliance on linguistic priors and insufficient visual feature integration, proposing heuristic decoding calibration strategies to mitigate them. However, the non-trainable nature of these strategies inherently limits their optimization potential. To this end, we propose an adversarial parametric editing framework for Hallucination mitigation in VLMs, which follows an \textbf{A}ctivate-\textbf{L}ocate-\textbf{E}dit \textbf{A}dversarially paradigm. Specifically, we first construct an activation dataset that comprises grounded responses (positive samples attentively anchored in visual features) and hallucinatory responses (negative samples reflecting LLM prior bias and internal knowledge artifacts). Next, we identify critical hallucination-prone parameter clusters by analyzing differential hidden states of response pairs. Then, these clusters are fine-tuned using prompts injected with adversarial tuned prefixes that are optimized to maximize visual neglect, thereby forcing the model to prioritize visual evidence over inherent parametric biases. Evaluations on both generative and discriminative VLM tasks demonstrate the significant effectiveness of ALEAHallu in alleviating hallucinations. Our code is available at https://github.com/hujiayu1223/ALEAHallu.

cs.CV

Discrete-Time CRLB-based Power Allocation for CF MIMO-ISAC with Joint Localization and Velocity Sensing

In this paper, we investigate integrated sensing and communication (ISAC) in a cell-free (CF) multiple-input multiple-output (MIMO) network, where each access point functions either as an ISAC transmitter or as a sensing receiver. We devote into the ISAC sensing metric using the discrete-time signal-based Cramer-Rao lower bounds (CRLBs) for joint location and velocity estimation under arbitrary power allocation ratios under the deterministic radar cross section assumption (RCS). Then, we consider the power allocation optimization problem for the CF MIMO-ISAC as the maximization of the communication signal-to-interference-plus-noise ratio (SINR), subject to CRLB-based sensing constraints and per-transmitter power limits. To solve the resulting nonlinear and non-convex problem, we propose a penalty function and projection-based modified conjugate gradient algorithm with inexact line search (PP-MCG-ILS), and an alternative method based on a modified steepest descent approach (PP-MSD-ILS). We show that the proposed algorithms are scalable and can be extended to a broad class of optimization problems involving nonlinear inequality constraints and affine equality constraints. In addition, we extend the PP-MCG-ILS algorithm to the pure sensing scenario, where a penalty function-based normalized conjugate gradient algorithm (P-NCG-ILS) is developed for sensing power minimization. Finally, we analyze the convergence behavior and qualitatively compare the computational complexity of the proposed algorithms. Simulation results confirm the accuracy of the derived CRLBs and demonstrate the effectiveness of the proposed power allocation strategies in enhancing both sensing and overall ISAC performance.

eess.SP

Joint Location and Velocity Estimation and Fundamental CRLB Analysis for Cell-Free MIMO-ISAC

This paper presents a fundamental performance analysis of joint location and velocity estimation in a cell-free (CF) MIMO integrated sensing and communication (ISAC) system. Unlike prior studies that primarily rely on continuous-time signal models, we consider a more practical and challenging scenario in the discrete-time digital domain. Specifically, we first formulate a logarithmic likelihood function (LLF) and corresponding maximum likelihood estimation (MLE) for both single- and multiple-target sensing. Building upon the proposed LLF framework, closed-form Cramer-Rao lower bounds (CRLBs) for joint location and velocity estimation are derived under deterministic, unknown, and spatially varying radar cross-section (RCS) models. These CRLBs can serve as a fundamental performance metric to guide CF MIMO-ISAC system design. To enhance tractability, we also develop a class of simplified closed-form CRLBs, referred to as approximate CRLBs, along with a rigorous analysis of the conditions under which they remain accurate. Furthermore, we investigate how the sampling rate, squared effective bandwidth, and time width influence CRLB performance. For multi-target scenarios, the concepts of safety distance and safety velocity are introduced to characterize the conditions under which the CRLBs converge to their single-target counterparts. Extensive simulations using orthogonal frequency division multiplexing (OFDM) and orthogonal chirp division multiplexing (OCDM) validate the theoretical findings and provide practical insights for CF MIMO-ISAC system design

eess.SP

DuSSS: Dual Semantic Similarity-Supervised Vision-Language Model for Semi-Supervised Medical Image Segmentation

Semi-supervised medical image segmentation (SSMIS) uses consistency learning to regularize model training, which alleviates the burden of pixel-wise manual annotations. However, it often suffers from error supervision from low-quality pseudo labels. Vision-Language Model (VLM) has great potential to enhance pseudo labels by introducing text prompt guided multimodal supervision information. It nevertheless faces the cross-modal problem: the obtained messages tend to correspond to multiple targets. To address aforementioned problems, we propose a Dual Semantic Similarity-Supervised VLM (DuSSS) for SSMIS. Specifically, 1) a Dual Contrastive Learning (DCL) is designed to improve cross-modal semantic consistency by capturing intrinsic representations within each modality and semantic correlations across modalities. 2) To encourage the learning of multiple semantic correspondences, a Semantic Similarity-Supervision strategy (SSS) is proposed and injected into each contrastive learning process in DCL, supervising semantic similarity via the distribution-based uncertainty levels. Furthermore, a novel VLM-based SSMIS network is designed to compensate for the quality deficiencies of pseudo-labels. It utilizes the pretrained VLM to generate text prompt guided supervision information, refining the pseudo label for better consistency regularization. Experimental results demonstrate that our DuSSS achieves outstanding performance with Dice of 82.52%, 74.61% and 78.03% on three public datasets (QaTa-COV19, BM-Seg and MoNuSeg).

cs.CV

Denoising Magnetic Resonance Spectroscopy (MRS) Data Using Stacked Autoencoder for Improving Signal-to-Noise Ratio and Speed of MRS

Background: Magnetic resonance spectroscopy (MRS) enables non-invasive detection and measurement of biochemicals and metabolites. However, MRS has low signal-to-noise ratio (SNR) when concentrations of metabolites are in the range of the million molars. Standard approach of using a high number of signal averaging (NSA) to achieve sufficient NSR comes at the cost of a long acquisition time. Purpose: We propose to use deep-learning approaches to denoise MRS data without increasing the NSA. Methods: The study was conducted using data collected from the brain spectroscopy phantom and human subjects. We utilized a stack auto-encoder (SAE) network to train deep learning models for denoising low NSA data (NSA = 1, 2, 4, 8, and 16) randomly truncated from high SNR data collected with high NSA (NSA=192) which were also used to obtain the ground truth. We applied both self-supervised and fully-supervised training approaches and compared their performance of denoising low NSA data based on improved SNRs. Results: With the SAE model, the SNR of low NSA data (NSA = 1) obtained from the phantom increased by 22.8% and the MSE decreased by 47.3%. For low NSA images of the human parietal and temporal lobes, the SNR increased by 43.8% and the MSE decreased by 68.8%. In all cases, the chemical shift of NAA in the denoised spectra closely matched with the high SNR spectra, suggesting no distortion to the spectra from denoising. Furthermore, the denoising performance of the SAE model was more effective in denoising spectra with higher noise levels. Conclusions: The reported SAE denoising method is a model-free approach to enhance the SNR of low NSA MRS data. With the denoising capability, it is possible to acquire MRS data with a few NSA, resulting in shorter scan times while maintaining adequate spectroscopic information for detecting and quantifying the metabolites of interest.

physics.med-ph