SearcharxivSearch

arXiv subjects

Chen Qin

Publications and source records attributed to Chen Qin.

At least 19 recordsLinked to original sources

Learning from Noisy Prompts: Saliency-Guided Prompt Distillation for Robust Segmentation with SAM

Segmentation is central to clinical diagnosis and monitoring, yet the reliability of modern foundation models in medical imaging still depends on the availability of precise prompts. The Segment Anything Model (SAM) offers powerful zero-shot capabilities, although it collapses under the weak, generic, and noisy prompts that dominate real clinical workflows. In practice, annotations such as centerline points are coarse and ambiguous, often drifting across neighboring anatomy and misguiding SAM toward inconsistent or incomplete masks. We introduce SPD, a Saliency-Guided Prompt Distillation framework that converts these unreliable cues into robust guidance. SPD first learns data-driven anatomical priors through a lightweight saliency head to obtain confident localization maps. These priors then drive Contextual Prompt Distillation, which validates and enriches noisy prompts using cues from anatomically adjacent slices, producing a consensus prompt set that matches the behavior of expert reasoning. A Pairwise Slice Consistency objective further enforces local anatomical coherence during segmentation. Experiments on four challenging MRI and CT benchmarks demonstrate that SPD consistently outperforms existing SAM adaptations and supervised baselines, delivering large gains in both region-based and boundary-based metrics. SPD provides a practical and principled path toward reliable foundation model deployment in clinical environments where only imperfect prompts are available.

cs.CV

Active Galactic Nuclei and STaR fOrmation in Nearby Galaxies AGNSTRONG. III. A Study on Ionized and Warm Molecular Gas Outflows of 6 Type-2 AGNs

Active galactic nucleus (AGN)-driven gas outflows are one of the best tracers of AGN feedback in action, as these powerful outflows expel/heat or compress the surrounding interstellar medium (ISM), thus quenching or enhancing star-forming activity in their hosts. Studying the kinematics of outflows in different gas phases is crucial for comprehending how AGNs impact the ISM within their host galaxies. However, the differences in the physical natures of ionized and warm molecular gas outflows remain largely unexplored. To obtain a complete picture of AGN outflows and their feedback effects, we present a study of both ionized and warm molecular gas outflows in six type-2 AGNs ($z<0.1$) that exhibit strong ionized outflows in previous optical observations. Utilizing the Triple Spectrograph and Double Spectrograph instruments on the Palomar 200-inch Hale Telescope, we conduct spatially resolved measurements in the slit direction of strong emission lines from both ionized and warm molecular gas, such as $\rm [O\ III]$, $\rm Pa\alpha$, $\rm H_{2}$ 1-0 S(1), etc., allowing for a direct comparison of their outflow properties. One out of six AGNs shows significant ionized and warm molecular outflows in near-infrared bands, exhibiting the most powerful kinematics and highest luminosity. A positive correlation between the kinematics and AGN luminosity is shown, suggesting that more luminous AGNs, which reflect higher levels of AGN activity, tend to have a greater impact on the gases, probably driving the outflows.

astro-ph.GA

Inference-Time Dynamic Modality Selection for Incomplete Multimodal Classification

Multimodal deep learning (MDL) has achieved remarkable success across various domains, yet its practical deployment is often hindered by incomplete multimodal data. Existing incomplete MDL methods either discard missing modalities, risking the loss of valuable task-relevant information, or recover them, potentially introducing irrelevant noise, leading to the discarding-imputation dilemma. To address this dilemma, in this paper, we propose DyMo, a new inference-time dynamic modality selection framework that adaptively identifies and fuses reliable recovered modalities, fully exploring task-relevant information beyond the conventional discard-or-impute paradigm. Central to DyMo is a novel selection algorithm that maximizes multimodal task-relevant information for each test sample. Since direct estimation of such information at test time is intractable due to the unknown data distribution, we theoretically establish a connection between information and the task loss, which we compute at inference time as a tractable proxy. Building on this, a novel principled reward function is proposed to guide modality selection. In addition, we design a flexible multimodal network architecture compatible with arbitrary modality combinations, alongside a tailored training strategy for robust representation learning. Extensive experiments on diverse natural and medical image datasets show that DyMo significantly outperforms state-of-the-art incomplete/dynamic MDL methods across various missing-data scenarios. Our code is available at https://github.com//siyi-wind/DyMo.

cs.CV

Self-Supervised Slice-to-Volume Reconstruction with Gaussian Representations for Fetal MRI

Reconstructing 3D fetal MR volumes from motion-corrupted stacks of 2D slices is a crucial and challenging task. Conventional slice-to-volume reconstruction (SVR) methods are time-consuming and require multiple orthogonal stacks for reconstruction. While learning-based SVR approaches have significantly reduced the time required at the inference stage, they heavily rely on ground truth information for training, which is inaccessible in practice. To address these challenges, we propose GaussianSVR, a self-supervised framework for slice-to-volume reconstruction. GaussianSVR represents the target volume using 3D Gaussian representations to achieve high-fidelity reconstruction. It leverages a simulated forward slice acquisition model to enable self-supervised training, alleviating the need for ground-truth volumes. Furthermore, to enhance both accuracy and efficiency, we introduce a multi-resolution training strategy that jointly optimizes Gaussian parameters and spatial transformations across different resolution levels. Experiments show that GaussianSVR outperforms the baseline methods on fetal MR volumetric reconstruction. Code is available at https://github.com/Yinsong0510/GaussianSVR-Self-Supervised-Slice-to-Volume-Reconstruction-with-Gaussian-Representations.

cs.CV

Active Galactic Nuclei and STaR fOrmation in Nearby Galaxies (AGNSTRONG). II: Results for Jetted Type-I AGNs with Strong Ionized Gas Outflows

We investigate the correlation between ionized gas outflows, jets, and star formation in a sample of 42 local type-I active galactic nuclei (AGNs) exhibiting significant [O III] outflows. This study uses both new submillimeter (sub-mm) observations and archival data from the James Clerk Maxwell Telescope. Our analysis, which includes a correction for jet emission in the sub-mm bands, fitting spectral energy distribution, and analyzing spectra, enables us to derive star-formation rates (SFRs) through various methods. By comparing radio power and SFRs, we select a sub-sample of jetted AGNs of which radio emission is mostly from the jets. We find that jetted AGNs predominantly lie above the main sequence of star-forming galaxies, suggesting a correlation between jet activity and star formation. By comparing dust extinction, we demonstrate that jetted AGNs do not have more dust which is the fuel of both star formation and AGN activity. Therefore, this correlation is more likely to arise from AGN feedback. We also find that the Eddington ratio does not impact the specific SFRs (sSFRs) of our sample. Additionally, for jetted AGNs, stronger radio emission corresponds to higher sSFRs, suggesting that jet emission may promote star formation, i.e., positive feedback. Our results not only shed light on the feedback mechanisms of AGNs but also underscore the complex interplay between black hole activity and star formation in galaxy evolution.

astro-ph.GA

Adaptive Conditional Contrast-Agnostic Deformable Image Registration with Uncertainty Estimation

Deformable multi-contrast image registration is a challenging yet crucial task due to the complex, non-linear intensity relationships across different imaging contrasts. Conventional registration methods typically rely on iterative optimization of the deformation field, which is time-consuming. Although recent learning-based approaches enable fast and accurate registration during inference, their generalizability remains limited to the specific contrasts observed during training. In this work, we propose an adaptive conditional contrast-agnostic deformable image registration framework (AC-CAR) based on a random convolution-based contrast augmentation scheme. AC-CAR can generalize to arbitrary imaging contrasts without observing them during training. To encourage contrast-invariant feature learning, we propose an adaptive conditional feature modulator (ACFM) that adaptively modulates the features and the contrast-invariant latent regularization to enforce the consistency of the learned feature across different imaging contrasts. Additionally, we enable our framework to provide contrast-agnostic registration uncertainty by integrating a variance network that leverages the contrast-agnostic registration encoder to improve the trustworthiness and reliability of AC-CAR. Experimental results demonstrate that AC-CAR outperforms baseline methods in registration accuracy and exhibits superior generalization to unseen imaging contrasts. Code is available at https://github.com/Yinsong0510/AC-CAR.

cs.CV

Enabling Ultra-Fast Cardiovascular Imaging Across Heterogeneous Clinical Environments with A Generalist Foundation Model and Multimodal Database

Multimodal cardiovascular magnetic resonance (CMR) imaging provides comprehensive and non-invasive insights into cardiovascular disease (CVD) diagnosis and underlying mechanisms. Despite decades of advancements, its widespread clinical adoption remains constrained by prolonged scan times, inconsistent image quality, and heterogeneity across medical environments. This underscores the urgent need for a generalist reconstruction foundation model for ultra-fast CMR imaging, one formulated for physics-constrained inverse problems in the sensor (k-space) domain, capable of adapting across diverse imaging scenarios and serving as the essential substrate for all downstream analyses. To enable this goal, we curate MMCMR-427K, the largest and most comprehensive multimodal CMR k-space database to date, comprising 427,465 multi-coil k-space data paired with structured metadata across 13 international centers, 12 CMR modalities, 15 scanners spanning four field strengths, and 17 CVD categories in populations across three continents. Building on this unprecedented resource, we introduce CardioMM, a generalist reconstruction foundation model capable of dynamically adapting to heterogeneous fast CMR imaging scenarios. CardioMM unifies semantic contextual understanding with physics-informed data consistency to deliver robust reconstructions across varied scanners, protocols, and patient presentations. Comprehensive evaluations demonstrate that CardioMM achieves state-of-the-art performance across internal centers and exhibits strong zero-shot generalization to unseen external settings. Importantly, CardioMM supports acceleration up to 24x, providing the first evidence that such extreme acquisition speed can preserve key cardiac phenotypes, quantitative myocardial biomarkers, and diagnostic image quality without compromising clinical integrity.

eess.IV

UPMRI: Unsupervised Parallel MRI Reconstruction via Projected Conditional Flow Matching

Reconstructing high-quality images from substantially undersampled k-space data for accelerated MRI presents a challenging ill-posed inverse problem. While supervised deep learning has revolutionized this field, it relies heavily on large datasets of fully sampled ground-truth images, which are often impractical or impossible to acquire in clinical settings due to long scan times. Despite advances in self-supervised/unsupervised MRI reconstruction, their performance remains inadequate at high acceleration rates. To bridge this gap, we introduce UPMRI, an unsupervised reconstruction framework based on Projected Conditional Flow Matching (PCFM) and its unsupervised transformation. Unlike standard generative models, PCFM learns the prior distribution of fully sampled parallel MRI data by utilizing only undersampled k-space measurements. To reconstruct the image, we establish a novel theoretical link between the marginal vector field in the measurement space, governed by the continuity equation, and the optimal solution to the PCFM objective. This connection results in a cyclic dual-space sampling algorithm for high-quality reconstruction. Extensive evaluations on the fastMRI brain and CMRxRecon cardiac datasets demonstrate that UPMRI significantly outperforms state-of-the-art self-supervised and unsupervised baselines. Notably, it also achieves reconstruction fidelity comparable to or better than leading supervised methods at high acceleration factors, while requiring no fully sampled training data.

eess.IV

The Faintest, Extremely Variable X-ray Tidal Disruption Event from a Supermassive Black Hole Binary?

Tidal disruption events (TDEs), which occur when stars enter the tidal radii of supermassive black holes (SMBHs) and are subsequently torn apart by their tidal forces, represent intriguing phenomena that stimulate growing research interest and pose an increasing number of puzzles in the era of time-domain astronomy. Here we report an unusual X-ray transient, XID 935, discovered in the 7 Ms Chandra Deep Field-South, the deepest X-ray survey ever. XID 935 experienced an overall X-ray dimming by a factor of more than 40 between 1999 and 2016. Not monotonically decreasing during this period, its X-ray luminosity increased by a factor $> 27$ within 2 months, from $L_{\rm 0.5-7\ keV}<10^{40.87}$ erg s$^{-1}$ (10 October 2014 -- 4 January 2015) to $L_{\rm 0.5-7\ keV}=10^{42.31\pm 0.20}$ erg s$^{-1}$ (16 March 2015). The X-ray position of XID 935 is located at the center of its host galaxy with a spectroscopic redshift of 0.251, whose optical spectra do not display emission characteristics associated with an active galactic nucleus. The peak 0.5--2.0 keV flux is the faintest among all the X-ray-selected TDE candidates to date. Thanks to a total exposure of $\sim 9.5$ Ms in the X-ray bands, we manage to secure relatively well-sampled, 20-year-long X-ray light curves of this deepest X-ray-selected TDE candidate. We find that a partial TDE model could not explain the main declining trend. An SMBH binary TDE model is in acceptable accordance with the light curves of XID 935; however, it fails to match short-timescale fluctuations exactly. Therefore, the exceptional observational features of XID 935 provide a key benchmark for refining quantitative TDE models and simulations.

astro-ph.HE

Extreme Cardiac MRI Analysis under Respiratory Motion: Results of the CMRxMotion Challenge

Deep learning models have achieved state-of-the-art performance in automated Cardiac Magnetic Resonance (CMR) analysis. However, the efficacy of these models is highly dependent on the availability of high-quality, artifact-free images. In clinical practice, CMR acquisitions are frequently degraded by respiratory motion, yet the robustness of deep learning models against such artifacts remains an underexplored problem. To promote research in this domain, we organized the MICCAI CMRxMotion challenge. We curated and publicly released a dataset of 320 CMR cine series from 40 healthy volunteers who performed specific breathing protocols to induce a controlled spectrum of motion artifacts. The challenge comprised two tasks: 1) automated image quality assessment to classify images based on motion severity, and 2) robust myocardial segmentation in the presence of motion artifacts. A total of 22 algorithms were submitted and evaluated on the two designated tasks. This paper presents a comprehensive overview of the challenge design and dataset, reports the evaluation results for the top-performing methods, and further investigates the impact of motion artifacts on five clinically relevant biomarkers. All resources and code are publicly available at: https://github.com/CMRxMotion

eess.IV

Insights from the "Red devil" AT 2022fpx: A Dust-reddened Family of Tidal Disruption Events Excluded by Their Apparent Red Color?

We report unnoticed but intriguing features in the peculiar nuclear transient AT 2022fpx, and investigate its type. These features include the constantly red optical color of $g-r>0$, a stable soft X-ray flare ($kT\sim100$ eV) in the past $\sim$550 days, a prominent mid-infrared echo peaked at $\sim$$10^{43.3}$ erg s$^{-1}$ and the confirmation of a weak active galactic nucleus by weak flares in pre-event Wide-field Infrared Survey Explorer mid-infrared light curves with no contemporary optical, radio or X-ray counterparts. The combination of the optical red color and possible origin of a tidal disruption event (TDE) of AT 2022fpx is particularly attractive, as it challenges the most widely accepted and adopted "blue color" criterion for optical TDE selection. Although we still cannot confirm whether the red color is intrinsic, we do find that the "blue color" criterion can filter out normal TDEs whose optical-UV spectral energy distributions (SEDs) are either severely contaminated by prominent emission lines (especially H$\alpha$) or heavily dust-reddened. Hence, its potential selection effect may have been imprinted on the whole optical TDE family. Blackbody fitting on the optical (rest-frame $\sim$$4000-7000$ \AA) and optical-UV ($\sim$$2000-7000$ \AA) SEDs of four TDEs with high-cadence UV observations shows that $T_\mathrm{bb}$ rise by $\sim$40$-$110 \% when the UV bands are included. The power-law models ($f_{\lambda}\propto\lambda^{-\alpha}$ with $\alpha=2-3$) can fit the rest-frame $\sim$$2000-7000$ \AA SEDs more consistently, indicating that SEDs should peak at shorter wavelengths, but not simple blackbodies. Hence, the estimated released energy for the optical-UV bright but X-ray faint TDEs based on blackbody SED fitting should be significantly lower than the intrinsic energy.

astro-ph.HE

Spectral Bias Correction in PINNs for Myocardial Image Registration of Pathological Data

Accurate myocardial image registration is essential for cardiac strain analysis and disease diagnosis. However, spectral bias in neural networks impedes modeling high-frequency deformations, producing inaccurate, biomechanically implausible results, particularly in pathological data. This paper addresses spectral bias in physics-informed neural networks (PINNs) by integrating Fourier Feature mappings and introducing modulation strategies into a PINN framework. Experiments on two distinct datasets demonstrate that the proposed methods enhance the PINN's ability to capture complex, high-frequency deformations in cardiomyopathies, achieving superior registration accuracy while maintaining biomechanical plausibility - thus providing a foundation for scalable cardiac image registration and generalization across multiple patients and pathologies.

eess.IV

STiL: Semi-supervised Tabular-Image Learning for Comprehensive Task-Relevant Information Exploration in Multimodal Classification

Multimodal image-tabular learning is gaining attention, yet it faces challenges due to limited labeled data. While earlier work has applied self-supervised learning (SSL) to unlabeled data, its task-agnostic nature often results in learning suboptimal features for downstream tasks. Semi-supervised learning (SemiSL), which combines labeled and unlabeled data, offers a promising solution. However, existing multimodal SemiSL methods typically focus on unimodal or modality-shared features, ignoring valuable task-relevant modality-specific information, leading to a Modality Information Gap. In this paper, we propose STiL, a novel SemiSL tabular-image framework that addresses this gap by comprehensively exploring task-relevant information. STiL features a new disentangled contrastive consistency module to learn cross-modal invariant representations of shared information while retaining modality-specific information via disentanglement. We also propose a novel consensus-guided pseudo-labeling strategy to generate reliable pseudo-labels based on classifier consensus, along with a new prototype-guided label smoothing technique to refine pseudo-label quality with prototype embeddings, thereby enhancing task-relevant information learning in unlabeled data. Experiments on natural and medical image datasets show that STiL outperforms the state-of-the-art supervised/SSL/SemiSL image/multimodal approaches. Our code is available at https://github.com/siyi-wind/STiL.

cs.CV

Towards Modality- and Sampling-Universal Learning Strategies for Accelerating Cardiovascular Imaging: Summary of the CMRxRecon2024 Challenge

Cardiovascular health is vital to human well-being, and cardiac magnetic resonance (CMR) imaging is considered the {clinical reference standard} for diagnosing cardiovascular disease. However, its adoption is hindered by long scan times, complex contrasts, and inconsistent quality. While deep learning methods perform well on specific CMR imaging {sequences}, they often fail to generalize across modalities and sampling schemes. The lack of benchmarks for high-quality, fast CMR image reconstruction further limits technology comparison and adoption. The CMRxRecon2024 challenge, attracting over 200 teams from 18 countries, addressed these issues with two tasks: generalization to unseen {modalities} and robustness to diverse undersampling patterns. We introduced the largest public multi-{modality} CMR raw dataset, an open benchmarking platform, and shared code. Analysis of the best-performing solutions revealed that prompt-based adaptation and enhanced physics-driven consistency enabled strong cross-scenario performance. These findings establish principles for generalizable reconstruction models and advance clinically translatable AI in cardiovascular imaging.

eess.IV

Unsupervised Accelerated MRI Reconstruction via Ground-Truth-Free Flow Matching

Accelerated magnetic resonance imaging involves reconstructing fully sampled images from undersampled k-space measurements. Current state-of-the-art approaches have mainly focused on either end-to-end supervised training inspired by compressed sensing formulations, or posterior sampling methods built on modern generative models. However, their efficacy heavily relies on large datasets of fully sampled images, which may not always be available in practice. To address this issue, we propose an unsupervised MRI reconstruction method based on ground-truth-free flow matching (GTF$^2$M). Particularly, the GTF$^2$M learns a prior denoising process of fully sampled ground-truth images using only undersampled data. Based on that, an efficient cyclic reconstruction algorithm is further proposed to perform forward and backward integration in the dual space of image-space signal and k-space measurement. We compared our method with state-of-the-art learning-based baselines on the fastMRI database of both single-coil knee and multi-coil brain MRIs. The results show that our proposed unsupervised method can significantly outperform existing unsupervised approaches, and achieve performance comparable to most supervised end-to-end and prior learning baselines trained on fully sampled MRI, while offering greater efficiency than the compared generative model-based approaches.

eess.IV

Uncertainty quantification for White Matter Hyperintensity segmentation detects silent failures and improves automated Fazekas quantification

White Matter Hyperintensities (WMH) are key neuroradiological markers of small vessel disease present in brain MRI. Assessment of WMH is important in research and clinics. However, WMH are challenging to segment due to their high variability in shape, location, size, poorly defined borders, and similar intensity profile to other pathologies (e.g stroke lesions) and artefacts (e.g head motion). In this work, we assess the utility and semantic properties of the most effective techniques for uncertainty quantification (UQ) in segmentation for the WMH segmentation task across multiple test-time data distributions. We find UQ techniques reduce 'silent failure' by identifying in UQ maps small WMH clusters in the deep white matter that are unsegmented by the model. A combination of Stochastic Segmentation Networks with Deep Ensembles also yields the highest Dice and lowest Absolute Volume Difference % (AVD) score and can highlight areas where there is ambiguity between WMH and stroke lesions. We further demonstrate the downstream utility of UQ, proposing a novel method for classification of the clinical Fazekas score using spatial features extracted from voxelwise WMH probability and UQ maps. We show that incorporating WMH uncertainty information improves Fazekas classification performance and calibration. Our model with (UQ and spatial WMH features)/(spatial WMH features)/(WMH volume only) achieves a balanced accuracy score of 0.74/0.67/0.62, and root brier score of 0.65/0.72/0.74 in the Deep WMH and balanced accuracy of 0.74/0.73/0.71 and root brier score of 0.64/0.66/0.68 in the Periventricular region. We further demonstrate that stochastic UQ techniques with high sample diversity can improve the detection of poor quality segmentations.

eess.IV

Sampling-Pattern-Agnostic MRI Reconstruction through Adaptive Consistency Enforcement with Diffusion Model

Magnetic Resonance Imaging (MRI) is a powerful, non-invasive diagnostic tool; however, its clinical applicability is constrained by prolonged acquisition times. Whilst present deep learning-based approaches have demonstrated potential in expediting MRI processes, these methods usually rely on known sampling patterns and exhibit limited generalisability to novel patterns. In the paper, we propose a sampling-pattern-agnostic MRI reconstruction method via a diffusion model through adaptive consistency enforcement. Our approach effectively reconstructs high-fidelity images with varied under-sampled acquisitions, generalising across contrasts and acceleration factors regardless of sampling trajectories. We train and validate across all contrasts in the MICCAI 2024 Cardiac MRI Reconstruction Challenge (CMRxRecon) dataset for the ``Random sampling CMR reconstruction'' task. Evaluation results indicate that our proposed method significantly outperforms baseline methods.

eess.IV

SGSR: Structure-Guided Multi-Contrast MRI Super-Resolution via Spatio-Frequency Co-Query Attention

Magnetic Resonance Imaging (MRI) is a leading diagnostic modality for a wide range of exams, where multiple contrast images are often acquired for characterizing different tissues. However, acquiring high-resolution MRI typically extends scan time, which can introduce motion artifacts. Super-resolution of MRI therefore emerges as a promising approach to mitigate these challenges. Earlier studies have investigated the use of multiple contrasts for MRI super-resolution (MCSR), whereas majority of them did not fully exploit the rich contrast-invariant structural information. To fully utilize such crucial prior knowledge of multi-contrast MRI, in this work, we propose a novel structure-guided MCSR (SGSR) framework based on a new spatio-frequency co-query attention (CQA) mechanism. Specifically, CQA performs attention on features of multiple contrasts with a shared structural query, which is particularly designed to extract, fuse, and refine the common structures from different contrasts. We further propose a novel frequency-domain CQA module in addition to the spatial domain, to enable more fine-grained structural refinement. Extensive experiments on fastMRI knee data and low-field brain MRI show that SGSR outperforms state-of-the-art MCSR methods with statistical significance.

eess.IV