Searcharxiv⌕ Search

arXiv subjects

Xuan Cuong Ngo

Publications and source records attributed to Xuan Cuong Ngo.

4 recordsLinked to original sources

DualCount: Structurally Consistent Density and Point Modeling for Zero-Shot Object Counting

Zero-shot object counting aims to estimate the number of objects specified by a text query without category-specific training. Recent approaches primarily rely on density regression or detection-style instance prediction. While effective, density-based models often suffer from spatial ambiguity and background leakage due to weakly regulated mass allocation, leading to fragmented or part-biased representations that increase counting error in complex scenes. In this work, we propose an instance-aware dual-decoder framework that structurally couples density and point representations for zero-shot object counting. Instead of treating density estimation as independent pixel-wise regression, we interpret it as a structured mass allocation problem over a latent set of object instances. Predicted instance centers induce a soft instance-wise decomposition of the density map, upon which we enforce two geometric constraints: (1) per-instance mass conservation, ensuring each object contributes approximately one unit of density mass, and (2) center-of-mass alignment, encouraging each density component to concentrate around its corresponding predicted center. These constraints introduce instance-level geometric consistency and lead to more accurate mass allocation, thereby reducing counting error. Extensive experiments on FSC-147, PUCPR+, and CARPK show that our approach consistently reduces counting error and establishes new state-of-the-art performance in zero-shot object counting.

cs.CV↗

Learning to Adapt and Calibrate: Score Distribution Alignment for Few-Shot Uncertainty Prediction in Medical VLMs

Uncertainty estimation for medical vision--language models (VLMs) using conformal prediction has gained increasing attention due to its distribution-free coverage guarantees. However, standard conformal prediction relies on exchangeability between calibration and test data and typically requires a sufficiently large calibration set to obtain reliable coverage. These assumptions are difficult to satisfy in few-shot transfer settings, where only a small labeled support set is available to adapt a pretrained VLM to a new medical task, while an unlabeled query set is used for evaluation. Supervised fine-tuning on the support set changes the model parameters and consequently shifts the nonconformity score distribution, breaking exchangeability between calibration and query samples and leading to unreliable coverage under distribution shift. Existing transductive conformal adaptation methods often preserve validity by avoiding supervised updates. While this helps maintain conformal assumptions, it underutilizes the scarce labeled support data and limits task adaptation, which is the primary objective in few-shot learning. In this setting, conformal prediction should serve as an uncertainty estimation layer that supports the adapted model, rather than preventing adaptation itself. To this end, we propose AlignCP, a framework that reconciles supervised few-shot adaptation with conformal uncertainty estimation under non-exchangeability. AlignCP learns a reweighted calibration distribution that reduces the score-level discrepancy between the labeled support set and the unlabeled query set. By aligning the one-dimensional nonconformity score distributions, AlignCP aims to close the coverage gap induced by adaptation without requiring query labels.

cs.CV↗

GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models

Activation steering provides a lightweight way to control large language models (LLMs) by modifying their hidden activations at inference time. Among these approaches, norm-preserving steering aims to change model behavior without altering the activation norm, reducing the risk of representation collapse and degradation. However, existing norm-preserving methods are limited by predefined steering trajectories and by their reliance on one-step updates, which may fail to capture the complex structure of activation distributions. We propose GeoSteer, an optimization-based method for norm-preserving activation steering. GeoSteer formulates steering as a Riemannian optimization problem and updates activations through a sequence of small geodesic steps on the representation manifold. To avoid fixed steering directions, GeoSteer learns a nonlinear activation-space objective that distinguishes desired from undesired activations, and uses this function to adaptively guide each steering step. This multistep formulation yields smoother, more stable, and more consistent steering behavior while preserving the activation norm. Across TruthfulQA, RealToxicityPrompts, and UltraFeedback benchmarks, GeoSteer consistently improves over state-of-the-art activation steering baselines. These results suggest that norm-preserving steering can be made more effective by replacing predefined one-step edits with adaptive, geometry-aware optimization.

cs.LG↗

Learning to See Locally and Align Clinically with Pathology Semantics for Radiology Report Generation

Recent radiology-adapted vision-language models have achieved strong performance on standard report generation benchmarks, yet their robustness and generalization remain constrained by imperfect alignment and correlation between visual and textual features. Existing methods connect image and text either implicitly through autoregressive report supervision or explicitly through contrastive learning. However, autoregressive supervision alone is insufficient to establish reliable image-text alignment, while contrastive learning can push apart unpaired reports that describe related pathologies simply because they are not paired with the same image. This is problematic in radiology, where different reports may share compatible pathology semantics rather than being true negatives. As a result, the learned representation may fail to organize images and reports around shared pathology concepts, causing the decoder to rely on pretrained language priors and generate clinically plausible reports that are not fully supported by radiographic evidence. To address this issue, we propose PALM, a pathology-aware alignment framework for radiology report generation. Instead of directly matching each image-report pair while separating all others, PALM aligns visual and textual features through shared pathology prototypes. These prototypes provide a clinically meaningful bridge between radiographic evidence and textual findings, allowing cases with similar pathology semantics to move toward common concepts without separating compatible cases. In addition, we introduce Masked Evidence Modeling to strengthen the image encoder sensitivity to local radiographic evidence by learning semantic changes caused by masked image regions. Experiments on MIMIC-CXR, IU X-Ray, and MIMIC-ABN show that PALM consistently improves both report generation and abnormality-focused robustness.

eess.IV↗