SearcharxivSearch

arXiv subjects

Alison D. Gernand

Publications and source records attributed to Alison D. Gernand.

4 recordsLinked to original sources

RulerNet: Learning Perspective-Invariant Ruler Representations for Robust Image Scale Estimation

Accurately converting pixel measurements into absolute real-world dimensions remains a fundamental challenge in computer vision, limiting progress in applications such as biomedicine, forensics, nutritional analysis, and e-commerce. We introduce RulerNet, a deep learning framework that robustly infers scale in the wild by reformulating ruler reading as a unified keypoint detection problem and representing rulers with geometric progression parameters that compactly approximate the non-uniform spacing induced by perspective transformations. Unlike traditional methods that rely on handcrafted thresholds or rigid, ruler-specific pipelines, RulerNet directly localizes centimeter markings using a mark-visibility-based annotation and training strategy that remains valid under mark-preserving transformations, enabling strong generalization across diverse ruler types and imaging conditions while mitigating data scarcity. Additionally, we introduce a scalable synthetic data generation pipeline that combines graphics-based ruler creation with ControlNet-enhanced realism, significantly expanding training diversity and improving model performance. Extensive experiments on both our datasets and public benchmarks demonstrate that RulerNet achieves accurate, consistent, and efficient scale estimation under challenging real-world conditions. Integration into a medical analysis pipeline further demonstrates its practical utility for scale-aware measurement. These results suggest that RulerNet can serve as a generalizable measurement component and be readily integrated with other vision modules, such as segmentation, monocular depth estimation, and diagnostic workflows, for automated analysis in medical and other domains. An online demo is provided. The project page is at https://github.com/ymp5078/RulerNet.

cs.CV

VLCD: Vision-Language Contrastive Distillation for Accurate and Efficient Automatic Placenta Analysis

Pathological examination of the placenta is an effective method for detecting and mitigating health risks associated with childbirth. Recent advancements in AI have enabled the use of photographs of the placenta and pathology reports for detecting and classifying signs of childbirth-related pathologies. However, existing automated methods are computationally extensive, which limits their deployability. We propose two modifications to vision-language contrastive learning (VLC) frameworks to enhance their accuracy and efficiency: (1) text-anchored vision-language contrastive knowledge distillation (VLCD)-a new knowledge distillation strategy for medical VLC pretraining, and (2) unsupervised predistillation using a large natural images dataset for improved initialization. Our approach distills efficient neural networks that match or surpass the teacher model in performance while achieving model compression and acceleration. Our results showcase the value of unsupervised predistillation in improving the performance and robustness of our approach, specifically for lower-quality images. VLCD serves as an effective way to improve the efficiency and deployability of medical VLC approaches, making AI-based healthcare solutions more accessible, especially in resource-constrained environments.

cs.CV

S2S2: Semantic Stacking for Robust Semantic Segmentation in Medical Imaging

Robustness and generalizability in medical image segmentation are often hindered by scarcity and limited diversity of training data, which stands in contrast to the variability encountered during inference. While conventional strategies -- such as domain-specific augmentation, specialized architectures, and tailored training procedures -- can alleviate these issues, they depend on the availability and reliability of domain knowledge. When such knowledge is unavailable, misleading, or improperly applied, performance may deteriorate. In response, we introduce a novel, domain-agnostic, add-on, and data-driven strategy inspired by image stacking in image denoising. Termed ``semantic stacking,'' our method estimates a denoised semantic representation that complements the conventional segmentation loss during training. This method does not depend on domain-specific assumptions, making it broadly applicable across diverse image modalities, model architectures, and augmentation techniques. Through extensive experiments, we validate the superiority of our approach in improving segmentation performance under diverse conditions. Code is available at https://github.com/ymp5078/Semantic-Stacking.

cs.CV

AI-SAM: Automatic and Interactive Segment Anything Model

Semantic segmentation is a core task in computer vision. Existing methods are generally divided into two categories: automatic and interactive. Interactive approaches, exemplified by the Segment Anything Model (SAM), have shown promise as pre-trained models. However, current adaptation strategies for these models tend to lean towards either automatic or interactive approaches. Interactive methods depend on prompts user input to operate, while automatic ones bypass the interactive promptability entirely. Addressing these limitations, we introduce a novel paradigm and its first model: the Automatic and Interactive Segment Anything Model (AI-SAM). In this paradigm, we conduct a comprehensive analysis of prompt quality and introduce the pioneering Automatic and Interactive Prompter (AI-Prompter) that automatically generates initial point prompts while accepting additional user inputs. Our experimental results demonstrate AI-SAM's effectiveness in the automatic setting, achieving state-of-the-art performance. Significantly, it offers the flexibility to incorporate additional user prompts, thereby further enhancing its performance. The project page is available at https://github.com/ymp5078/AI-SAM.

cs.CV