Searcharxiv⌕ Search

arXiv subjects

Glenda Reynolds

Publications and source records attributed to Glenda Reynolds.

2 recordsLinked to original sources

DentiAsk: A VQA Benchmark for Multimodal Reasoning in Panoramic Dental Radiographs

Accurate interpretation of panoramic dental radiographs requires the integration of multiple reasoning capabilities: detection, spatial localization, and quantitative assessment. Despite recent advances in multimodal learning, existing medical visual question answering (VQA) benchmarks do not fully capture this complexity, often reducing the task to simplified classification or templated queries. As a result, they provide limited coverage of the diverse reasoning processes required for clinically meaningful interpretation. We introduce DentiAsk, a large-scale dental VQA benchmark that pairs high-resolution panoramic dental radiographs with clinician-validated question-answer pairs spanning three reasoning tiers: descriptive recognition, spatial localization, and numerical quantification across three high-prevalence pathologies: periapical radiolucency (PARL), impacted teeth, and dental caries. DentiAsk comprises 1,000 high-resolution radiographs annotated with 10,000 expert-curated QA pairs. To our knowledge, it is the first dental VQA benchmark to unify categorical, spatial, and quantitative reasoning as separately scored tasks within a single evaluation framework. We benchmark 10 state-of-the-art vision-language models, including LLaVA-v1.5, LLaVA-v1.6, Qwen-VL, InternVL2, and LLaVA-Med, and find that models achieve stronger performance on descriptive queries, whereas they degrade sharply on spatial localization and counting, exposing limitations in compositional, multi-step reasoning. These findings reveal a gap between visual recognition and clinically meaningful reasoning, establishing DentiAsk as a challenging benchmark for advancing multimodal reasoning in medical imaging.

q-bio.QM↗

When CNNs Outperform Transformers and Mambas: Revisiting Deep Architectures for Dental Caries Segmentation

Accurate identification and segmentation of dental caries in panoramic radiographs are critical for early diagnosis and effective treatment planning. Automated segmentation remains challenging due to low lesion contrast, morphological variability, and limited annotated data. In this study, we present the first comprehensive benchmarking of convolutional neural networks, vision transformers and state-space mamba architectures for automated dental caries segmentation on panoramic radiographs through a DC1000 dataset. Twelve state-of-the-art architectures, including VMUnet, MambaUNet, VMUNetv2, RMAMamba-S, TransNetR, PVTFormer, DoubleU-Net, and ResUNet++, were trained under identical configurations. Results reveal that, contrary to the growing trend toward complex attention based architectures, the CNN-based DoubleU-Net achieved the highest dice coefficient of 0.7345, mIoU of 0.5978, and precision of 0.8145, outperforming all transformer and Mamba variants. In the study, the top 3 results across all performance metrics were achieved by CNN-based architectures. Here, Mamba and transformer-based methods, despite their theoretical advantage in global context modeling, underperformed due to limited data and weaker spatial priors. These findings underscore the importance of architecture-task alignment in domain-specific medical image segmentation more than model complexity. Our code is available at: https://github.com/JunZengz/dental-caries-segmentation.

cs.CV↗