OncoVision: Integrating Mammography and Clinical Data through Attention-Driven Multimodal AI for Enhanced Breast Cancer Diagnosis
OncoVision is a privileged-information training framework that uses mammography images and clinical features during training and performs inference from mammographic images alone. Employing an attention-based encoder-decoder backbone, it jointly segments four regions of interest (masses, calcifications, axillary findings, and breast tissue) with accuracy exceeding the nnU-Net baseline and predicts ten structured clinical features, including BI-RADS category. We developed two late-fusion strategies, Independent and Dependent, that integrate imaging, radiomic, and clinical information during training to improve diagnostic precision and potentially reduce inter-observer variability. Radiomic features extracted from predicted masks provide shape, intensity, and texture descriptors that complement the learned CNN representations. We evaluated OncoVision in a retrospective multi-reader study with six board-certified radiologists, assessing diagnostic confidence, reading time, and segmentation accuracy with and without AI assistance. In a paired reader-assistance evaluation, OncoVision was associated with higher diagnostic confidence for junior and senior radiologists, reduced reading time by up to 61%, and achieved segmentation accuracy comparable to or exceeding that of radiologists for mass lesions. We operationalized OncoVision as a secure web application, now deployed at a partner hospital, that generates structured reports with dual-confidence scoring and attention-weighted visualizations for real-time diagnostic support. The platform is designed for integration into clinical workflows, with the goal of supporting screening access in underprivileged regions. By combining accurate segmentation with clinical intuition, OncoVision advances AI-assisted mammographic interpretation, offering a scalable and accessible approach to earlier and more consistent image interpretation.