A scale-controlled super-resolution study with YOLOv11 on VinDr-Mammo
Medical Imaging 2026: Computer-Aided Diagnosis, Vancouver, Canada, 15 - 19 February 2026, vol.13926, (Full Text)
- Publication Type: Conference Paper / Full Text
- Volume: 13926
- Doi Number: 10.1117/12.3088089
- City: Vancouver
- Country: Canada
- Keywords: Mammography, Medical imaging, Object detection, Real-ESRGAN, Super-resolution, Task-driven evaluation, VinDrMammo, YOLOv11
- Hacettepe University Affiliated: No
Abstract
Deep learning-based super-resolution (SR) has recently emerged as a promising approach for enhancing medical images, offering the potential to improve image quality without the associated risks of increased radiation dose or extended acquisition time. However, its value for downstream tasks remains unclear in mammography, where clinically relevant targets like masses live at the limit of native sampling. We conduct a scale-controlled, taskoriented study on the VinDr-Mammo dataset by generating low-resolution variants (128, 256, 512, and 1024 pixels), training independent YOLOv11 detectors at each respective resolution, and subsequently evaluating super-resolution (2×/4×) to object detection SR → OD pipelines in comparison with bilinear up-sampling and native-resolution baselines. To reduce domain confounders, we restrict the dataset to Siemens Mammomat cases with mass annotations and form dual-view composites (CC+MLO) on a square canvas to expose cross-view context to the detector. Our results demonstrate that super-resolution (SR) consistently improves PSNR and SSIM relative to bilinear interpolation across all settings; however, the corresponding task-level gains are highly dependent on the native sampling resolution. Notably, at 128 px, SR fails to outperform the native detector. From 256 px, 4× SR boosts mAP from 0.493 (native) and 0.526 (bilinear) to 0.573. Similarly, from 512 px, 2× SR boosts mAP from 0.503 (native) and 0.574 (bilinear) to 0.592. These findings suggest that SR becomes advantageous once lesions subtend a minimal spatial footprint, and may even produce inputs that align more effectively with a detector's learned feature representations than those obtained through raw higher-resolution training.