A scale-controlled super-resolution study with YOLOv11 on VinDr-Mammo
Medical Imaging 2026: Computer-Aided Diagnosis, Vancouver, Kanada, 15 - 19 Şubat 2026, cilt.13926, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Cilt numarası: 13926
- Doi Numarası: 10.1117/12.3088089
- Basıldığı Şehir: Vancouver
- Basıldığı Ülke: Kanada
- Anahtar Kelimeler: Mammography, Medical imaging, Object detection, Real-ESRGAN, Super-resolution, Task-driven evaluation, VinDrMammo, YOLOv11
- Hacettepe Üniversitesi Adresli: Hayır
Özet
Deep learning-based super-resolution (SR) has recently emerged as a promising approach for enhancing medical images, offering the potential to improve image quality without the associated risks of increased radiation dose or extended acquisition time. However, its value for downstream tasks remains unclear in mammography, where clinically relevant targets like masses live at the limit of native sampling. We conduct a scale-controlled, taskoriented study on the VinDr-Mammo dataset by generating low-resolution variants (128, 256, 512, and 1024 pixels), training independent YOLOv11 detectors at each respective resolution, and subsequently evaluating super-resolution (2×/4×) to object detection SR → OD pipelines in comparison with bilinear up-sampling and native-resolution baselines. To reduce domain confounders, we restrict the dataset to Siemens Mammomat cases with mass annotations and form dual-view composites (CC+MLO) on a square canvas to expose cross-view context to the detector. Our results demonstrate that super-resolution (SR) consistently improves PSNR and SSIM relative to bilinear interpolation across all settings; however, the corresponding task-level gains are highly dependent on the native sampling resolution. Notably, at 128 px, SR fails to outperform the native detector. From 256 px, 4× SR boosts mAP from 0.493 (native) and 0.526 (bilinear) to 0.573. Similarly, from 512 px, 2× SR boosts mAP from 0.503 (native) and 0.574 (bilinear) to 0.592. These findings suggest that SR becomes advantageous once lesions subtend a minimal spatial footprint, and may even produce inputs that align more effectively with a detector's learned feature representations than those obtained through raw higher-resolution training.