NeuroFuser: Resource-aware neuromodulation of multi-scale fusion attention for domain adaptive segmentation
PATTERN RECOGNITION, cilt.180, sa.114167, ss.1-12, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 180 Sayı: 114167
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.patcog.2026.114167
- Dergi Adı: PATTERN RECOGNITION
- Derginin Tarandığı İndeksler: Applied Science & Technology Source, Academic Search Ultimate (EBSCO), Engineering Source (EBSCO), Scopus, Science Citation Index Expanded (SCI-EXPANDED), BIOSIS, Compendex, INSPEC, MLA - Modern Language Association Database, zbMATH, MLA International Bibliography
- Sayfa Sayıları: ss.1-12
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- Hacettepe Üniversitesi Adresli: Evet
Özet
Balancing multi-scale fusion and spatial precision under fixed resources remains a challenging bottleneck in semantic segmentation. Strong attention-based mechanisms become costly at high resolution while lightweight designs often lack precision and degrade under cross-domain transfer. To address this gap, NeuroFuser is proposed as a memory-efficient encoder–decoder that treats fusion as a budgeted, controllable process rather than a fixed architectural choice. NeuroFuser introduces plug-in attention-fusion modules and a neuromodulation controller that explicitly regulates multi-head, multi-scale fusion, windowing, and mixture-of-dilations to enforce predictable memory while preserving fine details by stable scaling across resolutions and improves robustness to domain shift without backbone-specific redesign. For training-free domain adaptation, NeuroFuser aligns label palettes using CLIP-ViT class prototypes and reuses the same controller on the target domain, keeping computational cost fixed while improving cross-domain performance. NeuroFuser achieves 86.4% mIoU on Cityscapes and 64.9% mIoU on ADE20K, and transfers to BDD100K with 74.2% mIoU, CamVid with 86.6% mIoU, NYU-D with 57.7% mIoU, and SUN RGB-D with 53.5% mIoU. With neuromodulation, peak GPU memory drops from 77.5 GB to 63.4 GB, enabling high-resolution training and inference on a single A100-80 GB GPU. NeuroFuser integrates with CNN and transformer backbones, making it suitable for multi-modal pipelines where segmentation acts as an anchor task. The results show consistent accuracy with tight resource control, establishing neuromodulated, memory-efficient multi-scale fusion as a state-of-the-art for domain-adaptive dense prediction.