NeuroFuser: Resource-aware neuromodulation of multi-scale fusion attention for domain adaptive segmentation
PATTERN RECOGNITION, vol.180, no.114167, pp.1-12, 2026 (SCI-Expanded, Scopus)
- Publication Type: Article / Article
- Volume: 180 Issue: 114167
- Publication Date: 2026
- Doi Number: 10.1016/j.patcog.2026.114167
- Journal Name: PATTERN RECOGNITION
- Journal Indexes: Applied Science & Technology Source, Academic Search Ultimate (EBSCO), Engineering Source (EBSCO), Scopus, Science Citation Index Expanded (SCI-EXPANDED), BIOSIS, Compendex, INSPEC, MLA - Modern Language Association Database, zbMATH, MLA International Bibliography
- Page Numbers: pp.1-12
- Open Archive Collection: AVESIS Open Access Collection
- Hacettepe University Affiliated: Yes
Abstract
Balancing multi-scale fusion and spatial precision under fixed resources remains a challenging bottleneck in semantic segmentation. Strong attention-based mechanisms become costly at high resolution while lightweight designs often lack precision and degrade under cross-domain transfer. To address this gap, NeuroFuser is proposed as a memory-efficient encoder–decoder that treats fusion as a budgeted, controllable process rather than a fixed architectural choice. NeuroFuser introduces plug-in attention-fusion modules and a neuromodulation controller that explicitly regulates multi-head, multi-scale fusion, windowing, and mixture-of-dilations to enforce predictable memory while preserving fine details by stable scaling across resolutions and improves robustness to domain shift without backbone-specific redesign. For training-free domain adaptation, NeuroFuser aligns label palettes using CLIP-ViT class prototypes and reuses the same controller on the target domain, keeping computational cost fixed while improving cross-domain performance. NeuroFuser achieves 86.4% mIoU on Cityscapes and 64.9% mIoU on ADE20K, and transfers to BDD100K with 74.2% mIoU, CamVid with 86.6% mIoU, NYU-D with 57.7% mIoU, and SUN RGB-D with 53.5% mIoU. With neuromodulation, peak GPU memory drops from 77.5 GB to 63.4 GB, enabling high-resolution training and inference on a single A100-80 GB GPU. NeuroFuser integrates with CNN and transformer backbones, making it suitable for multi-modal pipelines where segmentation acts as an anchor task. The results show consistent accuracy with tight resource control, establishing neuromodulated, memory-efficient multi-scale fusion as a state-of-the-art for domain-adaptive dense prediction.