NeuroFuser: Resource-aware neuromodulation of multi-scale fusion attention for domain adaptive segmentation


Creative Commons License

Erişen S.

PATTERN RECOGNITION, vol.180, no.114167, pp.1-12, 2026 (SCI-Expanded, Scopus)

  • Publication Type: Article / Article
  • Volume: 180 Issue: 114167
  • Publication Date: 2026
  • Doi Number: 10.1016/j.patcog.2026.114167
  • Journal Name: PATTERN RECOGNITION
  • Journal Indexes: Applied Science & Technology Source, Academic Search Ultimate (EBSCO), Engineering Source (EBSCO), Scopus, Science Citation Index Expanded (SCI-EXPANDED), BIOSIS, Compendex, INSPEC, MLA - Modern Language Association Database, zbMATH, MLA International Bibliography
  • Page Numbers: pp.1-12
  • Open Archive Collection: AVESIS Open Access Collection
  • Hacettepe University Affiliated: Yes

Abstract

Balancing multi-scale fusion and spatial precision under fixed resources remains a challenging bottleneck in semantic segmentation. Strong attention-based mechanisms become costly at high resolution while lightweight designs often lack precision and degrade under cross-domain transfer. To address this gap, NeuroFuser is proposed as a memory-efficient encoder–decoder that treats fusion as a budgeted, controllable process rather than a fixed architectural choice. NeuroFuser introduces plug-in attention-fusion modules and a neuromodulation controller that explicitly regulates multi-head, multi-scale fusion, windowing, and mixture-of-dilations to enforce predictable memory while preserving fine details by stable scaling across resolutions and improves robustness to domain shift without backbone-specific redesign. For training-free domain adaptation, NeuroFuser aligns label palettes using CLIP-ViT class prototypes and reuses the same controller on the target domain, keeping computational cost fixed while improving cross-domain performance. NeuroFuser achieves 86.4% mIoU on Cityscapes and 64.9% mIoU on ADE20K, and transfers to BDD100K with 74.2% mIoU, CamVid with 86.6% mIoU, NYU-D with 57.7% mIoU, and SUN RGB-D with 53.5% mIoU. With neuromodulation, peak GPU memory drops from 77.5 GB to 63.4 GB, enabling high-resolution training and inference on a single A100-80 GB GPU. NeuroFuser integrates with CNN and transformer backbones, making it suitable for multi-modal pipelines where segmentation acts as an anchor task. The results show consistent accuracy with tight resource control, establishing neuromodulated, memory-efficient multi-scale fusion as a state-of-the-art for domain-adaptive dense prediction.