CCASeg: Decoding Multi-Scale Context with Convolutional Cross-Attention for Semantic Segmentation

Citations

WEB OF SCIENCE

3
Citations

SCOPUS

4

초록

Capturing multi-scale context within feature maps is crucial for semantic segmentation. With the success of the Vision Transformer (ViT), recent models have been designed with transformer decoders to capture it. However, these models face limitations in utilizing diverse contextual information due to the inherent nature of the attention mechanism and structural constraints. Typically, multi-head attention which leads to similar receptive fields for each token feature is achieved at the expense of significantly increased computational cost. The nature of the structure can cause inconsistent combination of the information across different levels. To address this issue, in this paper, we propose a novel and effective decoding scheme, CCASeg, which is based on convolutional cross-attention (CCA). The proposed CCA along with the decoding structure is devised not only to capture both local and global context through convolutional kernels of various sizes, but also to achieve high efficiency by effective utilization of the cheap convolution operations. Moreover, the decoding structure, which ensures the successive combination of information across various levels, facilitates understanding of diverse contexts. Consequently, this novel decoding scheme enables feature maps to effectively learn the relationships between objects of different sizes. In this way, our proposed CCASeg outperforms previous state-of-the-art methods on popular semantic segmentation benchmarks, including ADE20K, Cityscapes, COCO-stuff, and iSAID.

제목
CCASeg: Decoding Multi-Scale Context with Convolutional Cross-Attention for Semantic Segmentation
저자
Yoo, JiwonKo, DamiKim, Gyeong hwan
DOI
10.1109/WACV61041.2025.00918
발행일
2025
유형
Proceedings Paper
저널명
IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
페이지
9479 ~ 9488