CNN-Based Encoder and Transformer-Based Decoder for Efficient Semantic Segmentation

Citations

SCOPUS

0

초록

Recent exploration of transformer-based networks have achieved significant success in the field of computer vision, particularly in dense prediction tasks. The potent attention mechanism necessitates quadratic computational complexity due to its consideration of interrelationships among all elements, thus constituting a significant limitation. Accordingly, there have been numerous studies focused on light-weighting transformer-based models. The key is to minimize computational cost while mitigating performance degradation as much as possible. In this paper, we propose LGNet, an efficient yet powerful semantic segmentation network that encodes contextual information through convolutional attention, while enforcing and aggregating global context of hierarchical features via neighborhood attention. The proposed model demonstrates significantly lower computational overhead compared to previous methods while achieving minimal performance degradation. © 2024 IEEE.

키워드

attentionCNNsemantic segmentationtransformer
제목
CNN-Based Encoder and Transformer-Based Decoder for Efficient Semantic Segmentation
저자
Moon, SeunghunKang, Suk-ju
DOI
10.1109/ICEIC61013.2024.10457284
발행일
2024
유형
Conference Paper
저널명
2024 International Conference on Electronics, Information, and Communication, ICEIC 2024