LMLT: Low-to-high Multi-Level Vision Transformer for Lightweight Image Super- Resolution

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

2

초록

Recent Vision Transformer (ViT)-based methods for Image Super-Resolution have demonstrated impressive performance. However, they suffer from significant complexity, resulting in high inference times and memory usage. Additionally, ViT models using Window Self-Attention (WSA) face challenges in processing regions outside their windows. To address these issues, we propose the Low-to-high Multi-Level Transformer (LMLT), which employs attention with varying feature sizes for each head. LMLT divides image features along the channel dimension, gradually reduces spatial size for lower heads, and applies self-attention to each head. This approach effectively captures both local and global information. By integrating the results from lower heads into higher heads, LMLT overcomes the window boundary issues in self-attention. Extensive experiments show that our model significantly reduces inference time and GPU memory usage while maintaining or even surpassing the performance of state-of-the-art ViT-based Image Super-Resolution methods. © 2025 IEEE.

키워드

image super-resolutionvision transformer
제목
LMLT: Low-to-high Multi-Level Vision Transformer for Lightweight Image Super- Resolution
저자
Kim, JeongsooNang, JonghoChoe, Junsuk
DOI
10.1109/ICCVW69036.2025.00583
발행일
2025-10
유형
Proceedings Paper
저널명
Proceedings - 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
페이지
5568 ~ 5578