Accelerated Block-Sparsity-Aware Matrix Reordering for Leveraging Tensor Cores in Sparse Matrix-Multivector Multiplication

Citations

WEB OF SCIENCE

4
Citations

SCOPUS

4

초록

Sparse Matrix-Multivector (SpMM) multiplication is a key kernel in deep learning models and scientific computing applications. However, achieving high performance for SpMM is challenging due to the irregular distribution of non-zero elements and memory access patterns. Therefore, several sparse matrix reordering algorithms have been developed to improve data locality for SpMM. However, existing approaches for reordering sparse matrix have not considered block sparsity during the reordering process. In this paper, we present a novel algorithm for sparse matrix reordering that considers block sparsity to enhance data locality for SpMM on Tensor Cores. To alleviate the main bottleneck of reordering, which involves substantial computations for measuring similarity between rows, we develop an efficient GPU implementation by adapting dynamic parallelism and synchronization schemes. Experimental results on a large number of sparse matrices demonstrate the effectiveness of our reordering algorithm and the benefits of leveraging Tensor Cores for SpMM. Our approach achieves a significant performance improvement over various state-of-the-art SpMM implementations.

키워드

SpMMsparse matrix reorderingTensor Cores
제목
Accelerated Block-Sparsity-Aware Matrix Reordering for Leveraging Tensor Cores in Sparse Matrix-Multivector Multiplication
저자
Lee, EunjiHan, YoonsangMoon, Gordon Euhyun
DOI
10.1007/978-3-031-69583-4_1
발행일
2024
유형
Proceedings Paper
저널명
Lecture Notes in Computer Science
14803
페이지
3 ~ 16