Exploiting Tensor Cores in Sparse Matrix-Multivector Multiplication via Block-Sparsity-Aware Clustering

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

2

초록

Sparse Matrix-Multivector (SpMM) multiplication is a key kernel for deep learning models and scientific computing applications. However, achieving high performance for SpMM on GPUs is challenging because of the irregular distribution of non-zero elements and irregular memory accesses in sparse matrices. In this paper, we propose a novel sparse matrix reordering algorithm based on the block sparsity patterns in rows to improve data locality for SpMM. To accelerate our reordering algorithm, we develop a parallel implementation based on GPUs. The high-density tiles in the reordered matrix are used to perform accelerated dense matrix multiplication by leveraging Tensor Cores on GPUs. Experimental results on a large number of sparse matrices demonstrate that our TC-SpMM achieves an average speedup of 3.4 x and a peak speedup of 20.77 x over cuSPARSE.

키워드

SpMMsparse matrix reorderingTensor Cores
제목
Exploiting Tensor Cores in Sparse Matrix-Multivector Multiplication via Block-Sparsity-Aware Clustering
저자
Lee, EunjiHan, YoonsangMoon, Gordon Euhyun
DOI
10.1109/IPDPSW63119.2024.00199
발행일
2024
유형
Proceedings Paper
저널명
2024 IEEE INTERNATIONAL PARALLEL AND DISTRIBUTED PROCESSING SYMPOSIUM WORKSHOPS, IPDPSW 2024
페이지
1181 ~ 1183