상세 보기
Fault-Tolerant Deep Learning Cache with Hash Ring for Load Balancing in HPC Systems
- Lee, Seoyeong;
- Khan, Awais;
- Kim, Yoochan;
- Park, Junghwan;
- Hwang, Soon;
- ... Kim, Youngjae;
- 외 3명
SCOPUS
0초록
Large-scale DL on HPC systems like Frontier and Summit uses distributed node-local caching to address scalability and performance challenges. However, as these systems grow more complex, the risk of node failures increases, and current caching approaches lack fault tolerance, jeopardizing large-scale training jobs. We analyzed six months of SLURM job logs from Frontier and found that over 30% of jobs failed after an average of 75 minutes. To address this, we propose fault-tolerance strategies that recache data lost from failed nodes using a hash ring technique for balanced data recaching in the distributed node-local caching, reducing reliance on the PFS. Our extensive evaluations on Frontier showed that the hash ring-based recaching approach reduced training time by approximately 25% compared to the approach that redirects I/O to the PFS after node failures and demonstrated effective load balancing of training data across nodes. © 2024 IEEE.
키워드
- 제목
- Fault-Tolerant Deep Learning Cache with Hash Ring for Load Balancing in HPC Systems
- 저자
- Lee, Seoyeong; Khan, Awais; Kim, Yoochan; Park, Junghwan; Hwang, Soon; Lee, Jae-Kook; Hong, Taeyoung; Zimmer, Christopher; Kim, Youngjae
- 발행일
- 2024
- 유형
- Conference Paper
- 저널명
- Proceedings of SC 2024-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis
- 페이지
- 1349 ~ 1357