상세 보기
FlexLLM: Flexible and Cost-Efficient LLM Serving with Heterogeneous GPUs
- Kim, Kihyun;
- Kim, Jinwoo;
- Kim, James J.;
- Li, Dong;
- Kim, Youngjae
WEB OF SCIENCE
0SCOPUS
0초록
The autoregressive nature of LLMs causes memory bottlenecks, requiring multi-GPU parallelization. Prior work has mainly optimized strategies for homogeneous setups, overlooking heterogeneous configurations and service-level objectives (SLOs) despite growing GPU cost-performance gaps. This paper introduces FlexLLM, a framework that predicts execution times and selects cost-efficient strategies in heterogeneous GPU environments while satisfying latency-per-token (LPT) SLO constraints. The proposed system (i) bridges theoretical predictions and actual performance through a Linear Correction Function (LCF) and (ii) performs SLO-aware cost-efficiency optimization based on human reading speeds (≤150 ms per token). Our evaluation demonstrates that FlexLLM identifies costefficient configurations meeting SLO requirements, achieving significant cost reductions compared to performance-oriented approaches. Heterogeneous GPU analysis reveals that SLO-aware parallelization strategy selection yields up to 2.28 × cost-efficiency differences, demonstrating that architecture selection under SLO constraints is as critical as hardware investment. © 2025 IEEE.
키워드
- 제목
- FlexLLM: Flexible and Cost-Efficient LLM Serving with Heterogeneous GPUs
- 저자
- Kim, Kihyun; Kim, Jinwoo; Kim, James J.; Li, Dong; Kim, Youngjae
- 발행일
- 2025-10
- 유형
- Proceedings Paper
- 페이지
- 74 ~ 81
- 언어
- ENG
- 출판사
- IEEE Computer Society
- 분량
- 8 페이지
- ISSN
- P 1526-7539