상세 보기
FlexLLM: Flexible and Cost-Efficient LLM Serving with Heterogeneous GPUs
- Kim, Kihyun;
- Kim, Jinwoo;
- Kim, James J.;
- Li, Dong;
- Kim, Youngjae
SCOPUS
0초록
The autoregressive nature of LLMs causes memory bottlenecks, requiring multi-GPU parallelization. Prior work has mainly optimized strategies for homogeneous setups, overlooking heterogeneous configurations and service-level objectives (SLOs) despite growing GPU cost-performance gaps. This paper introduces FlexLLM, a framework that predicts execution times and selects cost-efficient strategies in heterogeneous GPU environments while satisfying latency-per-token (LPT) SLO constraints. The proposed system (i) bridges theoretical predictions and actual performance through a Linear Correction Function (LCF) and (ii) performs SLO-aware cost-efficiency optimization based on human reading speeds (≤150 ms per token). Our evaluation demonstrates that FlexLLM identifies costefficient configurations meeting SLO requirements, achieving significant cost reductions compared to performance-oriented approaches. Heterogeneous GPU analysis reveals that SLO-aware parallelization strategy selection yields up to 2.28 × cost-efficiency differences, demonstrating that architecture selection under SLO constraints is as critical as hardware investment. © 2025 IEEE.
키워드
- 제목
- FlexLLM: Flexible and Cost-Efficient LLM Serving with Heterogeneous GPUs
- 저자
- Kim, Kihyun; Kim, Jinwoo; Kim, James J.; Li, Dong; Kim, Youngjae
- 발행일
- 2025-10
- 유형
- Conference paper
- 저널명
- Proceedings - IEEE Computer Society's Annual International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunications Systems, MASCOTS