FlexLLM: Flexible and Cost-Efficient LLM Serving with Heterogeneous GPUs

  • Kim, Kihyun; 
  • Kim, Jinwoo; 
  • Kim, James J.; 
  • Li, Dong; 
  • Kim, Youngjae
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

The autoregressive nature of LLMs causes memory bottlenecks, requiring multi-GPU parallelization. Prior work has mainly optimized strategies for homogeneous setups, overlooking heterogeneous configurations and service-level objectives (SLOs) despite growing GPU cost-performance gaps. This paper introduces FlexLLM, a framework that predicts execution times and selects cost-efficient strategies in heterogeneous GPU environments while satisfying latency-per-token (LPT) SLO constraints. The proposed system (i) bridges theoretical predictions and actual performance through a Linear Correction Function (LCF) and (ii) performs SLO-aware cost-efficiency optimization based on human reading speeds (≤150 ms per token). Our evaluation demonstrates that FlexLLM identifies costefficient configurations meeting SLO requirements, achieving significant cost reductions compared to performance-oriented approaches. Heterogeneous GPU analysis reveals that SLO-aware parallelization strategy selection yields up to 2.28 × cost-efficiency differences, demonstrating that architecture selection under SLO constraints is as critical as hardware investment. © 2025 IEEE.

키워드

Cost-Efficiency Optimization; Heterogeneous GPU Computing; Large Language Model Serving; Parallelization Strategies; Performance Prediction
제목
FlexLLM: Flexible and Cost-Efficient LLM Serving with Heterogeneous GPUs
저자
Kim, Kihyun; Kim, Jinwoo; Kim, James J.; Li, Dong; Kim, Youngjae
DOI
10.1109/MASCOTS67699.2025.11283352
발행일
2025-10
유형
Proceedings Paper
저널명
Proceedings - IEEE Computer Society's Annual International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunications Systems, MASCOTS
페이지
74 ~ 81