상세 보기
Disk-Based Shared KV Cache Management for Fast Inference in Multi-Instance LLM RAG Systems
- Lee, Hyung Woo;
- Kim, Ki Hyun;
- Kim, Jin Woo;
- So, Jung Min;
- Cha, Myung Hoon;
- ... Kim, Young Jae;
- 외 2명
WEB OF SCIENCE
1SCOPUS
1초록
Recent large language models (LLMs) face increasing inference latency as input context length and model size grow. Retrieval-augmented generation (RAG) exacerbates this by significantly increasing input tokens, leading to higher computational overhead during the prefill stage and prolonged time-to-first-token (TTFT). To address this, the paper proposes using a disk-based key-value (KV) cache to reduce the prefill computational burden, thereby shortening TTFT. We also introduce a disk-based shared KV cache management system, called Shared RAG-DCache, for multi-instance LLM RAG service environments. This system leverages the locality of documents related to user queries in RAG and queueing delays in LLM inference services to proactively generate and store disk KV caches for query-related documents, sharing them across multiple LLM instances to enhance inference performance. In experiments on a single host with 2 GPUs and 1 CPU, Shared RAG-DCache achieved a 15-71% increase in throughput and up to a 12-65% reduction in latency, depending on the resource configuration.
키워드
- 제목
- Disk-Based Shared KV Cache Management for Fast Inference in Multi-Instance LLM RAG Systems
- 저자
- Lee, Hyung Woo; Kim, Ki Hyun; Kim, Jin Woo; So, Jung Min; Cha, Myung Hoon; Kim, Hong Yeon; Kim, James J.; Kim, Young Jae
- 발행일
- 2025
- 유형
- Proceedings Paper
- 저널명
- IEEE International Conference on Cloud Computing, CLOUD
- 페이지
- 199 ~ 209