Disk-Based Shared KV Cache Management for Fast Inference in Multi-Instance LLM RAG Systems

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

Recent large language models (LLMs) face increasing inference latency as input context length and model size grow. Retrieval-augmented generation (RAG) exacerbates this by significantly increasing input tokens, leading to higher computational overhead during the prefill stage and prolonged time-to-first-token (TTFT). To address this, the paper proposes using a disk-based key-value (KV) cache to reduce the prefill computational burden, thereby shortening TTFT. We also introduce a disk-based shared KV cache management system, called Shared RAG-DCache, for multi-instance LLM RAG service environments. This system leverages the locality of documents related to user queries in RAG and queueing delays in LLM inference services to proactively generate and store disk KV caches for query-related documents, sharing them across multiple LLM instances to enhance inference performance. In experiments on a single host with 2 GPUs and 1 CPU, Shared RAG-DCache achieved a 15-71% increase in throughput and up to a 12-65% reduction in latency, depending on the resource configuration.

키워드

LLMKV-CacheRAGVector DBTTFT
제목
Disk-Based Shared KV Cache Management for Fast Inference in Multi-Instance LLM RAG Systems
저자
Lee, Hyung WooKim, Ki HyunKim, Jin WooSo, Jung MinCha, Myung HoonKim, Hong YeonKim, James J.Kim, Young Jae
DOI
10.1109/CLOUD67622.2025.00029
발행일
2025
유형
Proceedings Paper
저널명
IEEE International Conference on Cloud Computing, CLOUD
페이지
199 ~ 209