상세 보기
CALL: Context-Aware Low-Latency Retrieval in Disk-Based Vector Databases
- Jeong, Yeonwoo;
- Cho, Hyunji;
- Park, Kyuli;
- Kim, Youngjae;
- Park, Sungyong
WEB OF SCIENCE
0SCOPUS
0초록
Embedding models capture both semantic and syntactic structures of queries, often mapping different queries to similar regions in vector space. This results in nonuniform cluster access patterns in modern disk-based vector databases. While existing approaches optimize individual queries, they overlook the impact of cluster access patterns, failing to account for the locality effects of queries that access similar clusters. This oversight increases cache miss penalty. To minimize the cache miss penalty, we propose CALL, a context-Aware query grouping mechanism that organizes queries based on shared cluster access patterns. Additionally, CALL incorporates a group-Aware prefetching method to minimize cache misses during transitions between query groups and latency-Aware cluster loading. Experimental results show that CALL reduces the 99th percentile tail latency by up to 33 % while consistently maintaining a higher cache hit ratio, substantially reducing search latency. © 2025 IEEE.
키워드
- 제목
- CALL: Context-Aware Low-Latency Retrieval in Disk-Based Vector Databases
- 저자
- Jeong, Yeonwoo; Cho, Hyunji; Park, Kyuli; Kim, Youngjae; Park, Sungyong
- 발행일
- 2025-12
- 유형
- Proceedings Paper
- 저널명
- Proceedings - 2025 IEEE 32nd International Conference on High Performance Computing, Data, and Analytics, HiPC 2025
- 페이지
- 172 ~ 182