Integrating Distributed SQL Query Engines with Object-Based Computational Storage

  • Ryu, Junghyun
  • Hwang, Soon
  • Park, Junhyeok
  • Ahn, Seonghoon
  • Park, JeoungAhn
  • ... Kim, Youngjae
  • 외 7명
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Existing object storage systems like AWS S3 and MinIO offer only limited in-storage compute capabilities, typically restricted to simple SQL WHERE-clause filtering. Consequently, high-impact operators such as aggregation and top-N are still executed entirely at the compute layer. Recent advances in Object-based Computational Storage (OCS) enable these complex operators to run natively within storage, creating opportunities for substantial reductions in data movement and query time. To demonstrate these benefits in distributed SQL engines, we used Presto as a case study and developed the Presto-OCS connector, which analyzes execution plans to identify pushdown-eligible operators and offloads them to OCS for efficient in-storage execution. Evaluations with real-world HPC analytics queries and the TPC-H benchmark show that our approach achieves up to 4.07x speedup and 99% data movement reduction compared to filter-only pushdown. When combined with compression techniques, our approach delivers 1.39x speedup over compressed filter-only pushdown, demonstrating that advanced query pushdown complements existing optimizations.

키워드

Computational StorageObject StorageSQL Query EnginesBig Data AnalyticsCOMPRESSION
제목
Integrating Distributed SQL Query Engines with Object-Based Computational Storage
저자
Ryu, JunghyunHwang, SoonPark, JunhyeokAhn, SeonghoonPark, JeoungAhnLee, JeongjinYang, JinnaYang, SoonyealNoh, JungkiZheng, QingChung, WoosukKim, HoshikKim, Youngjae
DOI
10.1145/3731599.3767371
발행일
2025-11
유형
Proceedings Paper
저널명
PROCEEDINGS OF 2025 WORKSHOPS OF THE INTERNATIONAL CONFERENCE ON HIGH PERFORMANCE COMPUTING, NETWORK, STORAGE, AND ANALYSIS, SC25 WORKSHOPS
페이지
290 ~ 299