E-Flash: Energy-Efficient LLM Mapping on NAND Flash-Based In-Storage Inference Computing

  • Ji, Gisan
  • Shin, Sanghun
  • Baik, Jangho
  • Shim, Wonbo
  • Ryu, Sungju
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Transformer-based deep neural networks (DNNs) have achieved remarkable success across a wide range of applications such as image and text generation tasks. However, The continuous growth of model size and memory demands imposes significant pressure on energy consumption and memory bandwidth, especially during weight access operations. To deal with such a challenge, prior studies have investigated NAND flash-based processing-in-memory (PIM) architectures, but it still experiences large energy consumption due to the significant increase in the recent model sizes. In this work, we present E-Flash, a digital NAND flash-based architecture for energy-efficient DNN weight access. E-Flash introduces a novel state-switching algorithm that reallocates frequently occurring weight patterns to low-power cell states in triple-level cell (TLC) flash memory. In addition, a cell-first allocation scheme further amplifies energy savings by aligning bit patterns within cells. Evaluation results on quantized BERT and Llama 2 models demonstrate up to 37.73% and 16.74% reduction in read energy, respectively, with negligible hardware overhead.

키워드

Computer architectureMicroprocessorsFlash memoriesEnergy efficiencyReflective binary codesArtificial intelligenceResource managementTransformersHardwareOptimizationNAND flashquadraple-level cell (QLC)state-switching algorithmtransformerstriple-level cell (TLC)
제목
E-Flash: Energy-Efficient LLM Mapping on NAND Flash-Based In-Storage Inference Computing
저자
Ji, GisanShin, SanghunBaik, JanghoShim, WonboRyu, Sungju
DOI
10.1109/TVLSI.2026.3657777
발행일
2026-04
유형
Article
저널명
IEEE Transactions on Very Large Scale Integration (VLSI) Systems
34
4
페이지
1124 ~ 1133