Mapreduce-Based Distributed Clustering Method Using CF<sup> plus </sup> Tree

Citations

WEB OF SCIENCE

3
Citations

SCOPUS

3

초록

Clustering exceptionally large data sets is becoming a major challenge in data analytics with the continuous increase in their size. Summary-based clustering methods and distributed computing frameworks such as MapReduce can efficiently handle this challenge. These methods include BIRCH and its extension CF+-ERC. CF+-ERC can reduce the clustering time of large data sets by utilizing the structure of a CF+ tree. However, CF+-ERC is a sequential clustering method, so it cannot be used with multiple machines to reduce the clustering time. In this study, we propose a novel MapReduce-based distributed clustering method called CF+-ERC on MapReduce (CF+ERC_MR). It builds a CF+ tree for clustering an exceptionally large data set with a given threshold and finds the final clusters using MapReduce, which significantly reduces the clustering time. Further, our method is scalable with respect to the number of machines. The efficacy of this method is validated through not only its theoretical analysis but also in-depth experimental analysis of exceptionally large synthetic and real data sets. The experimental results demonstrate that the clustering speed of our approach is far superior to that of the existing clustering methods.

키워드

ClusteringBIRCHCF treerange queryvery large data setsMapReduceALGORITHMS
제목
Mapreduce-Based Distributed Clustering Method Using CF<sup> plus </sup> Tree
저자
Ryu, Hyeong-CheolJung, Sungwon
DOI
10.1109/ACCESS.2020.2999085
발행일
2020
유형
Article
저널명
IEEE Access
8
페이지
104232 ~ 104246