Clustering Non-Ordered Discrete Data

  • Watve, Alok
  • Pramanik, Sakti
  • Jung, Sungwon
  • Jo, Bumjoon
  • Kumar, Sunil
  • 외 1명
Citations

WEB OF SCIENCE

5
Citations

SCOPUS

5

초록

Clustering in continuous vector data spaces is a well-studied problem. In recent years there has been a significant amount of research work in clustering categorical data. However, most of these works deal with market-basket type transaction data and are not specifically optimized for high-dimensional vectors. Our focus in this paper is to efficiently cluster high-dimensional vectors in non-ordered discrete data spaces (NDDS). We have defined several necessary geometrical concepts in NDDS which form the basis of our clustering algorithm. Several new heuristics have been employed exploiting the characteristics of vectors in NDDS. Experimental results on large synthetic datasets demonstrate that the proposed approach is effective, in terms of cluster quality, robustness and running time. We have also applied our clustering algorithm to real datasets with promising results.

키워드

clusteringdata miningcategorical datanon-ordered discrete datavector dataCATEGORICAL-DATAALGORITHM
제목
Clustering Non-Ordered Discrete Data
저자
Watve, AlokPramanik, SaktiJung, SungwonJo, BumjoonKumar, SunilSural, Shamik
발행일
2014-01
유형
Article
저널명
Journal of Information Science and Engineering
30
1
페이지
1 ~ 23