FGCH: a fast and grid based clustering algorithm for hybrid data stream

被引:25
作者
Chen, Jinyin [1 ]
Lin, Xiang [1 ]
Xuan, Qi [1 ]
Xiang, Yun [1 ]
机构
[1] Zhejiang Univ Technol, Coll Informat Engn, Hangzhou, Zhejiang, Peoples R China
关键词
Data stream; Clustering analysis; Non-uniform attenuation; Grid clustering;
D O I
10.1007/s10489-018-1324-x
中图分类号
TP18 [人工智能理论];
学科分类号
140502 [人工智能];
摘要
Streaming large volumes of data has a wide range of real-world applications, e.g., video flows, internet calls, and online games etc. Thus, fast and real-time data stream processing is important. Traditionally, data clustering algorithms are efficient and effective to mine information from large data. However, they are mostly not suitable for online data stream clustering. Therefore, in this work, we propose a novel fast and grid based clustering algorithm for hybrid data stream (FGCH). Specifically, we have made the following main contributions: 1), we develop a non-uniform attenuation model to enhance the resistance to noise; 2), we propose a similarity calculation method for hybrid data, which can calculate the similarity more efficiently and accurately; and 3), we present a novel clustering center fast determination algorithm (CCFD), which can automatically determine the number, center, and radius of clusters. Our technique is compared with several state-of-art clustering algorithms. The experimental results show that our technique can achieve more than better clustering accuracy on average. Meanwhile, the running time is shorter compared with the closest algorithm.
引用
收藏
页码:1228 / 1244
页数:17
相关论文
共 46 条
[1]
Aggarwal C., 2004, P 30 INT C VER LARG, V30, P852
[2]
[Anonymous], 2003, P 29 INT C VER LARG
[3]
[Anonymous], 2010, P IEEE C EV COMP
[4]
SNCStream+ : Extending a high quality true anytime data stream clustering algorithm [J].
Barddal, Jean Paul ;
Gomes, Heitor Murilo ;
Enembreck, Fabricio ;
Barthes, Jean-Paul .
INFORMATION SYSTEMS, 2016, 62 :60-73
[5]
SNCStream: A Social Network-based Data Stream Clustering Algorithm [J].
Barddal, Jean Paul ;
Gomes, Heitor Murilo ;
Enembreck, Fabricio .
30TH ANNUAL ACM SYMPOSIUM ON APPLIED COMPUTING, VOLS I AND II, 2015, :935-940
[6]
Clustering data streams using grid-based synopsis [J].
Bhatnagar, Vasudha ;
Kaur, Sharanjit ;
Chakravarthy, Sharma .
KNOWLEDGE AND INFORMATION SYSTEMS, 2014, 41 (01) :127-152
[7]
Variational Inference for Dirichlet Process Mixtures [J].
Blei, David M. ;
Jordan, Michael I. .
BAYESIAN ANALYSIS, 2006, 1 (01) :121-143
[8]
Bodyanskiy YV, 2017, NEUROCOMPUTING
[9]
Density-Based Clustering over an Evolving Data Stream with Noise [J].
Cao, Feng ;
Ester, Martin ;
Qian, Weining ;
Zhou, Aoying .
PROCEEDINGS OF THE SIXTH SIAM INTERNATIONAL CONFERENCE ON DATA MINING, 2006, :328-+
[10]
Chatzis SP, 2011, FUZZY C MEANS TYPE A