On change diagnosis in evolving data streams

被引:33
作者
Aggarwal, CC [1 ]
机构
[1] IBM Corp, TJ Watson Res Ctr, Hawthorne, NY 10532 USA
关键词
data streams; evolution; change detection;
D O I
10.1109/TKDE.2005.78
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
In recent years, the progress in hardware technology has made it possible for organizations to store and record large streams of transactional data. This results in databases which grow without limit at a rapid rate. This data can often show important changes in trends over time. In such cases, it is useful to understand, visualize, and diagnose the evolution of these trends. In this paper, we introduce the concept of velocity density estimation, a technique used to understand, visualize, and determine trends in the evolution of fast data streams. We show how to use velocity density estimation in order to create both temporal velocity profiles and spatial velocity profiles at periodic instants in time. These profiles are then used in order to predict three kinds of data evolution: dissolution, coagulation, and shift. Methods are proposed to visualize the changing data trends in a single online scan of the data stream and a computational requirement which is linear in the number of data points. The visualization techniques can also be used to provide online animations which show the changes in the data characteristics while they occur. In addition, batch processing techniques are proposed in order to quantify the level of change across different combinations of dimensions. This quantification is then used in order to determine dimensional combinations with significant evolution. The techniques discussed in this paper can be easily extended to spatiotemporal data, changes in data snapshots at fixed instances in time, or any other data which has a temporal component during its evolution.
引用
收藏
页码:587 / 600
页数:14
相关论文
共 24 条
[1]  
AGGARWAL CC, 2004, P VER LARG DAT BAS C
[2]  
AGGARWAL CC, 2003, P VER LARG DAT BAS C
[3]  
AGGARWAL CC, 2004, P ACM INT C KNOWL DI
[4]  
AGGARWAL CC, 2003, P ACM SIGMOD C
[5]  
Agrawal R., 1994, P VER LARG DAT BAS C
[6]  
ANDRIENKO N, 2000, P 3 AGILE C GEOGR IN, P137
[7]  
Bonachea D, 1999, USENIX ASSOCIATION PROCEEDINGS OF THE 2ND CONFERENCE ON DOMAIN-SPECIFIC LANGUAGES (DSL'99), P163
[8]  
Brodsky BE., 1993, Nonparametric Methods in Change Point Problems
[9]  
CHAWATHE S, 1997, P ACM SIGMOD C
[10]  
CHEUNG D, 1996, P IEEE INT C DAT ENG