Data discretization unification

被引:46
作者
Jin, Ruoming [1 ]
Breitbart, Yuri [1 ]
Muoh, Chibuike [1 ]
机构
[1] Kent State Univ, Dept Comp Sci, Kent, OH 44241 USA
关键词
Discretization; Entropy; Gini index; MDLP; Chi-square test; G(2) test; MIXED-MODE DATA; CONTINUOUS ATTRIBUTES; ALGORITHM;
D O I
10.1007/s10115-008-0142-6
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Data discretization is defined as a process of converting continuous data attribute values into a finite set of intervals with minimal loss of information. In this paper, we prove that discretization methods based on informational theoretical complexity and the methods based on statistical measures of data dependency are asymptotically equivalent. Furthermore, we define a notion of generalized entropy and prove that discretization methods based on Minimal description length principle, Gini index, AIC, BIC, and Pearson's X (2) and G (2) statistics are all derivable from the generalized entropy function. We design a dynamic programming algorithm that guarantees the best discretization based on the generalized entropy notion. Furthermore, we conducted an extensive performance evaluation of our method for several publicly available data sets. Our results show that our method delivers on the average 31% less classification errors than many previously known discretization methods.
引用
收藏
页码:1 / 29
页数:29
相关论文
共 38 条
[1]  
Agresti A., 1990, CATEGORICAL DATA ANA
[2]  
[Anonymous], P EUR WORK SESS LEAR
[3]  
[Anonymous], PRINCIPLES DATA MINI
[4]  
[Anonymous], 1995, P 7 IEEE INT C TOOLS
[5]  
[Anonymous], 2001, STAT INFERENCE
[6]  
[Anonymous], 1993, Proceedings of the 13th International Joint Conference on Artificial Intelligence
[7]  
[Anonymous], 1998, CLASSIFICATION REGRE
[8]  
Auer P., 1995, P 12 INT C MORG KAUF
[9]   Multivariate Discretization for Set Mining [J].
Stephen D. Bay .
Knowledge and Information Systems, 2001, 3 (4) :491-512
[10]   Khiops: A statistical discretization method of continuous attributes [J].
Boulle, M .
MACHINE LEARNING, 2004, 55 (01) :53-69