A neural network-based framework for the reconstruction of incomplete data sets

被引:71
作者
Gheyas, Iffat A. [1 ]
Smith, Leslie S. [1 ]
机构
[1] Univ Stirling, Dept Comp Sci & Math, Stirling FK9 4LA, Scotland
关键词
Missing values; Imputation; Single imputation; Multiple imputation; Generalized regression neural networks; MULTIPLE IMPUTATION; GENERAL REGRESSION; GENETIC ALGORITHM; ENSEMBLE; CLASSIFIERS;
D O I
10.1016/j.neucom.2010.06.021
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
The treatment of incomplete data is an important step in the pre-processing of data. We propose a novel nonparametric algorithm Generalized regression neural network Ensemble for Multiple Imputation (GEMI). We also developed a single imputation (SI) version of this approach-GESI. We compare our algorithms with 25 popular missing data imputation algorithms on 98 real-world and synthetic datasets for various percentage of missing values. The effectiveness of the algorithms is evaluated in terms of (i) the accuracy of output classification: three classifiers (a generalized regression neural network, a multilayer perceptron and a logistic regression technique) are separately trained and tested on the dataset imputed with each imputation algorithm, (ii) interval analysis with missing observations and (iii) point estimation accuracy of the missing value imputation. GEMI outperformed GESI and all the conventional imputation algorithms in terms of all three criteria considered. (C) 2010 Elsevier B.V. All rights reserved.
引用
收藏
页码:3039 / 3065
页数:27
相关论文
共 34 条
[1]   An empirical comparison of voting classification algorithms: Bagging, boosting, and variants [J].
Bauer, E ;
Kohavi, R .
MACHINE LEARNING, 1999, 36 (1-2) :105-139
[2]   Improving classification performance on real data through imputation [J].
Bratu, C. Vidrighin ;
Muresan, T. ;
Potolea, R. .
2008 IEEE INTERNATIONAL CONFERENCE ON AUTOMATION, QUALITY AND TESTING, ROBOTICS (AQTR 2008), THETA 16TH EDITION, VOL III, PROCEEDINGS, 2008, :464-469
[3]  
CARLO G, 2003, BIOMETRIKA, V0090, P00643
[4]   Principal component analysis of neuronal ensemble activity reveals multidimensional somatosensory representations [J].
Chapin, JK ;
Nicolelis, MAL .
JOURNAL OF NEUROSCIENCE METHODS, 1999, 94 (01) :121-140
[5]  
Chen Jiahua., 2000, J OFF STAT, V16, P113
[6]   A novel ensemble of classifiers for microarray data classification [J].
Chen, Yuehui ;
Zhao, Yaou .
APPLIED SOFT COMPUTING, 2008, 8 (04) :1664-1669
[7]   Flexible neural trees ensemble for stock index modeling [J].
Chen, Yuehui ;
Yang, Bo ;
Abraham, Ajith .
NEUROCOMPUTING, 2007, 70 (4-6) :697-703
[8]   Online adaptive policies for ensemble classifiers [J].
Dimitrakakis, C ;
Bengio, S .
NEUROCOMPUTING, 2005, 64 :211-221
[9]  
Dondeti S, 2005, ACTA CHIM SLOV, V52, P440