Noise reduction for instance-based learning with a local maximal margin approach

被引:21
作者
Segata, Nicola [1 ]
Blanzieri, Enrico [1 ]
Delany, Sarah Jane [2 ]
Cunningham, Padraig [3 ]
机构
[1] Univ Trento, Dipartimento Ingn & Sci Informaz, Trento, Italy
[2] Dublin Inst Technol, Dublin, Ireland
[3] Univ Coll Dublin, Dublin 2, Ireland
关键词
Noise reduction; Editing techniques; k-NN; SVM; Locality; NEIGHBOR; CLASSIFIERS;
D O I
10.1007/s10844-009-0101-z
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
To some extent the problem of noise reduction in machine learning has been finessed by the development of learning techniques that are noise-tolerant. However, it is difficult to make instance-based learning noise tolerant and noise reduction still plays an important role in k-nearest neighbour classification. There are also other motivations for noise reduction, for instance the elimination of noise may result in simpler models or data cleansing may be an end in itself. In this paper we present a novel approach to noise reduction based on local Support Vector Machines (LSVM) which brings the benefits of maximal margin classifiers to bear on noise reduction. This provides a more robust alternative to the majority rule on which almost all the existing noise reduction techniques are based. Roughly speaking, for each training example an SVM is trained on its neighbourhood and if the SVM classification for the central example disagrees with its actual class there is evidence in favour of removing it from the training set. We provide an empirical evaluation on 15 real datasets showing improved classification accuracy when using training data edited with our method as well as specific experiments regarding the spam filtering application domain. We present a further evaluation on two artificial datasets where we analyse two different types of noise (Gaussian feature noise and mislabelling noise) and the influence of different class densities. The conclusion is that LSVM noise reduction is significantly better than the other analysed algorithms for real datasets and for artificial datasets perturbed by Gaussian noise and in presence of uneven class densities.
引用
收藏
页码:301 / 331
页数:31
相关论文
共 85 条
[1]  
AHA DW, 1991, MACH LEARN, V6, P37, DOI 10.1007/BF00153759
[2]   Fast nearest neighbor condensation for large data sets classification [J].
Angiulli, Fabrizio .
IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2007, 19 (11) :1450-1464
[3]  
[Anonymous], P 10 INT C MACHINE L
[4]  
[Anonymous], [No title captured]
[5]  
[Anonymous], MACHINE LEARNING /
[6]  
[Anonymous], 2007, Uci machine learning repository
[7]  
Bakir G.H., 2005, ADV NEURAL INFORM PR, P81
[8]  
Bello-Tomás JJ, 2004, LECT NOTES COMPUT SC, V3155, P32
[9]  
BEYGELZIMER A, 2006, 23 INT C MACH LEARN, P97
[10]  
BLANZIERI E, 2007, 4 C EM ANT CEAS 07 M